2 karma · joined November 22, 2024
STT -> Ambient Retrieval(Moss) -> LLM [+ Tool calls -> On-Demand Retrieval(Moss)] -> TTS
Now STT, TTS and LLM output generation are fixed cost and independent of data scales. In reality, a typical landing page and public-facing website content will range from 100's of docs (for startups) to 100K's of docs (for enterprises).
Moss's retrieval stack runs sub-10 ms with the following internal benchmarks -
- P99 of ~5.4 ms for 100K docs in a shared container
- P99 of ~4 ms for 1M docs in a dedicated VM
our R&D team is cranking it to 200M+ docs with sub-10ms promise but sky is the limit for our scale.
We love the idea of integrating MOSS with VitePress - it's exactly the kind of high-performance, client-side experience where on-device semantic search could shine. If anyone here is connected with the VitePress maintainers or community, we'd appreciate an introduction! We'd be happy to collaborate or contribute if there’s interest in exploring an integration together.