HNHacker News
TopNewBestAskShowJobs

srimalireddi

2 karma · joined November 22, 2024

submissionscomments
srimalireddi··on Show HN: We put voice agent on our website, learned retrieval isn't bottleneck
Thank you!
srimalireddi··on Show HN: We put voice agent on our website, learned retrieval isn't bottleneck
You asked the right question that's blocking many people from productionizing this kind of solution on their website. If we break down the anatomy of the Voice Agent, it looks like this

STT -> Ambient Retrieval(Moss) -> LLM [+ Tool calls -> On-Demand Retrieval(Moss)] -> TTS

Now STT, TTS and LLM output generation are fixed cost and independent of data scales. In reality, a typical landing page and public-facing website content will range from 100's of docs (for startups) to 100K's of docs (for enterprises).

Moss's retrieval stack runs sub-10 ms with the following internal benchmarks -

- P99 of ~5.4 ms for 100K docs in a shared container

- P99 of ~4 ms for 1M docs in a dedicated VM

our R&D team is cranking it to 200M+ docs with sub-10ms promise but sky is the limit for our scale.

srimalireddi··on Show HN: We put voice agent on our website, learned retrieval isn't bottleneck
Yes it does! It all boils down to the retrieval speed and quality. And Moss is primarily built for this purpose which is now powering the Founding Agent.
srimalireddi··on Show HN: We put voice agent on our website, learned retrieval isn't bottleneck
Yes, founding agent is powered by Moss’s sub-10 ms retrieval under the hood. Typical retrieval systems can take anywhere between 200-500 ms per turn which kills the experience of live conversation. With the help of Moss, we are able to make Founding Agent converse like human.
srimalireddi··on Show HN: Moss – AI-Powered Semantic Search Running In-Browser (No Cloud)
We provide MOSS as a lightweight TypeScript/WASM library that developers can drop directly into their applications. It's designed for easy frontend integration - no backend services or Docker setups required. With just a few lines of setup, you can instantly index your multi-modal data and start running real-time semantic search entirely on-device.

We love the idea of integrating MOSS with VitePress - it's exactly the kind of high-performance, client-side experience where on-device semantic search could shine. If anyone here is connected with the VitePress maintainers or community, we'd appreciate an introduction! We'd be happy to collaborate or contribute if there’s interest in exploring an integration together.

srimalireddi··on Show HN: Moss – AI-Powered Semantic Search Running In-Browser (No Cloud)
MOSS benefits any team or developer who wants to add "ask anything" semantic search and personalization without standing up extra backend services. That includes `indie devs`, `small SaaS teams`, `mobile apps`, and `privacy-conscious or regulated products` where sending data to the cloud is problematic. By bundling the entire stack - embeddings, vector DB store, and retrieval - into a lightweight on-device module, you avoid months of backend integration, reduce ongoing costs, and unlock private, instant UX even offline. Beyond search, we see this enabling per-user personalization without the privacy trade-offs of cloud-based systems.