177 karma · joined January 10, 2019
With larger-scale real-world enterprise RAG-based applications, you soon realize the enormous time and effort required to experiment with all these levers to optimize the RAG pipeline: which vector DB to use and how, which embedding model to use, pure vector search or hybrid search, chunking strategies, and on and one...
With Vectara's RAG-as-a-service (www.vectara.com) we try to help address exactly this issue: you get an optimized, high performance, secure and scalable RAG pipeline, so you don't need to go through this massive hyper-parameter tuning exercise. Yes, there are still some very useful levers you can experiment with, but only where it really matters.
Clearly many of us see the need here. I have also been working on a similar demo: https://search-hackernews.vercel.app/ 1. Stack: Vectara for RAG, Vercel for hosting 2. Results show the main story and top 3-4 comments from the story 3. Focused mostly on the search aspect - so if you click it redirects you to the HN page itself. No summaries although it'd be easy to add.
Would love to get some feedback and any suggestions for improvement. I'm still working on this as a side project.
Example query to try: "What did Nvidia announce in GTC 2024?" (regular HN search returns empty)
https://huggingface.co/spaces/vectara/Hallucination-evaluati...
Looking forward to seeing GPT4-Turbo there soon.
Embeddings are also super important in retrieval-augmented-generation (RAG), and getting the best "embeddings model" is important to achieve the best RAG performance. At Vectara we recently launched our new Boomerang model that pushes the limit on performance on embedding models, and I hope will spur more innovation and further improvements in this space.
https://vectara.com/introducing-boomerang-vectaras-new-and-i...