177 karma · joined January 10, 2019
It scores at 9.6% hallucination rate, similar to qwen3-next-80b-a3b-thinking (9.3%) but of course it is much smaller.
We just evaluated it for Vectara's grounded hallucination leaderboard: it scores at 10.9% hallucination rate, better than Gemini-3, GPT-5.1-high or Grok-4.
I built this repository to be a community-curated list of failure modes, techniques to mitigate, and other resources, so that we can all learn from each other and build better agents.
Contributions encouraged!
Lots of valuable use-cases: compliance monitoring, sales enablement, onboarding, legal, and many others.
What use-cases would you use this for?
Repo: https://github.com/vectara/open-rag-eval and a nice UI to use this with: openevaluation.ai
Would love to hear feedback on this after you try it out and what you might want to see on the roadmap.
We made Open-RAG-Eval to solve this - RAG Eval that only requires the question, yet provides great metrics for retrieval, generation, hallucination and citations for any RAG setup.
This was in collaboration with Jimmy Lin and his students at UWaterloo.
It has connectors to LangChain, LlamaIndex and Vectara, and hoping others can contribute more connectors to other RAG systems.
repo: https://github.com/vectara/open-rag-eval
UI for reviewing eval results: https://openevaluation.ai/
Papers: https://arxiv.org/pdf/2406.06519 and https://arxiv.org/abs/2504.15068
Check the numbers on the hallucination leaderboard: https://github.com/vectara/hallucination-leaderboard
I work at Vectara, and we see this all the time. Wondering how others are experiencing this?
Curious to hear from the YC community - anyone else did systemic testing and if so what did you find?
I wanted to ask advice from the HN community: what are some real use-cases you have that can benefit from UDF reranking in RAG?
TL;DR: * 90B-Vision: 4.3% hallucination rate * 11B-Vision: 5.5% hallucination rate
At <10B parameters it's an LLM trained to provide optimal results for RAG and structured outputs. Although significantly smaller (and thus faster) than GPT-4 or Gemini-1.5 Pro, it performs at a comparable level for generation, citations, and structured outputs.
https://huggingface.co/spaces/vectara/hacker-news-chat
Feel free to ask it some things and let me know how it works.