HNHacker News
TopNewBestAskShowJobs

ofermend

177 karma · joined January 10, 2019

Developer Relations, AI Agents, LLMs
submissionscomments
ofermend··on Gemini 3 Flash: Frontier intelligence built for speed
Gemini-3-flash is now on Vectara hallucination leaderboard, and rated at 13.5% grounded hallucination rate.

https://github.com/vectara/hallucination-leaderboard

ofermend··on Nvidia Nemotron 3 Family of Models
We just evaluated Nemotron-3 for Vectara's hallucination leaderboard.

It scores at 9.6% hallucination rate, similar to qwen3-next-80b-a3b-thinking (9.3%) but of course it is much smaller.

https://github.com/vectara/hallucination-leaderboard

ofermend··on GPT-5.2
GPT-5.2 just added to Vectara Hallucination Leaderboard. Definitely an improvement over GPT-5.1 - congrats to the team

https://github.com/vectara/hallucination-leaderboard

ofermend··on Claude Opus 4.5
Can't wait to try Opus 4.5

We just evaluated it for Vectara's grounded hallucination leaderboard: it scores at 10.9% hallucination rate, better than Gemini-3, GPT-5.1-high or Grok-4.

https://github.com/vectara/hallucination-leaderboard

ofermend··on [dead]
If you have built AI agents in the last 6-12 months you know they fail a lot.

I built this repository to be a community-curated list of failure modes, techniques to mitigate, and other resources, so that we can all learn from each other and build better agents.

Contributions encouraged!

ofermend··on [dead]
Enterprise Deep Research is like "consumer" deep research just pointed at your private data, and I think may become the "killer app" of Agentic AI for business.

Lots of valuable use-cases: compliance monitoring, sales enablement, onboarding, legal, and many others.

What use-cases would you use this for?

ofermend··on About AI Evals
One of the biggest challenges in RAG Evaluation is the assumption that you somehow can get the "source of truth" generated, specifically the set of "golden answers" (or golden chunks/documents). In practice that is extremely difficult and non scalable. Open-RAG-Eval is a new open source project that aims to address that via reference-free evaluation such as UMBRELA and AutoNuggetizer scores.

Repo: https://github.com/vectara/open-rag-eval and a nice UI to use this with: openevaluation.ai

Would love to hear feedback on this after you try it out and what you might want to see on the roadmap.

ofermend··on Trust in AI
Well, we expect AI to become AGI sometime in the future. Some say it's here, others say it's in 5 years or 50 years or whatever. So imagine AGI is here already (for sake of argument), and really has superintelligence, and will be able to have agency. How do we need to treat "it"? Over history, humans and society created mechanism to overcome distrust, and our ability to collaborate is what helped us thrive. Should we think about our upcoming "relationship" with AI from that perspective as well?
ofermend··on [dead]
RAG Evaluation is difficult, primarily because it's hard to come up with "golden answers" (or golden chunks).

We made Open-RAG-Eval to solve this - RAG Eval that only requires the question, yet provides great metrics for retrieval, generation, hallucination and citations for any RAG setup.

This was in collaboration with Jimmy Lin and his students at UWaterloo.

It has connectors to LangChain, LlamaIndex and Vectara, and hoping others can contribute more connectors to other RAG systems.

repo: https://github.com/vectara/open-rag-eval

UI for reviewing eval results: https://openevaluation.ai/

Papers: https://arxiv.org/pdf/2406.06519 and https://arxiv.org/abs/2504.15068

ofermend··on The Llama 4 herd
A great day for open source, and so glad to see llama4 out. However, I'm a bit disappointed that the hallucination rates of Llama4 are not as low as I would have liked (TL;DR slightly higher than Llama3).

Check the numbers on the hallucination leaderboard: https://github.com/vectara/hallucination-leaderboard

ofermend··on Gemini 2.5
This model is quite impressive. Not just useful for math/research with great reasoning, it also maintained a very low hallucination rate of 1.1% on Vectara Hallucination Leaderboard: https://github.com/vectara/hallucination-leaderboard
ofermend··on [dead]
It is common these days to see in large companies multiple teams developing isolated RAG applications. This is similar to the problem of "Shadow IT" back in the early cloud era - causes a big headache to IT teams.

I work at Vectara, and we see this all the time. Wondering how others are experiencing this?

ofermend··on [dead]
DeepSeek-R1 is an amazing reasoning LLM, but it seems to hallucinate more than we might expect.
ofermend··on Gemini 2.0: our new AI model for the agentic era
Gemini-2.0-Flash does extremely well on the Hallucination Evaluation Leaderboard, at 1.3% hallucination rate https://github.com/vectara/hallucination-leaderboard
ofermend··on [dead]
We've done a study (see link) that shows that - unlike common belief - semantic chunking is not always the best approach.

Curious to hear from the YC community - anyone else did systemic testing and if so what did you find?

ofermend··on IBM Granite 3.0: open enterprise models
Check out Granite 3.0 on the hallucination leaderboard: https://github.com/vectara/hallucination-leaderboard
ofermend··on [dead]
We recently launched UDF reranking as part of the RAG stack, and we think this supports a lot of interesting use-cases to go beyond simple relevance. For example, it supports ranking by distance (geo-location), by recency, and more.

I wanted to ask advice from the HN community: what are some real use-cases you have that can benefit from UDF reranking in RAG?

ofermend··on Grandmaster expelled from team chess championship after phone found in toilet
I remember the Magnus/Niemann controversy from 2023 - that was quite a drama... https://en.wikipedia.org/wiki/Carlsen%E2%80%93Niemann_contro...
ofermend··on Llama 3.2: Revolutionizing edge AI and vision with open, customizable models
Great release. Models just added to Hallucination Leaderboard: https://github.com/vectara/hallucination-leaderboard.

TL;DR: * 90B-Vision: 4.3% hallucination rate * 11B-Vision: 5.5% hallucination rate

ofermend··on Show HN: Airbyte 1.0, Marketplace, AI Assist, GenAI Support and Enterprise GA
About a year ago we launched in partnership with the Airbyte team the Vectara Destination, to help developers accelerate Generative AI applications - congrats on the Airbyte team on this great launch and looking forward to 2.0
ofermend··on Show HN: Vectara-Agentic
Oh, and there's a demo of an AI assistant for hacker news here: https://huggingface.co/spaces/vectara/hacker-news-chat
ofermend··on Llama 3.1
I'm excited to try it with RAG and see how it performs (the 405B model)
ofermend··on Mistral NeMo
Congrats. Very exciting to see continued innovation around smaller models, that can perform much better than larger models. This enables faster inference and makes them more ubiquitous.
ofermend··on [dead]
I'm happy to share that today we are releasing the Mockingbird LLM.

At <10B parameters it's an LLM trained to provide optimal results for RAG and structured outputs. Although significantly smaller (and thus faster) than GPT-4 or Gemini-1.5 Pro, it performs at a comparable level for generation, citations, and structured outputs.

ofermend··on Show HN: I made a search engine for Hacker News
Thank you everyone for the feedback. For those interested in asking questions about articles, here's another nice demo (still in beta): an Agentic RAG chatbot demo (hosted on Huggingface, using streamlit for UI):

https://huggingface.co/spaces/vectara/hacker-news-chat

Feel free to ask it some things and let me know how it works.

ofermend··on Show HN: I made a search engine for Hacker News
This should be fixed now. Thanks for the find.
ofermend··on Show HN: I made a search engine for Hacker News
Yes, it's just about 6 months back. If requested by folks here, we can certainly crawl back more years - this was just the first crawl I did.
ofermend··on Show HN: I made a search engine for Hacker News
For sure. Will test that.
ofermend··on Show HN: I made a search engine for Hacker News
Thank you - these are great suggestions. Will work to add these...
ofermend··on Show HN: I made a search engine for Hacker News
Good find. let me check why that occurs.
Page 1 of 3Next →