HNHacker News
TopNewBestAskShowJobs

ofermend

177 karma · joined January 10, 2019

Developer Relations, AI Agents, LLMs
submissionscomments
ofermend··on Show HN: I made a search engine for Hacker News
Currently it's indexing the content of stories and comments. Are you suggesting also to index the content of the main story link (which is outside HN)?
ofermend··on Systematically Improving Your RAG
Building RAG can be easy for a simple example, but it's much more nuanced than you might think when you try to do it at larger scale.

With larger-scale real-world enterprise RAG-based applications, you soon realize the enormous time and effort required to experiment with all these levers to optimize the RAG pipeline: which vector DB to use and how, which embedding model to use, pure vector search or hybrid search, chunking strategies, and on and one...

With Vectara's RAG-as-a-service (www.vectara.com) we try to help address exactly this issue: you get an optimized, high performance, secure and scalable RAG pipeline, so you don't need to go through this massive hyper-parameter tuning exercise. Yes, there are still some very useful levers you can experiment with, but only where it really matters.

ofermend··on Show HN: Hacker Search – A semantic search engine for Hacker News
Super cool.

Clearly many of us see the need here. I have also been working on a similar demo: https://search-hackernews.vercel.app/ 1. Stack: Vectara for RAG, Vercel for hosting 2. Results show the main story and top 3-4 comments from the story 3. Focused mostly on the search aspect - so if you click it redirects you to the HN page itself. No summaries although it'd be easy to add.

Would love to get some feedback and any suggestions for improvement. I'm still working on this as a side project.

Example query to try: "What did Nvidia announce in GTC 2024?" (regular HN search returns empty)

ofermend··on Snowflake Arctic Instruct (128x3B MoE), largest open source model
This model is great. Jumped immediately to 2nd place on HHEM leaderboard: https://github.com/vectara/hallucination-leaderboard
ofermend··on Ask HN: How does AI transform UI design?
Yes agree there, and that was my question too: more focused on how design might change due AI and ChatGPT coming in. My colleague Deryk wrote a nice article about it too: https://vectara.com/blog/user-interfaces-for-ai-applications...
ofermend··on Pg_vectorize: Vector search and RAG on Postgres
have you tried Vectara?
ofermend··on Pg_vectorize: Vector search and RAG on Postgres
Totally agree - the "R" in RAG is about retrieval which is a complex problem and much more than just similarity between embedding vectors.
ofermend··on Claude 3 model family
Exciting to see the competition yield better and better LLMs. Thanks Anthropic for this new version of Claude.
ofermend··on Launch HN: Danswer (YC W24) – Open-source AI search and chat over private data
Nice to see yet another open source approach to LLM/RAG. For those who do not want to meddle with the complexity of do-it-youself, Vectara (https://vectara.com) provides a RAG-as-a-service approach - pretty helpful if you want to stay away from having to worry about all the details, scalability, security, etc - and just focus on building your RAG application.
ofermend··on Gemma.cpp: lightweight, standalone C++ inference engine for Gemma models
Awesome work on getting this done so quickly. We just added Gemma to the HHEM leaderboard - https://huggingface.co/spaces/vectara/leaderboard, and as you can see there its doing pretty good in terms of low hallucination rate, relative to other small models.
ofermend··on Gemma: New Open Models
Gemma-7B (instruction tuned version) is now on the Vectara HHEM leaderboard, with 100% answer rate and 7.5% hallucination rate. Pretty good for a model with 7B params.

https://huggingface.co/spaces/vectara/leaderboard

ofermend··on Show HN: I built a vector database API on Cloudflare
Somewhat related: I actually think vector databases are not as important as people are led to believe. Read more here: https://vectara.com/blog/vector-database-do-you-really-need-...
ofermend··on Show HN: I built a vector database API on Cloudflare
Curious to hear what criteria each team considers important (top-3) for choosing the VDB they chose. There are so many vector databases available and in my experience it's actually not the most critical component in the overall GenAI/RAG architecture, although it gets the most attention.
ofermend··on Show HN: Chat-focused RAG with automated memory management
I think DSPy is more a framework to build LLM programs; this is more of chat functionality as a service.
ofermend··on Ask HN: Challenges with RAG
What about data types? Are your RAG pipelines mostly using text data from structured data, from document stores, or unstructured data like PDF files, website content and the like? Or maybe enterprise applications like Salesforce, HR, project management, etc?
ofermend··on Ask HN: Challenges with RAG
So simplicity and ease-of-use. I assume you mean for builders/developers here, right?
ofermend··on Ask HN: Challenges with RAG
We do see that hallucination varies between LLMs https://huggingface.co/spaces/vectara/leaderboard
ofermend··on Ask HN: Challenges with RAG
Yes I agree it's complex. So large companies with large tech teams and expertise can handle this and will likely build teams to develop and maintain their own RAG pipelines. I believe simplifying this to non-experts is a necessary next step. I'm curious though about more specific challenges - what is the most difficult part of building RAG applications? Is it educating yourself about the various components (embeddings, vector databases, LLM, prompts), is it scaling, data ingest, security, data privacy? Something else?
ofermend··on Ask HN: Challenges with RAG
Do you think there is a level of hallucination that may be "acceptable" for an enterprise deployment? if so - what would that be?
ofermend··on [dead]
We've updated the Hughes Hallucination Evaluation Model (HHEM) with results about Phi-2. TL;DR: slightly better than Mixtral-8x7B
ofermend··on LLMs by Hallucination Rate
We just updated the Hughes Hallucination Evaluation model (HEM) with the results of using the latest (beta) version of Palm 2. TL;DR - much better than before.
ofermend··on Claude 2.1
Really impressed with the progress of Anthropic with this release. I would love to see how this new version added to Vectara's Hallucination Evaluation Leaderboard.

https://huggingface.co/spaces/vectara/Hallucination-evaluati...

ofermend··on New models and developer products
Excited to see GPT4-Turbo and longer sequence lengths from OpenAI. We just released Vectara's "Hallucination Evaluation Model" (aka HEM) today https://huggingface.co/vectara/hallucination_evaluation_mode... (along with this leaderboard: https://github.com/vectara/hallucination-leaderboard). GPT-4 was already in the lead. Looking forward to seeing GPT4-Turbo there soon.
ofermend··on GPTs: Custom versions of ChatGPT
Excited about GPT4-Turbo and longer sequence lengths. Looking forward very much for faster inference. We just released Vectara's "Hallucination Evaluation Model" (aka HEM) today https://huggingface.co/vectara/hallucination_evaluation_mode..., with a leaderboard: https://github.com/vectara/hallucination-leaderboard GPT-4 was already in the lead.

Looking forward to seeing GPT4-Turbo there soon.

ofermend··on Embeddings: What they are and why they matter
First time I ran into embeddings was with word2vec and could not resist showing that, similar to the "king - man + woman ~ queen", it also the case that "yoda - good + evil ~ vader". It's also cool that the semantic meaning is the same in vector space, regardless of the language.

Embeddings are also super important in retrieval-augmented-generation (RAG), and getting the best "embeddings model" is important to achieve the best RAG performance. At Vectara we recently launched our new Boomerang model that pushes the limit on performance on embedding models, and I hope will spur more innovation and further improvements in this space.

https://vectara.com/introducing-boomerang-vectaras-new-and-i...

ofermend··on [dead]
I compared the embedding models of OpenAI, Cohere and Vectara using Llama_Index for an end-to-end RAG question-answering flow. Here are the results.
ofermend··on Show HN: Boomerang, a new embedding model for RAG and semantic search
Using Boomerang can significantly improve your end-to-end RAG performance: retrieving the most relevant facts (or chunks) matters, a lot!
ofermend··on A Comprehensive Guide for Building Rag-Based LLM Applications
RAG is a very useful flow but I agree the complexity is often overwhelming, esp as you move from a toy example to a real production deployment. It's not just choosing a vector DB (last time I checked there were about 50), managing it, deciding on how to chunk data, etc. You also need to ensure your retrieval pipeline is accurate and fast, ensuring data is secure and private, and manage the whole thing as it scales. That's one of the main benefits of using Vectara (https://vectara.com; FD: I work there) - it's a GenAI platform that abstracts all this complexity away, and you can focus on building your application.
ofermend··on LLMs, RAG, and the missing storage layer for AI
Yes totally agree with that (and other comments below). Moving from a toy example to production deployment requires all the things we are used to having in robust/mature products like postgres.
ofermend··on Bob and Juliet? Rag vs. Finetuning an LLM
Agreed. Utilizing the power of LLMs with RMs (retrieval models) can be much more powerful, and I expect RAG implementations to progress in that direction in the coming years.
← PreviousPage 2 of 3Next →