HNHacker News
TopNewBestAskShowJobs

zmccormick7

233 karma · joined February 5, 2018

RAG/LLM enthusiast
submissionscomments
zmccormick7··on Show HN: Runner – the anti-vibe coding agent
Good to know. I've heard great things about Context7, but haven't experimented with it yet.
zmccormick7··on Show HN: Runner – the anti-vibe coding agent
As in the download itself didn't happen when you clicked the download button, or the installation failed?
zmccormick7··on Show HN: Runner – the anti-vibe coding agent
Cool, I hadn't heard of Traycer. That does look quite similar!

Completely agree. I basically built Runner to codify the process I was already using manually with Claude Code and Gemini. A lot of developers seem to be settling on a similar workflow, so I'm betting that something like Runner or Traycer will be useful for a lot of devs.

I'll be curious to see how far the existing players like Cursor push in this direction.

zmccormick7··on Show HN: Runner – the anti-vibe coding agent
I agree that's a major problem. It's not something I've solved yet. I suspect a web research sub-agent is likely what's needed, so it can pull in up-to-date docs for whatever library you need to work with.
zmccormick7··on Show HN: Runner – the anti-vibe coding agent
Gemini is required for the context management sub-agent. You can use any of OpenAI, Anthropic, or Gemini for the main planning and coding agents, but GPT-5 performs the best in my experience. Claude 4 Sonnet works well too, but it's about twice as expensive.
zmccormick7··on Show HN: MinDB – an extremely memory-efficient vector database
The main thing we need to add is metadata filtering, as that's required for a lot of use cases. We're also thinking about adding hybrid search support and multi-factor ranking.
zmccormick7··on Show HN: MinDB – an extremely memory-efficient vector database
We've only done full benchmarking with the FIQA dataset, comparing minDB with Chroma. We're going to try it with Qdrant and Weaviate soon too, since they both have support for quantization, which will be a more apples-to-apples comparison with our approach.

We did test uploading and querying a Wikipedia dump, which was ~35M vectors. Query latency was around 150ms and peak memory usage was 1.5GB. We couldn't test recall, though, because we didn't have queries with ground truths.

zmccormick7··on Solving the out-of-context chunk problem for RAG
Agreed that thresholds don't work when applied to the cosine similarity of embeddings. But I have found that the similarity score returned by high-quality rerankers, especially Cohere, are consistent and meaningful enough that using a threshold works well there.
zmccormick7··on Solving the out-of-context chunk problem for RAG
Agreed. Retrieval performance is very dependent on the quality of the search queries. Letting the LLM generate the search queries is much more reliable than just embedding the user input. Also, no retrieval system is going to return everything needed on the first try, so using a multi-step agent approach to retrieving information is the only way I've found to get extremely high accuracy.
zmccormick7··on Show HN: I made a search engine for Hacker News
I had the same issue when searching for specific companies/products. It feels like a pretty basic vector search with no hybrid search component or reranking.
zmccormick7··on Show HN: SpRAG – Open-source RAG implementation for challenging real-world tasks
That should work well with the default parameters, so you shouldn't have to do anything special.
zmccormick7··on Show HN: SpRAG – Open-source RAG implementation for challenging real-world tasks
That description is a little vague, so I need to improve that. The use cases we're focused on are ones with 1) dense unstructured text, like legal documents, financial reports, and academic papers; and 2) challenging queries that go beyond simple factoid question answering. Those kinds of use cases are where we see existing RAG systems struggle the most.
zmccormick7··on Show HN: SpRAG – Open-source RAG implementation for challenging real-world tasks
I think the difference is that they're building for end users, not developers.
zmccormick7··on Show HN: SpRAG – Open-source RAG implementation for challenging real-world tasks
Agreed. I've gotten a lot of feedback along those lines today, so that's my top priority now.
zmccormick7··on Show HN: SpRAG – Open-source RAG implementation for challenging real-world tasks
So the point of AutoContext is so you don't have to do that two-step process of first finding the right document, and then finding the right section of that document. I think it's cleaner to do it this way, but it's not necessarily going to perform any better or worse. But then spRAG also has the RSE part which is what identifies the right section(s) of the document. Whether or not that helps in your case is going to depend on how good of a solution you already have for that.
zmccormick7··on Show HN: SpRAG – Open-source RAG implementation for challenging real-world tasks
That's a great question. I'll start with a little context: most of the users of our existing hosted platform are no-code/low-code developers who choose us because we're the simplest solution for building what they want to build (primarily because we have end-to-end workflows like Chat built-in). The improved retrieval performance is a nice-to-have for this group, but usually not the primary reason they choose us.

For the larger companies and venture-backed startups we've talked to, they almost universally want to own their RAG stack in-house or build on open-source frameworks, rather than outsource it to an end-to-end solution provider. So open-sourcing our core retrieval tech is our bid to appeal to these developers.

zmccormick7··on Show HN: SpRAG – Open-source RAG implementation for challenging real-world tasks
I think spRAG should be pretty well suited for that use case. I think the biggest challenge will be generating specific search queries off of more general user inputs. You can look at the `auto_query.py` file for a basic implementation of that kind of system, but it'll likely require some experimentation and customization for your use case.
zmccormick7··on Show HN: SpRAG – Open-source RAG implementation for challenging real-world tasks
That's great feedback. I actually went back and forth between those two descriptions. I agree that "dense text, like financial reports and legal documents" is more precise. Those are the kinds of use cases this project is built for.

I want to keep this project tightly scoped to just retrieval over dense unstructured text, rather than trying to build a fully-featured RAG framework.

zmccormick7··on Show HN: SpRAG – Open-source RAG implementation for challenging real-world tasks
In our AutoContext implementation, the document title gets included with the generated summary. So if you have files that are organized into nested folders with descriptive names, you can input that full file path as the `document_title`. I did this with one of our internal benchmarks and it worked really well.
zmccormick7··on Show HN: SpRAG – Open-source RAG implementation for challenging real-world tasks
Those are just the defaults, and spRAG is designed to be flexible in terms of the models you can use with it. For AutoContext (which is just a summarization task) Haiku offers a great balance of price and performance. Llama 3-8B would also be a great choice there, especially if you want something you can run locally. For reranking, the Cohere v3 reranker is by far the best performer on the market right now. And for embeddings, it's really a toss-up between OpenAI, Cohere, and Voyage.