You're likely to get better results from vector-based semantic search though, just because it takes you beyond needing exact matches on search terms.
I've found that internal enterprise projects tend to be very keyword based, and vector search often produces weird, head-scratcher results that users hate - whereas term-based search does a better job of capturing the right terms, if you do the proper synonym/abbreviation expansions.
That said, I use them both, usually with vector search as a fallback after the initial keyword-based RAG pass
Arguably, for many use cases (e.g. searching through a document with ~200 passages), loading embeddings in memory and running a simple linear search would be fast enough.
[1]- https://towardsdatascience.com/in-context-learning-approache...
Full disclosure, I just joined Zilliz this week as a Dev Advocate.
A vector storage could help in reduce the time it takes to retrieve the most similar hit. I used faiss as a local vector store quite a bit to retrieve vectors fast. Though I had 1.5 million vectors to work through.
An in memory index is about as good as it gets for a single node performance, and fitting that many vectors into memory on a single machine is easy.
We then just fetch up to the vectors related to a customer's schema in memory (largest is ~200MB) and run cosine similarity in a few ms in Go (handwritten, ~25 lines of code), and then we've got out top N things to place in our prompt.
Primitive? You betcha. Works extremely well for our entire customer base? Yup. You definitely don't need a Vector DB unless you have an enormous amount of vectors. For us it means having to run our own Redis clusters, but we know how to do that, and so we don't need to involve another vendor.
You definitely do need information retrieval. It just shouldn't be limited to vector dbs. Unfortunately vector db companies and the VCs that back them have flooded the internet with propaganda suggesting vector db is the only choice. https://colinharman.substack.com/p/beware-tunnel-vision-in-a...
For most serious use cases, you'll have far too much data to fit into 1 (or several) inference contexts.