I am working quite a bit with normal and chained LLMs but so far haven't explored the retrieval route.
I am working quite a bit with normal and chained LLMs but so far haven't explored the retrieval route.
Some tips:
- Most vector search is basically kNN under the hood, with some kind of compression. If you have too many embeddings in your DB, this starts to pull up irrelevant text very quickly. The key is to segment the DB using other data before doing the embedding search. Postgres extensions are good for this.
- The quality of the data you put into your embedding DB matters a lot.
- How you chunk text matters. Chunking by paragraph is much better than naive chunking, for example.
- This is a good benchmark for embedding models [2]
[1] https://www.endless.academyBy segmenting data, is this as simple as adding a further condition to the SQL? E.g. if your embeddings belonged to a "client" you would start with WHERE client_id - x?
> How you chunk text matters. Chunking by paragraph is much better than naive chunking
Would love some more info on this. SO if you had a 400 word slab of text, it would be better to create m3 or 4 embeddings from that?
In my opinion, until LLMs learn to say I don't know, this is the way to do it in accuracy-first tasks.
1. User types in the planned work (e.g., "replace discharge line on water pump")
2. The app embeds that and look for similar historical work in a vector database
3. We retrieve the most similar historical work and the safety observations that were made on those jobs
4. We summarize the major hazards or mistakes from those previous jobs using gpt-3.5-turbo and deliver that to the user