https://github.com/jerryjliu/gpt_index is a particularly interesting implementation under very active development at the moment.
https://github.com/jerryjliu/gpt_index is a particularly interesting implementation under very active development at the moment.
One of the common failure modes of RAG (and I assume your technique as well) is hallucination: basically, making stuff up that isn't in any of the docs [2].
[2] https://twitter.com/sjwhitmore/status/1617318051455840258
The dot product between the query and key representations is similar to computing the cosine similarity between two vectors. The cosine similarity is a measure of the similarity between two vectors in a multi-dimensional space, and is defined as the dot product of the vectors normalized by their magnitudes.
The dot product of the query and key representations can be seen as an un-normalized version of the cosine similarity, in the sense that it computes the dot product of the two vectors. The result is a scalar value, which represents the similarity between the two vectors, the larger the scalar, the more similar the vectors are.This could be solved with using different search methodologies, using multiple gpt requests to summarize available info or using a structured knowledge framework to prepare prompts (instead of just raw text).
Any other ideas that I'm missing?
The tree is then traversed to find the most relevant chunk asking GPT to compare entries based on relevance to the question. This results in an original document chunk, which is given as context in a final prompt asking to answer the query.
This is great and powerful, but very not cost effective. Log(n) requests to completion API, for n documents.
The embedding search is probably necessary for bigger datasets.
[1] https://lukesalamone.github.io/posts/rolling-my-own-blog-sea...