Normally RAG just sends your query `q` to a information retrieval function which searches a database of documents using full-text search or vector search. Those documents are then passed to a generative model along with your query to give you your final answer.
MemoRAG instead immediately passes `q` to a generative model to generate some uninformed response `y`. `y` is then passed to the information retrieval function. Then, just like vanilla RAG, `q` and the retrieved documents are sent to a generative model to give you your final answer.
Not sure how this is any more "memory-based" than regular RAG, but it seems interesting.
Def check out the pre-print, especially eq. 1 and 2. https://arxiv.org/abs/2409.05591
EDIT: The "memory" part comes from the first generative model being able to handle larger context, covered in Section 2.1
It would be interesting to see a performance comparison, it certainly seems the most relevant one (that or an ablation of their "memory model" with the LLMs upon which they are based).
Section 2.2 of the paper[1] goes into this in more detail. They pretrain the draft model using the redpajama dataset, followed by a supervised fine-tuning step. The training objective "aims to maximize the generation probability of the next token given the KV cache of the previous memory tokens".
This suggests that any model with long context and good retrieval performance could do the same job (and maybe better in the case of the SOTA frontier models).
Not sure how this is any more "memory-based" than regular RAG, but it seems interesting.
I can't remember where I read this joke, but as a self-proclaimed Cognitive Engineer I think about it every day: "An AI startup's financial evaluation is directly proportional to how many times they can cram 'mind' into their pitch deck!"ollama and langchain can do something simimlar.