Normally RAG just sends your query `q` to a information retrieval function which searches a database of documents using full-text search or vector search. Those documents are then passed to a generative model along with your query to give you your final answer.
MemoRAG instead immediately passes `q` to a generative model to generate some uninformed response `y`. `y` is then passed to the information retrieval function. Then, just like vanilla RAG, `q` and the retrieved documents are sent to a generative model to give you your final answer.
Not sure how this is any more "memory-based" than regular RAG, but it seems interesting.
Def check out the pre-print, especially eq. 1 and 2. https://arxiv.org/abs/2409.05591
EDIT: The "memory" part comes from the first generative model being able to handle larger context, covered in Section 2.1