Not exactly, they use a small but long-context model that has the whole dataset in its context (or a large part of it) to generate the chunks as elements of the reply, before passing those to the final model. So the retrieval itself is different, there is no embeddeding model or vector db.
Agreed. In the RAG space there are a million Open Source projects on GitHub all calling it memory, recreating the same thing over and over again.
There's a lot there about the generative model ("Memory Models") in the paper, so perhaps I've misrepresented it, but generally speaking yah I agree with you. It doesn't sound like a fundamental change to how we think about RAG, but it might be a nice formalization of an incremental improvement :)