This is interesting, but unless I'm misreading the paper, it looks like they're training an LLM on the corpus. I can easily see why that would result in better performance than an off-the-shelf embeddings model, but... it won't work for a corpus that changes frequently, since you'll constantly have to retrain the LLM. That's sort of the point of RAG: how do you get the right information into an LLM as context, for data that changes so frequently that you can't directly train on it?