The "RAG" part (over the training set) is by far the smallest contribution to the performance gains reported (see the ablation study in section 5.2). I don't think the model is actually learning in-context from the selected samples, but rather is continuing to be better conditioned to sample from the right part of the pre-training distribution here, which does a slightly better job when the samples are on topic (vs. random)