Other modes of RAG that retrieve multiple times per generation (e.g. every so-many tokens) can improve performance, and intuitively feel closer to the 'sparse attention over the entire corpus' that RAG is approximating.
However, fine-tuning on relevant, high quality, knowledge-rich question/answer pairs seems dominant, when such examples are available or can be generated.