This was my takeaway as well. None of the other retrieval methods this paper benchmarks against are specifically trained on the corpus. I think it would be more fair to compare "Self-Retrieval" against models which have been fine-tuned on the corpus.