We can’t be critical of the datasets used for testing if we don’t know what’s in them.
The same applies to training data - it should be disclosed.
The same applies to training data - it should be disclosed.
TriviaQA: https://huggingface.co/datasets/trivia_qa
HotPotQA: https://huggingface.co/datasets/hotpot_qa
Those two are used in most of the question-and-answer based metrics for LLMs in academic publications.
The same applies for other datasets used for other metrics. HuggingFace has most of if not everything of any relevance.