First line of the abstract of MMLU: "We propose a new test to measure a text model's __multitask accuracy__."
Fourth line of the abstract of MATH: "To facilitate future research and __increase accuracy__ on MATH"
Second line of GPQA abstract: "We ensure that the questions are high-quality and extremely difficult: experts who have or are pursuing PhDs in the corresponding domains reach __65% accuracy__ [...] while highly skilled non-expert validators only reach __34% accuracy__"
Fifth line of the DROP abstract: "We apply state-of-the-art methods from both the reading comprehension and semantic parsing literature on this dataset and show that the best systems only achieve 32.7% F1 on __our generalized accuracy metric__"
From the MGSM paper: "MGSM __accuracy__ with different model scales."
Models are designed to output accurate information in a reasonable amount of time. That's literally the whole goal. The entire thing. A math-specific model wants to provide accurate math answers. A general model wants to provide accurate answers to general questions. That's the whole point.