You'll see that AI companies, including openai, are generally not competing on accuracy benchmarks.
For example, here are the benchmarks on which open ai seem to be trying to compete.
MMLU: Measuring Massive Multitask Language Understanding,
MATH: Measuring Mathematical Problem Solving With the MATH Dataset,
GPQA: A Graduate-Level Google-Proof Q&A Benchmark,
DROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs,
MGSM: Multilingual Grade School Math Benchmark (MGSM), Language Models are Multilingual Chain-of-Thought Reasoners,
HumanEval: Evaluating Large Language Models Trained on Code,