I think if you take a few moments to read carefully.
You'll see that AI companies, including openai, are generally not competing on accuracy benchmarks.
For example, here are the benchmarks on which open ai seem to be trying to compete.
MMLU: Measuring Massive Multitask Language Understanding,
MATH: Measuring Mathematical Problem Solving With the MATH Dataset,
GPQA: A Graduate-Level Google-Proof Q&A Benchmark,
DROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs,
MGSM: Multilingual Grade School Math Benchmark (MGSM), Language Models are Multilingual Chain-of-Thought Reasoners,
HumanEval: Evaluating Large Language Models Trained on Code,