> No benchmarks, no info on which models are used, […]
The benchmarks are here → https://echo.tracerml.ai/eval/
They are not good benchmarks but at least they exist.
The benchmarks are here → https://echo.tracerml.ai/eval/
They are not good benchmarks but at least they exist.
In my project, I wasted a huge amount of time trying to improve GPQA Diamond results above ~93% range. I realized my mistake when Fable dropped and made no improvement on this benchmark vs. Opus.