Then why are they (US frontier models) still so far ahead whenever I test them against the latest Chinese models? No bias here, I'd love them to be better for my own personal gain, but I haven't seen it
Same with local LLMs, I'd love to use them for my day-to-day software engineering, and I'm not exactly GPU poor, then people with 12GB VRAM try to convince me their local setup is perfectly fine running latest Qwen and it does real engineering but whenever I try, they're a far cry from what Codex+GPT 5.x would do.
Only way to be sure is creating your own private benchmarks and use those, and the difference in quality becomes very apparent, very quickly, for your specific use cases.