The semantics they are trained on include latent representations of the world that are superior to any pre-existing PL corpus. That is my guess and it's not any less rigorous than your guess.
I think I get your point but I think it's reductionist to the point of being incorrect. LLMs must be better at some semantics than others. Programming languages don't have random semantics, they have what matches the world and what matches our languages and so on. And the current frontier LLMs aren't so generic that they can predict any phrase no matter the quality of the content and grammar. Concretely I mean some languages are harder to reason about (predict) than others.
Cheap shot so take it for what it's worth: Maybe the tests measure the wrong things in 2026. Also maybe the tests pre-suppose that the entrenched format of schooling is right.
The high score of Gemini 3.8 Flash vibes with my experience anecdotally. While it often goes off the rails with open-ended questions (which is a strength if taken with care), it is also a good at solving issues in a well-defined environment like an enterprise codebase.
Like you are hand crafting a C++ game engine because you like to, I am crafting a category theory based foundation of accounting in various odd languages. Mine is made possible by AI because I'm only an accountant, not a mathematician. So your article is inspiring, humane, and I appreciate you sharing it.