If the code isn't readable to humans, there's no particular reason to think it's going to be magically readable by LLMs either.
If the code isn't readable to humans, there's no particular reason to think it's going to be magically readable by LLMs either.
You might argue they're still capped at the "best" quality seen in the input. Not so. Take typos. Human text has a certain base rate of typographical errors. LLM output contains almost none. Why? Because there are many more ways to be wrong than right. LLMs are not just averaging machines, they also denoise. That should also give you pause.
The view that they are only statistical prediction machines is becoming increasingly disconnected from their current abilities.
I probably should have put the word 'easily' before readable. After all, if it is valid code, it can be read.
Also, I'm really suspicious of any result that comes from OpenAI. AFAIK frontier labs hire people to solve problems for the AI to train the AI, not to mention it scraping the internet. I think it's easily possible that a human might have already provided a 90% of the solution and AI filled in the gaps. I don't know that for sure, but I do not trust these companies published results at all. I'm only interested in 3rd party independent results.
The problem was unsolved, it solved it when it was not possible for it to be trained on the solution, therefore your other claim that "they're always going to produce the most "average" code" does not hold up, they can come up with better code than they were trained on precisely because they can apply known techniques to write better code than existed in their training data. The book you have is out of date and does not apply to recent improvements in the field.
The unit distance problem had no known answer and no roadmap. The model identified the problem, chose its own approach, and produced the proof.
I am not convinced by either your logic or your argument from incredulity. Neither prove your position.