Can’t wait to see the human verifying the results and then figure out that the AI model actually cheated and the results are not correct.
Edit: at least ~600,000 lines
https://stanfordtechreview.com/articles/openai-buckmaster-na...
[1] https://leodemoura.github.io/blog/2026-8-1-postmortem-for-ke...