I've been seeing lots of demos like this, especially with games. Yes, they are impressive, but I'm surprised nobody is talking about some of the things I've noticed with GPT-6: the code it writes, by default, is surprisingly messy and obviously unmaintainable. It's also very...not human. No human would write code like it does. I get the sense that the optimization of these models on benchmarks is causing their behavior to morph into a very brute-force type of approach. It makes you think: when code is suddenly very cheap, does good style and organization even matter as much? Or were those just important for humans? (I think they still do, for the record).