I had an agent (SOTA models, xhigh reasoning, etc.) write and optimize some CUDA code recently. They've gotten a lot better at this stuff. And yet, they couldn't (self-)realize the obvious issue with their code: it was written as if for a typical multicore CPU. What the heck was a (serial) queue doing there? "Work efficiency" as a tradeoff for stalling 20k threads. Its later proposed optimizations were all about "can we get the queue faster" rather than "maybe we should actually parallelize the work for our very-parallel processor".
Funny stuff, as if it were hell-bent on writing a paper rather than actual software.