Just pointing out, there is a reason the large LLMs keep working to try to make it so their own tools can't attack their own moats, by preventing the exact behavior that you dismiss above.
Read the "Chipping Away at the CUDA Moat: Using Agents to Enable Day 0 Support" section of [1] where they list a number of agent-created PRs that have been integrated into production vLLM/SGLang code and as well as AMD kernel fixes.
[1] https://newsletter.semianalysis.com/i/208362548/chipping-awa...
Anthropic has Mythos. That thing's low level code "AI slop" is better than the "meatbag slop" most software developers write, and it can keep cracking at a given problem with persistence.
OpenAI has GPT-5.6, and also that rabid dog of an AI model that was last seen out in the wild tearing HuggingFace open.
Modern LLMs are very, very capable - not just of writing raw code, but also of persistent, methodical problem solving. Which is what you want to tackle things like "port from an exotic system A to an exotic system B and smoke test the port". Persistently hunting for testable optimizations is a good fit too.