I point it to example snippets and webdocumentation but the code it gens won't work at all, not even close
Opus4.6 is a tiny bit less wrong than Codex 5.4 xhigh, but still pretty useless.
So, after reading all the success stories here and everywhere, I'm wondering if I'm holding it wrong or if it just can't solve everything yet.
Obviously it cannot. But if you give the AI enough hints, clear spec, clear documentation and remove all distracting information, it can solve most problems.
What you're doing is more specialized and these models are useless there. It's not intelligence.
Another NFT/Crypto era is upon us so no you're not holding it wrong.
That's also how you can get the LLM to do stuff outside of the training data in a reasonably good way, by not just including the _what_ in the prompt, but also the _how_.
That sort of GPU code has a lot of concepts and machinery, it’s not just a syntax to express, and everything has to be just right or you will get a blank screen. I also use them differently than most examples; I use it for data viz (turning data into meshes) and most samples are about level of detail. So a double whammy.
But once I pointed either LLM at my own previous work — the code from months of my prior personal exploration and battles for understanding, then they both worked much better. Not great, but we could make progress.
I also needed to make more mini-harnesses / scaffolds for it to work through; in other words isolating its focus, kind of like test-driven development.
(Don’t get mad at me, I’m a webshit developer)
Such as:
Adding fine curl noise to a volumetric smoke shader
Fixing an issue with entity interpolation in an entity/snapshot netcode
Find some rendering bugs related to lightmaps not loading in particular cases, and it actually introduced this bug.
Just basic stuff.
I'm not saying it's the hardest thing but I also wouldn't consider it trivial.
On the plus side, I got to see first-hand how Postgres handles deadlocks and read up on how to avoid them.
Xhigh can also perform worse than High - more frequent compaction, and "overthinking".
I'm not accusing anyone of foul play and I don't have financial interests in either company, but it feels like "something" within Code Claude/Anthropic models is optimizing to make you spend more tokens instead of helping you complete the task.
All of the major models have been getting worse lately, not just Opus.
Can't wait, I need to buy some RAM for my local model server.
One of these is better.