Hm. I've not had any problems using GPT4 with code that it has never seen. But sure, I'll try that...
Edit: I took the GIO Hello World sample. I fed it into the GPT-4 API.
And I swear to God, first try, first click:
https://gist.github.com/FeepingCreature/bde583aa8a2e8e5e1c56... Near-perfect success. One extra import, sure, I'll give it that one, LLMs can't correct when they notice they don't actually need an import in the code. But then it works.
Gist link cause I use OpenRouter against the GPT-4 API, maybe that's why? Maybe the "stochastic parrot" vs "oncoming apocalypse" perspective genuinely represents that ..... GPT just doesn't work for some people?? Did they fuck the ChatGPT4 web interface enough somehow so that it's just inept, but somehow not affect the API?
Edit:
> In my research, 35/37 attempts don't even compile. It tries the same mistakes over and over again. It fails to make reasonable assessment of the compilation errors.
Important note: there's a theory that if you've got it making mistakes a few times, it thinks that it's playing "the sort of AI that fucks up a lot" and starts making more mistakes, reinforcing its role. Start over fresh if it seems strangely incapable. Like, it is absolutely possible to use GPT-4 in a way that results in it systematically being incapable of the most basic tasks. These things are not reliably competent; but they're occasionally competent! In my experience, even with completely novel tasks in completely novel environments, the competence is "in there" and can often be elicited with dedicated poking. That's why I think LLMs are enough for a takeoff with enough scale (and basic online learning probably): in my opinion, it's not a matter of attaining the skills but removing the roadblocks.
It's like the joke about seeing an old man playing chess with a labrador, and the labrador just carefully picking up figures and moving them around the board, every time in a legal move, and you say "that's amazing, a chess-playing dog!" And the old man scoffs and says, "Nonsense. His endgame is rubbish." Three years ago we didn't think a dog could understand pawn promotion at all.
edit: Hey, share your repros? I'm genuinely curious what's going on here now.
edit:
> Watch how it can't even spell.
This is a specific issue with the current generation of LLMs and should be thought of more as a disorder than a fundamental inability. LLMs don't actually ever see letters, they see BPEs, which consist of one up to n letters. For instance, the word "corporation" looks to the AI more like "<27768>". So like- yes. It can not spell. It fundamentally, architecturally, cannot perceive letters in the words you give it; its ability to manually split words into letters is based on chance memorization. Instead, try getting it to split the word into letters in code.