The GPT-4 I use now feels like a shadow of the GPT-4 I used during the early access program. GPT-4, back then, ported dirbuster to POSIX compliant multi-threaded C by name only. It required three prompts, roughly:
* “port dirbuster to POSIX compliant c”
* “that’s great! You’re almost there, but wordlists are not prefixed with a /, you’ll need to add that yourself. Can you generate a diff updating the file with the fix?”
* “This is pretty slow, can we make it more aggressive at scanning?”
It helped me write a daemon for FreeBSD that accepted jails definitions as a declarative manifest and managed them in a diff-reconciliation loop kubernetes style. It implemented a binpacking algorithm to fit an arbitrary number of randomly sized images onto a single sheet of A1 paper.
Most programming tasks I threw at it, it could work its way through with a little guidance if it could fit in a prompt.
Now, it’s basically worthless at helping me with programming tasks beyond trivial problems.
Folks who had early access before it was public have commented on how exceptional it was back then. But also how terrifyingly unaligned it was. And the more they aligned the model with acceptable social behavior, the worse the model performed. Alignment goals like “don’t help people plan mass killings” seem to cause regressions in the models performance.
I wouldn’t dismiss these comments. If they’re true, it means there is a hyper intelligent early GPT-4 model sitting on a HDD somewhere that dwarfs what we’ve seen publicly. A poorly aligned model that’s down to help no matter what your request is.