edit: having trouble finding the tweet I saw recently, it might have been from their lead engineer and not the CTO.
edit: having trouble finding the tweet I saw recently, it might have been from their lead engineer and not the CTO.
But I also think it's partially a psychological phenomenon, just people getting used to the magic and finding more bad edge-cases as it is used more.
EDIT: It seems that they do claim that the layers on top also didn't change https://twitter.com/alexalbert__/status/1780707227130863674
(I think it was just the settings for how ChatGPT calls the GPT4 model, and not affecting use of GPT4 by API, though I may be misremembering.)
And they do.
my $0.02: it makes me very uncomfortable that people misunderstand LLMs enough to even think this is possible
But even F16 -> llama.cpp Q4 (3.8 bits) has negligible perplexity loss.
Theoratically, a leading AI lab could quantize absurdly poorly after the initial release where they know they're going to have huge usage.
Theoratically, they could be lying even though they said nothing changed.
At that point, I don't think there's anything to talk about. I agree both of those things are theoratically possible. But it would be very unusual, 2 colossal screwups, then active lying, with many observers not leaking a word.
Prompt engineering is surprisingly fragile.
* I don't think you are! I've looked up to you a lot over last year on LLMs btw, just vagaries of online communication, can't tell if you're ignoring the tweet & introducing me to idea of system prompts, or you're suspicious it changed recently. (in which case, I would want to show off my ability to extract system prompt to senpai :)
Thanks for liking my work! :)
Claude 3 hasn't dropped Elo on the lmsys leaderboard which supports the CTO's claim.
When GPT4 got lobotomized, you had to work hard to avoid the new behavior, it popped up everywhere. People claiming claude got lobotomized seem to be cherry picking example.