A GPT-4 talk on youtube by personnel from Microsoft has documented this phenomenon with the 'Tikz Unicorn' evolution shown in the GPT-4 technical paper. The model gets qualitatively better with more training, and then degrades when trained to be safer (against racism sexism, etc), but it is not entirely clear why. These would seem very unrelated, especially when considering work done in LM editing (ROME/MEMIT) and the decent localization of knowledge seen there.
So, perhaps both the "I'm sorry I can't..." and 'strange errors' are not entirely orthogonal.
I find capitalism idiotic and broken, but I’m rarely allowed to say it, even if many people secretly agree with me, it might mean I’m a “communist” :)
Reasoning has degraded. To the point it was sometimes weirdly losing context and hallucinating...
Like its brain got fatigued or something...
The API is where it's at. There are wrappers on it that create the same chat look and feel, that can run on vercel or other very low cost providers, some with simpler UI, some with more features,some replicating the UI exactly.
I was under the impression that it was mostly GPU vram based but once the model is loaded, it could produce output quickly? I'm probably over-simplifying things...
The latest gpt-3.5-turbo model generates very quickly and cheaply (in part to some recently-discoverd optimization techniques... older versions cost 10x more). While the required hardware to run GPT-4 is currently unknown, it generates considerably slower on average and its much higher cost points to a higher hardware cost.
And this is per request. It's bananas.
[0] https://www.servethehome.com/chatgpt-hardware-a-look-at-8x-n...
Just highlighting the tactic :)