The prevailing commentary on Reddit is that GPT4 has been getting progressively nerfed for weeks.
Well if the recent leaks are true it's because they're running a smaller, likely heavily quantized model to do the bulk of the generation now, with the full float one only stepping in on occasion to save GPU time.
The interesting thing is that it seems that 3.5-turbo has slightly improved in recent months, while performance of 4 has deteriorated, it's like they're converging them to a middle ground or something.
> However, these approaches would be a lot more complex and beyond the scope of this assistant's capabilities
Which seems very strange