I imagine some people have their own personal in depth benchmarks they could do this for.
I imagine some people have their own personal in depth benchmarks they could do this for.
Those software settings are being changed in real-time by an algorithm and those algorithms are being tweaked and A/B tested daily by the ~~performance~~ revenue optimization teams. On the hardware side the footprint a particular model is running on is materially changing, growing or being re-distributed across DCs ~weekly.
Your benchmark didn't get 100 ? It's normal, it's not deterministic, and also your harness is wrong, and also you didn't do it when US users were offline, and also you got it wrong, and also we don't care about your results, the hivemind is speaking louder than you (also our bots are spamming more than you and drowning you out).
This very website has, at all times, a group of people saying "<Previous model> was never good enough for coding, but <current model> is the best thing and a game changer!" while the other goes "<current model> bad, <previous model> was better!". It's all vibes.