At their scale, you’d have to be setting money on fire if you’re not doing dynamic inference optimisations based on load.
API and consumer subscriptions are treated differently; all trackers measuring via API won’t notice this.
At their scale, you’d have to be setting money on fire if you’re not doing dynamic inference optimisations based on load.
API and consumer subscriptions are treated differently; all trackers measuring via API won’t notice this.
Where is the evidence they are "nerfing" the models due to request volume?
Edit: I don't know they do, I mean they could repurpose systems if they are idle. Inference demand is global, and providers like Azure have global routing options that are cheaper. Night time in the USA could be serving inference demand on the other side of the globe.
But you seem adamant that there's no chance the providers serve slightly quantized models for subscription users during high loads, or otherwise tweak models for requests from those users.
It's tricky to prove either way, but the chance is not zero.
Nerfing conspiracy doesn't need to be proven false. Where is the evidence it's true?
They just want some evidence. It should be pretty easy to measure, shouldn't it?