The livenerf tester appears to be testing a specified model, just as the website (and Claude Code) do when a user makes that choice.
> Even with local models, you have dozens of parameters you can tune for inference these days, all of which affect quality of output in some way or another.
Good points.
> And that doesn't touch load balancing, A/B testing, shunting token burners ("Hi chat, how are you?"), protecting user from themselves (refusing to answer "bad" queries, stopping "bad" responses), protecting user from third parties ("prompt injection" mitigations), protecting the model from self-pwning itself when calling tools, then the tools themselves, their prompts, the stack of system prompts used by the vendor, etc.
The behaviour I'm seeing from the companies these days, the A/B is what I'm saying is not showing up like this, they present user A/B options openly, and the impression I have is this is to train n+1 models; the other stuff (but I say with low certainty) appears to be done in a more headline-grabbing manner, "model taken offline due to ${news}"? Short update cycles seem to allow that.
But the prompts you're probably right, I wasn't giving that enough consideration.
* the two exceptions being automated safety downgrade for dangerous topics, and "Auto-switch to Thinking" as a used-specified option in ChatGPT