I wonder what their official explanation for this behavior is.
I wonder what their official explanation for this behavior is.
(Now, if TFA is actually measuring reasoning tokens, that's quite different! It's not entirely obvious to me how he is measuring.)
What is actually stopping these model companies from running a model at full capacity on release then once its name rings out, start serving users quantized garbage?
...I mean, if they were actually doing this despite saying that they don't—promising one product and delivering something else—I think that would be fraud, no?
And, maybe it's one thing to secretly defraud normies like us (although class action lawsuits do exist), but I don't think major enterprises or the US military would take too kindly to it.
Sorry there for the smarminess but fraud is just a standard business practice these days and fines are the cost of doing business.
And I really am all for someone suing these companies forcing discovery so we can see how the sausage is made and how many eyeballs are in it.
The reputational hit, if this was to be confirmed, would also be massive. And I do think it would leak! Some employee would say something.
nothing on the fine print tells you what the weights are, you're just getting Fable 5, whatever that is
> What is actually stopping these model companies
You can say this about any company in the world, selling anything.
It's trivially measurable, and there are people running the same benchmark on the leading models every day and measuring if they degrade. Spoiler: they don't.
But you can always say "the conspiracy goes higher", and that the companies know about these daily benchmarks and are routing them to "quality" envs.
As far as the API goes, it would be really obvious. I run a small service that uses LLMs extensively, and if a model suddenly dropped in performance it would be straightforward for us to prove it. We regularly run comparisons where we generate completions with alternative models to e.g. see if we could get away with using cheap models for easy cases, if the baseline outputs deteriorated it would be all over our metrics.
Anthropic Walks Back Policy That Could Have ‘Sabotaged’ AI Researchers Using Claude https://www.wired.com/story/anthropic-responds-to-backlash-o...
But it's still happening: https://github.com/anthropics/claude-code/issues/81759
Do you know why nobody outside the companies knows what's going on? Because they sell a black box with magic inside while steadfastly refusing to tell you if they are pushing buttons on said box while it is running.
Can you imagine how much fraud would exist in the gambling industry if the gambling commission didn't exist at all? Everytime an industry is unregulated and has high costs of entry the entities in the industry abuse their customers. The incentives are much too high for them not to.
This still leaves an absurd amount of wiggle room for arguments like "oh no, our evals show that this quantization has no detectable effect on performance (in the eval distribution) therefore running the quant doesn't degrade quality"
At least that's their explanation. Either way, it wasn't a good look for "vibecoding" but it got brushed over.