while here it outperforms Fable by a significant margin:
but if the latter is true, will people still say it was "distilled" from Fable?
while here it outperforms Fable by a significant margin:
but if the latter is true, will people still say it was "distilled" from Fable?
It is a capable small model, but it’s not frontier level. The interesting part will be seeing the model size, how it responds to quantization, and how fast it runs on the kind of non-server hardware that we can buy without selling a kidney.
They claim an "independent community benchmark" (pass-fail evaluation on tasks) here: https://oxalpha.com/ox-alpha-vs-fable-5
Source: https://twitterwebviewer.com/?tweet=2091116504787935350
Many people and even software engineers fall for this all the time.
Most of these people are from crypto pivoting to AI doing this.
AI has made this easier and cheaper and it is going to get a LOT worse.
Imagine lots of websites with typosquatting and looking exactly the same as another website, vibe coded and cloned within seconds.
The public have no chance.
This one that is listed is designed to rank on Google as an informational source (although unofficial and not from z.ai which is why I said it is phishing)
Assuming you are technical you are able to discern this, imagine the average person.
No chance.
Yes, and its these sorts of websites I am asking about.
> This one that is listed is designed to rank on Google as an informational source (although unofficial and not from z.ai which is why I said it is phishing)
Again, what is there to phish?
The number one problem in LLM benchmarking is that people try to draw conclusions from sample sizes far too small to conclude anything but "it works sometimes, it fails sometimes, hard to say which is better." (The number two problem is that people run benchmarks blindly without checking that they measure something meaningful.)
> 'GAD consistently surpasses standard sequence-level distillation, delivering superior generalization and achieving performance that rivals the proprietary teacher. These results validate GAD as an effective and robust solution for black-box LLM distillation.'
No RL, although I'm a little bit surprised to see MS Research publishing a paper on distilling GPT5?
> We reserve 500 samples of LMSYS-Chat-1M-Clean as the primary test set. We also include test datasets consisting of a 500-sample subset split from Dolly [6], the 252-sample SelfInst dataset [37], and the 80-question Vicuna benchmark [3] to evaluate out-of-distribution generalization. We report the GPT-4o evaluation scores [45, 10], where GPT-4o first generates reference answers and then scores the output of the student model against them. We also conduct human evaluations on the LMSYS-Chat-1M-Clean test set for qualitative assessment.
The metric used there is me screaming at my screen per operating hours.
Does it matter? IMO not really. Weights are open after all. (Or.. soon at least for 5.3)
And critically, like contracts in general, Anthropic's terms of service is only binding upon the user/counterparty. So even if a company say specifically sought out 'claude-like' content, and claude code traces available on the internet, if they don't use the Anthropic platform there is no ToS claim.
One thing that I think matters a lot for the non developers, all three of us (sol, oxa, and me) usually agree that oxa's write up is far far better. It explains the situation very well, and has great structure for its write ups. Sol gets the job done, but it's terrible at re-explaining the problem for humans, at laying out information. It also doesn't show it's thinking, so it's imo a terrible peer to work with!