IMO the biggest problem with distillation is that not enough people are openly doing it. I would love to see more small, competitive US labs instead of having the eggs in 2~4 baskets (depending on how you count).
IMO the biggest problem with distillation is that not enough people are openly doing it. I would love to see more small, competitive US labs instead of having the eggs in 2~4 baskets (depending on how you count).
An even smaller fraction of the cost if they do it by buying AI access at as much of a discount as they can find, including black market resellers, and then reselling that access to paying users again with a proxy. As is common.
This gives ruthless "fast followers" an economic edge over the innovator that's putting in the real work.
The dynamics are very much alike to what patents and copyright law are supposed to prevent. Same type of "we took the products of your work and used them to undercut you". Except there are no laws against distillation - so most of the enforcement happens on model provider level.
Is there actually that much capability transfer from non-logit-matched distillation, or is Anthropic just another unwilling source of data?
Even the early papers on distillation techniques found that surprisingly small distillation datasets can improve task performance noticeably on some specific task types - and that valuable adaptations like SFT/RLHF instruction following can be distilled from one-hot non-logit traces.
A big part of what distillation really gets you is: paving over the mismatch between pre-training and final performance. A base model is trained to spit out fitting text, but not to instruction follow, reason autoregressively, self-check or use tool calls - like an AI has to. There is transfer straight from the "text prediction" pre-training objective, and pre-training sets the foundation for all that follows - but the capabilities you get "out of the box" with it are often unrefined and fragile. Which makes some sense - internet text doesn't often include raw chain-of-thought autoregressive reasoning. It's not the kind of thing humans tend to write.
Reasoning traces? They let an AI learn proven techniques and adaptations directly, from an AI that was already taught "how to be an AI" in other ways.
It's why this kind of distillation typically plugs into mid-training and post-training, not pre-training.
Now, I'm not saying that all Chinese companies do is eat tokens, distill and lie. That just isn't the case. They developed or refined numerous training techniques and architectural adaptations - like deep fusion for high performance visual input, RLVR with GRPO, trunked MoE, storage-efficient and bandwidth-efficient attention formulations, or residual routing techniques like AttnRes. Some of those are used widely now, and some are still on the uptake but show good promise.
But Chinese labs are enjoying massive efficiency gains from being able to distill from the frontier instead of doing things the hard way. It's a leg up. It lets them put their supply of R&D effort and RL compute elsewhere. They wouldn't be nearly as advanced if they couldn't do it.
"You're trying to kidnap what I've rightfully stolen."
The issue is doing.. exactly what OpenAI and Anthropic have done to get where they are?
No, there is no issue.
As long as labs do not heavily kneecap model outputs, practically all this applies to the training corpus as well.
AI gives ruthless users of AI a leg up over the people who's data it was trained on. "We took the products of your work and used them to undercut you". It's all the same.
The only way I'd be against distilling would be if AI models became owned by the public who's work is used to create them. Of course the AI labs should be paid well, but these models are a product of the entire world's efforts, not only the labs.
That's the moat. Mistral has the capability but not the legal protections.
Couldn't I simply give a Chinese friend my key on Open router?
I say just let them duke it out. After a decade of regulatory capture and enshittification, it’s nice to see some actual competition again.
Probably. But if OpenAI or Anthropic stole your credit card to purchase tokens you could sue them. You won't get a cent from any Chinese labs.
> it’s nice to see some actual competition again.
Competition benefits everyone. But this isn't fair competition. A German startup cannot legally do any of these tactics required to bypass Anthropic/OAI's counter-measures. Which makes EU less competitive and therefore less investment in European AI.
A German startup cannot legally do what Anthropic/OpenAI have done. And neither could Anthropic/OAI themselves. What's your point again?
I said it in my original comment. That is the moat. The reason frontier AI is a two horse race. European labs cannot gain ground because the only way to do it is illegally.
Europe gets barely any funding because it is so far behind and cannot use the illegal strategies Chinese labs use to distill Opus/Astra. Mistral cannot compete because the competition is literally breaking every law.
> That is only possible in China because any other US/EU lab doing the same would get into massive legal trouble.