>Aren’t Anthropic’s models not just distilling down other people’s work?
Can you elaborate on that? I mean my direct answer would be no, of course not. But why do you think frontier models are distilled? I think maybe there is an equivocation over the word “distillation.”
Frontier labs train on their own pretraining data, human feedback, synthetic data, and research. A distilled model is specifically optimized to reproduce another model's behavior.
Meanwhile R1-Distill-Qwen-32B was distilled from DeepSeek-R1.
If you want to say a frontier model is "distilled" from the world's data and R1-Distill-Qwen-32B is distilled from DeepSeek-R1 then you are equivocating two very different things.