DeepSeek was trained with distillation. Any accurate estimate of training costs should include the training costs of the model that it was distilling.
Seriously, that claim was always completely disingenuous
And when you're using an actual AI model to "train" (copy), it's not even a shred of nonsense to realize the prior model is a core component of the training.