Yeah, I'd expect model performance to be super spiky on SWE work, at least they admit it with the name of the model. It's distilled from an already-distilled model.
Maybe still worth it if their "64% cheaper" figure holds.
Maybe still worth it if their "64% cheaper" figure holds.
Fine tuning you have the actual model weights of the original model, you then train that model to answer in a different (or better) way.
What you're describing is just synthetic data.
Note Anthropic misused the term in their post about Chinese model distillation, deliberately I assume.