Distillation is also a broad term - I think most specifically, it refers to training a smaller model on a larger/better model's full output token distribution rather than normal pretraining, which only can access the next token in the data that was actually used.
It's also used to describe the SFT bootstrapping for posttraining, which is what people generally refer to as Chinese labs "distilling".
I would almost guarantee that smaller US frontier models (ex Luna/Sonnet) are distilled from their respective large models.