that's because they are distilling the frontier models
And the Chinese have been a huge source of innovation in the field.
They have been a source of innovation but probably not in training them.
It was much easier when companies had models on the /completion style APIs, because you could actually get the logits for each generation step, and use that as a dataset to fit your model to.
That isn't to diminish the efforts of the Chinese developers though, they are great.
My intuition that one need ALOT api credits to distill such large models.