Certainly anyone who knows something about inference is going to speculate, looking at the change in token pricing (and particularly how the % drop in cache read pricing is much larger than the % drops in pricing for other token types), that there is some kind of KV cache optimization behind these newer models. But even if is true, I don't think anyone can say with certainty what it may be. It may be the labs making their own innovations (they have some very smart people, and this is probably an area where having unlimited pre-release access to frontier LLMs like Fable and Astra gives an additional research edge), it may indeed be the direct application of Chinese labs' methods, or it may be some combination of the two. Sure, it is fun to speculate about, but beyond the facts of the token pricing changes and the increased inference speed, it's just speculation. The certainty the author displays here is not very helpful.
The author also seems to have a bit of an axe to grind agains the US labs, judging by the tone. I think that detracts from the discussion too.
They also seem confused about why Chinese labs have released these optimizations recently. Well, you have to release them (with or without explanation) if you are going to release an open weight architecture, and that is what the Chinese labs have been doing for a long time. Sure, there are reasons behind that to discuss too, but this isn't exactly new.
So, this is an interesting topic to think about, and the Chinese labs do indeed deserve credit for some very clever new attention and inference techniques, but I would read it with a skeptical eye.