500B params performing worse than other OSS of the same size is pretty meaningless if no one will use it.
> That only means their training regime is inferior if their predecessors did so much more with so much less
Hard to imagine how that wouldn’t be the case. They probably missed the boat on distilling Claude (or their lawyers said no), they probably didn’t hire an army of math PhDs to write reasoning traces, they don’t have millions of DAUs in a coding agent to train from, and they probably have less money, less experience, fewer top tier researchers, and fewer resources for experiments. They are an underdog without a doubt.
None of that means they shouldn’t release their model.
Says who? We know Grok does at the least. They admitted it openly.
Alternative explanation is that the Chinese have far more technical talent than anyone else, along with the infra and capital to build out these models.
My money is on the latter explanation, tbh.