If by "training corpus" you mean the actual data I already acknowledged that that's currently legally impossible.
In fact everything I just said I said in my original comment. It's like you didn't read it at all.
DeepSeek's GRPO Infrastructure, multi-stage training pipeline, and their "cold start" phase have been massively influential in LLM research.
Did you read the site this very post links to? The entire point is that the training corpus, recipe, and scripts, as well as intermediate checkpoints, will be made available for K2 Horizon. Chinese models these days don't even release pre-trained weights anymore; all you get now is the finished post-trained product.
I've read your comment. I'm doubting you've even read the thing you were commenting on.
The labs that are actually competing with frontier models are using data would usually be a violation of copyright to release openly.
> Chinese models these days don't even release pre-trained weights anymore; all you get now is the finished post-trained product.
No? That's absolutely not true. Qwen, GLM, Kimi, DeepSeek, etc all consistently release both the post-trained "Instruct/Chat" versions and the underlying "Base" (pre-trained) weights.
Which specific Chinese models are you thinking about?
The only thing they don't release is their data-filtering pipelines but they detail even that in their public-access papers.
DeepSeek is truly as open source as you can possibly legally get. Besides the data itself, it's completely reproducible by anyone else.
I don't think americans yet acknowledge just how radically transparent Chinese labs are being (and how much even the west benefits from it).