What's most impressive here is that that DeepSeek keeps showing how much of a performance improvement can come from post-training alone while the architecture stayed the same. It's a strong reminder that we're probably underestimating how much optimization is still left after pretraining