It would be interesting to know how many optimizations of the Chinese models were incorporated back into Claude and Codex.
A month before the R1 paper came out, they released the Deepseek math paper which described their method for MoE load balancing.
If we go down the line of dead internet theory which I’m becoming more convinced of these days, the volume of information that’s not necessarily original or extracted from reality and just interpolated and extrapolated from existing information in different ways by LLMs will greatly outnumber human information coming up.
In which case these models should.. start to converge on the same data I imagine, with slightly different behaviors within that. One big generative orgy feedback loop.