I get the hype, even if I don’t necessarily agree with it. TLDR: it’s technically impressive in training infrastructure and a geopolitical surprise.
As others commenters said, it compares favorably against American Flagship models in benchmarks. This is geopolitically interesting since the model is Chinese, and subject to trade restrictions and America thinks of itself as the world’s best model builders.
What makes it interesting technically, is the trade restrictions required them to focus on efficiency and reusing old hardware. This made it wildly cheaper to run the training. The They make a ton of very low-level optimizations to make training efficient. This is impressive engineering, and shows how a small amount of effort can free up a nvidia-lock-in. They had to totally bypass a lot of CUDA libraries and drop down even lower to control the GPUs. Nvidia has been capturing a huge fraction of industry wide AI spend, and surely no one wants that. This is, IMO, the actual part to watch. I suspect it’ll drive a new wave of efficiency-focused LLM tools which will also unlock competitors GPUs.
They also had some novel-for-LLMS training techniques but it’s suspected that the big AI companies elsewhere are doing it now too, but not disclosed. (Mostly reinforcement learning).
What I think is hype, meanwhile, is the actual benchmarks. Most models are trained on their competitors output. This is just a reality of the industry. Especially non-flagships being trained against flagships data. DeepSeek was almost certainly trained against OpenAI models, so it makes sense that it would approach the output quality. That’s very different from being capable of outperforming or “taking the lead” or whatever. We’ll need to wait longer and see how the future goes to make that determination. “China” has long had a great history of machine learning tech, so there is no reason to think that it’s structurally impossible for a Chinese organization to be on the leading edge, but it has to happen before we can say it happened.
What is also hype is calling this a “side project” of a financial firm. The firm spun it out as a dedicated company. China cracked down on hedge funds, so the company looked for ways to repurpose the talent and compute infrastructure. This isn’t some side project during lunch breaks for a few bored quants.
PS, thinking models are very different in use-case than normal models. It’s great for tasks with a “right answer” like math problems. Consider a simple example, in calculus, your teacher made you “show your work”, and doing so certainly reduced the likelihood of errors. That’s the purpose of a thinking model, and it excels at similar tasks.