DeepSeek and the Effects of GPU Export Controls
vincentschmalbach.com
vincentschmalbach.com
Even so, if we look at Groq / Cerebras, the fastest LLM inference companies:
They're both based on architectures that are 7nm+, and so architectures that China can produce locally despite the export restrictions.
Ultimately, the export controls are mainly just inconvenience. Not a real blocker.
The Chinese don't need to achieve state of the art chip manufacturing to achieve SOTA AI outcomes.
They just need to make custom silicon specialized for the kinds of AI algorithms they want to scale.
Of course, at scale, that's going to mean that the US should eventually have both lower production costs, and energy use in consumer use of AI models, and that Chinese products will likely be more dependent on the cloud for at least the near future.
The whole strategy seems ultimately meh in a long term sense... Mainly good for building up a sense of mutual enmity and dividing the world ... Which is also going to result in higher cost of living around the world as trade falters.
Sad stuff.
What do you mean by fastest LLM inference companies? Is there a leaderboard for this?
To be determined. PRC can shift compute costs by building cheaper energy. Also arguable they can drive hardware costs to be competitive - % of manufacturing cost of hardware is ~30-40%, PRC save 100s of billions not paying western IP and service fees with indigenous hardware depending on how much central gov wants to squeeze margins. Would not be surprised if they can roll out 4x-8x more 14nm for compute parity and still have cost advantage once they get domestic fabs up at scale.
China can produce much cheaper electronics that can compete even when they aren't as powerful as NVIDIA's
> Novel training frameworks
Where can one find more information about these? I keep seeing hand-wavy language like this w.r.t. DeepSeek’s innovation
You can find more papers from the attached author: https://arxiv.org/search/cs?searchtype=author&query=DeepSeek... or title https://arxiv.org/search/?query=DeepSeek&searchtype=title&ab... and go through citations for more.
Of course, you could just search by some of the attached authors as well. Daya Guo, the lead author for the R1 paper has 36 papers on Arxiv: https://arxiv.org/search/cs?query=Guo%2C+Daya&searchtype=aut...
Besides the papers, DeepSeek has an active Github https://github.com/deepseek-ai and https://huggingface.co/deepseek-ai
I will concede that I have not read all their papers or looked through their code, but that's why I asked the question: I hoped someone here might be able to point me to specific places in specific papers instead of a axvix search.
How is that useful?
I would assume nothing, similarly to how exports of western tech from western countries somehow magically exploded overnight to Russia's neighbors and everyone is pretending not to notice because it makes money.
However there is nothing stopping some company setting up a company in a third country, funding it indirectly and getting them to build a cluster for deepseek/others to access.
After all the location of the servers isn't really an in surmountable problem, so long as the training data is closeby.
I suspect these export restrictions are less black-and-white than you imagine. If Nvidia shipped a lot of GPU's to, say, Brazil, and they ended up being rented to american startups for AI, all would be fine.
But if those same GPU's in Brazil ended up rented to Chinese companies who used them to make state of the art models, then Nvidia would get a big fine and the datacenter would magically catch fire[1].
[1]: https://www.elinfor.com/news/asml-supplier-is-caught-in-a-fi...
They definitely are, but things like Golden Sentry and Blue Lantern (amongst other Dual Use Monitoring regimes) can also still look for these sorts of uses. But yes, there's lots of examples of "Country X can't do Y, so we go to country Z and work with them to do Y" sorts of bypasses. Still increases the amount of work required if they want something NATSEC related to work on.
Oh I agree wholeheartedly. But that also is a risk, because if its not binary its much harder to prove that you didn't know the ultimate destination.
There are direct sales to high volume customers, like meta, aws, HP, Dell et al, then there are sales to large resellers. Those will all do due diligence checks to make sure that its not going to china.
Then there is the channel which are smaller players that either buy smaller volumes (ie a few thousand) or buy from resellers. whilst nominally those companies are vetted and audited, there are lots of them, so its perfectly possible to buy from lots of channel partners and avoid suspicion
And overall, the controls on sales are being expanded.
The simpler option for them realistically, is not so much that they buy the latest GPUs, but rather that they manage to use them on western cloud services.
The US is looking to track money flows in greater detail to see if funds are ultimately coming from China, but that's some majorly invasive stuff, and not entirely easy to implement on the scale of the planet.
Once you have trained models, inference is almost always less of a hassle.
China has a graphics processor company thats apparently good enough to land it on an entity list.
https://en.wikipedia.org/wiki/Moore_Threads
The sheer number of Chinese companies the US as entity listed for export controls is comical as its basically a blacklist of the entire PRC's tech sector.
export controls work well to do one thing: create a US competitor. china already fabs domestic 3nm chips. theres no reason to think they wont emerge as a serious competitor to NVidia.
[1] https://www.cnas.org/publications/reports/secure-governable-...
[2] https://www.iaps.ai/research/location-verification-for-ai-ch...
https://exportcontrol.lbl.gov/a-bigger-yard-a-higher-fence-u...
You just described the core of problem of export controls in general in the 21st century: they are freakishly hard (expensive) to enforce, especially for high dollar density goods (dollar/volume).
If you try to control how many millions of barrels of crude travels from one place to another, because of the volume, that's still somewhat feasible (although in the case of the Russian embargo: still does not work).
If you try to prevent a few crates packed to the gills with H100's from eventually reaching china ... LOL.
The only actual effect of embargos is that it raises the price of the item at final destination, not a cut-off of the supply.
And when a nation-state the size of China is ultimately financing you because of how strategic what you're doing is, I suspect that is not a real problem.
A secondary effect of the export controls is obviously something similar to a Streisand effect: it's going to accelerate innovation around GPUs in China, producing the exact opposite of the initial intent.
Famous historical instances of the problem:
https://history.stackexchange.com/questions/8093/when-did-th...
Should have been obvious but now somehow isn't?
It's similar to Mixtral having gotten good performance while not having anywhere near OpenAI class money / compute.
Is this described in the paper or was this inferred from the model itself ?
Just curious, especially if the latter.
If there’s news of DeepSeek behaving badly and I missed it, then I take that back, but AFAIK they are at or near the top of the rankings on being good actors.
According to who? You?
What does AGI mean today?
What it used to mean, we passed by a while back.
The old definition in the Research community was that AGI would be able to formulate solutions to novel problems that had not been defined by the programmers. Thus, the “general” intelligence. We’re talking about mouse level intelligence. That’s what AGI meant.
LLMs have demonstrated broad problem solving capabilities across domains, and are capable of making inferences about things and developing a type of internal world model, all by encoding and parsing the cultural-linguistic framework recorded by humanity. We’re way past the old mark.
Now, it seems, the qualification for AGI has been expanded to require:
1 a kind of agency
2 Superhuman accuracy and width/depth of knowledge
3 Vastly superhuman capacity to maintain thousands of simultaneous conversations with thousands of pages of context
4 A level of artistic proficiency, at least in imitation of artistic style
So what does AGI mean now? In 1990, GPT4 would have been called a “limited super-intelligence” in the parlance of the day. Hell even an 8b model could have hit that mark, based on the breadth of accessible knowledge and ability to reason alone.
I would venture to say that my uncle bob, or 10,000 uncle bobs, operating terminals in a pungent call center somewhere, would be deemed “not AGI yet” by current standards, and would be a hell of a lot less useful than an api for deepseek r1 32B.
So, do humans of bog average intelligence not qualify as General intelligences? Or is it their capacity to do other things besides the call center job that makes them GI?
If it’s the second option, how should we measure “call center intelligence” differently?
We are using a benchmark we obviously have no metric for, at all.