China's DeepSeek Shows Why Trump's Trade War Will Be Hard to Win
bloomberg.com
bloomberg.com
Deepseek trained its DeepSeek-V3 Mixture-of-Experts (MoE) language model with 671 billion parameters using a cluster containing 2,048 Nvidia H800 GPUs in just two months, which means 2.8 million GPU hours, according to its paper. For comparison, it took Meta 11 times more compute power (30.8 million GPU hours) to train its Llama 3 with 405 billion parameters using a cluster containing 16,384 H100 GPUs over the course of 54 days.
https://www.tomshardware.com/tech-industry/artificial-intell...
In contrast, Trump's approach of having the US piss off our 4 biggest trading partners simultaneously is both terrible and terribly stupid. Even just the threats (or blathering trial-ballons of "very fine people are telling me I oughta") will still affect companies' choices of how and where to structure their supply-chains.