you can host the relay and mailbox servers yourself as well
685 karma · joined April 18, 2022
you can host the relay and mailbox servers yourself as well
https://www.bloomberg.com/news/articles/2026-06-17/microsoft...
Bytedance which runs China’s most popular Doubao AI chatbot; is spending $70B in CapEx this year, most of it outside of China (Malaysia, Thailand, Brazil, etc; and they are allowed to lease NVIDIA chips). This is roughly 50% of Microsoft CapEx.
https://www.tomshardware.com/pc-components/gpus/chinas-byted...
Crazy
Image generation and veo models I’d imagine quite effective for creators; new Instagram accounts with AI content that are garnering millions of followers in spans of weeks are quite common now
Hahahaha, this is like saying, the world wars didn’t impact Europe, because it also impacted the whole world! Europeans, the war didn’t happen!! Anyways… this entire thread is more evidence that European stereotypes are valid for the most part
You cannot be kidding right? Those that remember the Apollo missions will undoubtedly agree their country is different, first but not least they are most likely using a smartphone assembled and designed with technology unimaginable by NASA planning the Apollo missions; not only that, the smartphone is assembled half way around the world by a country previously in such dire poverty and famine that over 30M died due to Marxist central planning.
Your comment is again another anecdote confirming European stereotypes. It’s not a “trap”, it’s a different world view.
On industrial infrastructure
On technology innovation
On internet regulation
On central planning
Otherwise, your comment becomes an anecdote supporting the common stereotypes (assuming you’re from Europe).
Feels like a Llama 4 type release. Benchmarks are not apples to apples. Reasoning effort is across the board higher, thus uses more compute to achieve an higher score on benchmarks.
Also notes that some may not be producible.
Also, vision benchmarks all use Python tool harness, and they exclude scores that are low without the harness.
It is possible, just may be not in the U.S.
Note: given renewables can't provide base load, capacity factor is 10-30% (lower for solar, higher for wind), so actual energy generation will vary...
Exhibit 1: Tariff revenues to bail out American farmers: https://www.ft.com/content/0267b431-2ec9-4ca4-9d5c-5abf61f2b...
[1]: https://www.cecc.gov/agencies-responsible-for-censorship-in-...
$100k is a big pizzo (protection fee)!
[1]: https://www.bloomberg.com/news/articles/2025-09-19/ted-cruz-...
> “That’s right outta ‘Goodfellas,’ that’s right out of a mafioso going into a bar saying, ‘Nice bar you have here, it’d be a shame if something happened to it,’” Cruz said, using the iconic New York accent associated with the Mafia.
This must be what AI hype actually is. Complete incoherent language to explain a very straight forward concept.
This is just: LLMs judging intermediate node outputs, and reverse traversing the graph while doing so until it modifies the original prompt.
Dario makes an astounding implicit assumption:
- China originating labs cannot acquire chips providing 80-90% similar utility without the US within the next 2-3 years.
I'll make an observation, re: DeepSeek's incentives that drove them to create the innovations from the V2 and V3 papers.
DeepSeek, compared to American AI labs, are much more compute constrained, but in a unique way. Their chips are more memory bandwidth constrained (depending on type anywhere from 50% to 80% less bandwidth).
Therefore, each dollar/hour of investment towards memory optimization is worth MORE to DeepSeek than to American labs.
In the V2/3 paper, they've demonstrated exactly that with these memory optimization techniques.
1. MLA -> reduces KV cache by nearly 80% compared to GQA. By the way, this was published in V2 in May 2024.
2. FP8 matmul (while still accumlating in FP32 gradients) without losing significant quality.
3. DualPipe scheduling and reworking of Hopper SM's allocation on communication vs. computation -> DeepSeek's V3 paper has 2 full pages of hardware suggestions for "hardware designers" (read NVIDIA)
Export controls in a global market create different incentives in parties. The resulting incentives will change, and agents (using it as an traditional economics term) will change their capital allocation strategy.
1. ChatGPT data is widely on the internet, just google Sharegpt dataset and you can scrap 200k+ conversations with a few stroke of huggingface commands. These were then used by the open source community like Vicuña models, there was a period of several months in the open source community where RLAIF was all the rage; so this data populated the internet. So if a company is crawling and scraping the internet, this will eventually be in the dataset.
2. The v3 deepseek model was trained on 15T tokens. Please educate yourself and calculate how long (in latency, inference for 1k token output will take almost 30seconds) and cost it would be to extract 15T tokens from ChatGPT / Azure API. Granted API accounts all have spend limits, and will trip fraud detection on OAI billing, how long would the subterfuge had to take place? With which model? At what time? Wouldn’t they have to keep repeating this for subsequent generation of OAI models?
3. OAI didn’t invent MLA, they didn’t invent multi token prediction with disconnected ROPE, they didn’t invent FP8 matmul training dynamics (while accumulating in FP32) without losing significant quality.
So go away
or the primary source twitter thread: https://x.com/jiayi_pirate/status/1882839370505621655
If your portfolio is green, you can still be a poor performer.
The rate of ML progress is spectacularly compute constrained today. Every step in today’s scaling program is setup to de-risked the next scale up, because the opportunity cost of compute is so high. If the opportunity cost of compute is not so high, you can skip the 1B to 8B scale ups and grid search data mixes and hyperparameters.
The market/concentration risk premium drove most of the volatility today. If it was truly value driven, then this should have happened 6 months ago when DeepSeek released V2 that had the vast majority of cost optimizations.
Cloud data center CapEx is backstopped by their growth outlook driven by the technology, not by GPU manufacturers. Dollars will shift just as quickly (like how Meta literally teared down a half built data center in 2023 to restart it to meet new designs).
Bryan is in his last year of high school.
</end>
Keep building!