https://newsletter.semianalysis.com/p/tpuv7-google-takes-a-s...
https://newsletter.semianalysis.com/p/tpuv7-google-takes-a-s...
"OpenAI’s leading researchers have not completed a successful full-scale pre-training run that was broadly deployed for a new frontier model since GPT-4o in May 2024, highlighting the significant technical hurdle that Google’s TPU fleet has managed to overcome."
Given the overall quality of the article, that is an uncharacteristically convoluted sentence. At the risk of stating the obvious, "that was broadly deployed" (or not) is contingent on many factors, most of which are not of the GPU vs. TPU technical variety.
The would have taken some time to calculate the efficiency gains of pretraining vs RL. Resumed the GPT-4.5 for whatever budget made sense and then spent the rest on RL.
Sure they chose to not serve the large base models anymore for cost reasons.
But I’d guess Google is doing the same. Gemini 2.5 samples very fast and seems way to small to be their base pre train. The efficiency gains in pertaining scale with model scale so it makes sense to train the largest model possible. But then the models end up super sparse and oversized and make little sense to serve in inference without distillation.
In RL the efficiency is very different because you have to inference sample the model to draw online samples. So small models start to make more sense to scale.
Big model => distill => RL
Makes the most theoretical sense for training now days for efficient spending.
So they already did train a big model 4.5. Not using it would have been absurd and they have a known recipe they could return scaling on if the returns were justified.
It kind of explains a coding issue I had with tradingview who update their pinescript thing quite frequently. ChatGPT seemed to have issues with v4 vs v5.
Valuation isn’t available money; they'd have to raise more money in the current, probably tighter for them, investment environment to enter the TPU race, since the money they have already raised that that valuation is based on is already needed to provide runway for what they are already doing without putting money into the TPU race
The bigger issue is that entering a 'race' implies a race to the bottom.
I've noted this before, but one of NVDA's biggest risks is that its primary customers are also technical, also make hardware, also have money, and clearly see NVDA's margin (70% gross!!, 50%+ profit) as something they want to eliminate. Google was first to get there (not a surprise), but Meta is also working on its own hardware along with Amazon.
This isn't a doom post for NVDA the company, but its stock price is riding a knifes edge. Any margin or growth contraction will not be a good day for their stock or the S&P.
Of course Huang will lean on the software being key because he sees the hardware competition catching up.
Google, Meta, Amazon do “shallow and broad” software. They are quite fast at capturing new markets swiftly, they frequently repackage OpenSource core and add the large amount of business logic to make it work, but essentially follow the market cycles - they hire and layoff on a few year cycle, and the people who work there typically also will jump around industries due to both transferable skills and relatively competitive competitors.
NVDA is roughly in the same bucket as HFT vendors. They retain talent on a 5-10y timescales. They build software stacks that range from complex kernel drivers and hardware simulators all the way to optimizing compilers and acceleration libraries.
This means they can build more integrated, more optimal and more coherent solutions. Just like Tesla can build a more integrated vehicle than Ford.
Maintaining a web browser requires about 1000 full-time developers (about the size of the Chrome team at Google) i.e., about $400 million a year.
Why would Microsoft incur that cost when Chromium is available under a license that allows Microsoft to do whatever it wants with it?
And so on all under licenses that allows Microsoft do whatever it wants with?
They should be embarrassed to do better, not spin it into a “wise business move” aka transfer that money into executive bonuses.
In contrast, basically no one derives any significant revenue from the sale of licenses or subscriptions for web browsers. As long as Microsoft can modify Chromium to have Microsoft's branding, to nag the user into using Microsoft Copilot and to direct search queries to Bing instead of Google Search, why should Microsoft care about web browsers?
It gets worse. Any browser Microsoft offers needs to work well on almost any web site. These web sites (of which there are 100s of 1000s) in turn are maintained by developers (hi, web devs!) that tend to be eager to embrace any new technology Google puts into Chrome, with the result that Microsoft must responding by putting the same technological capabilities into its own web browser. Note that the same does not hold for Windows: there is no competitor to Microsoft offering a competitor to Windows that is constantly inducing the maintainers of Windows applications to embrace new technologies, requiring Microsoft to incur the expense of applying engineering pressure to Windows to keep up. This suggests to me that maintaining Windows is actually significantly cheaper than it would be to maintain an independent mainstream browser. An independent mainstream browser is probably the most expensive category of software to create and to maintain excepting only foundational AI models.
"Independent" here means "not a fork of Chromium or Firefox". "Mainstream" means "capable of correctly rendering the vast majority of web sites a typical person might want to visit".
Potentially these last two points are related.
The prosecution rests.
You must have an amazing CV to think these are shallow projects.
I’d say I have an average CV in the EECS world, but also relatively humble perspective of what is and isn’t bleeding edge. And as the industry expands, the volume „inside” the bleeding edge is exploitation, while the surface is the exploration.
Waymo? Maybe; but that’s acquisition and they haven’t done much deep work since. Tensorflow is a handy and very useful DSL, but one that is shallow (builds heavily on CUDA and TPUs etc); Android is another acquisition, and rather incremental growth since; Go is a nth C-like language (so neither Dennis Richie nor Bjarne Stroustrup level work); MapReduce is a darn common concept in HPC (SGI had libraries for it in the 1990s) and implementation was pretty average. AlphaGo - another acquisition, and not much deep work since; Kubernetes is a layer over Linux Namespaces to solve - well - shallow and broad problems; Chrome/Chromium is the 4th major browser that reached dominance and essentially anyone with a 1B to spare can build one.. gVisor is another thin, shallow layer.
What I mean by deep software, is a product that requires 5-10y of work before it is useful, that touches multiple layers of software stack (ideally all from hardware to application) etc. But these types of jobs are relatively rare in the 2020s software world (pretty common in robotics and new space) - they were common in the 1990s where I got my calibration values ;) Netscape and Palm Pilot was a „whoa”. Chromium and Android are evolutions.
I get that bashing on Google is fun, but TensorFlow was the FIRST modern end-user ML library. JAX, an optimizing backend for it, is in its own league even today. The damn thing is almost ten years old already!
Waymo is literally the only truly publicly available robotaxi company. I don't know where you get the idea that it's an acquisition; it's the spun-off incarnation of the Google self-driving car project that for years was the butt of "haha, software engineers think they're real engineers" jokes. Again, more than a decade of development on this.
Kubernetes is a refinement of Borg, which Google was using to do containerized workloads all the way back in 2003! How's that not a deep project?
Waymo is an acquihire from ‘05 DARPA challenges, and I’d say Tesla got there too (but with a much stricter hardware to user stack, which ought to bear fruits)
I’d say Kubernetes would be impressive compared to 1970s mainframes ;) Jokes aside, it’s a neat tool to use crappy PCs as server farms, which was sort of Google’s big insight in 2000s when everyone was buying Sun and dying with it, but that makes it not deep, at least not within Google itself.
But this may change. I think Brin recognizes this during the Code Red, and they start very heavily on building a technical moat since OpenAI was the first credible threat to the user behavior moat.
Come on, man.
> Google's TPUs change this equation a bit
Google has been using TPUs to serve billions of customers for a decade. They were doing it at that scale before anyone else. They use them for training, too. I don't know why you say they don't own the stack "from silicon to apps" because THEY DO. Their kernels on their silicon to serve their apps. Their supply chain starts at TSMC or some third-party fab, exactly like NVIDIA.
Google's technical moat is a hundred miles deep, regardless of how dysfunctional it might look from the outside.
They're building it for themselves and employ world-class experts across the entire stack.
How can NVIDIA develop "more integrated" solutions when they are primarily building for these companies, as well as many others?
Examples of these companies doing things you mention as being somehow unique to or characteristic of NVIDIA:
Complex kernel drivers or modules:
- AWS: Nitro, ENA/EFA, Firecracker, NKI, bottlerocket
- Google: gasket/apex, gve, binder
- Meta: Katran, bpfilter, cgroup2, oomd, btrfs
Hardware simulators:
- AWS: Neuron, Annapurna builds simulations for nitro, graviton, inferentia and validates aws instances built for EDA services
- Google: Goldfish, Ranchu, Cuttlefish
- Meta: Arcadia, MTIA, CFD for thermal management
Optimizing Compilers:
- Amazon: NNVM, Neo-AI
- Google: MLIR, XLA, IREE
- Meta: Glow, Triton, LLM Compiler
Acceleration Libraries:
- Amazon: NeuronX, aws-ofi-nccl
- Google: Jax, TF
- Meta: FBGEMM, QNNPACK
Meta builds hardware from chip to cluster to datacenter scale, and drives research into simulation at every scale, all the way to CFD simulation of datacenter thermal management.
They have the money and talent to do it. As you point out, they do have major successes in areas that take real engineering. But they also have a lot of failures. It will depend how the internal politics play out, I imagine.
Everything.
They can easily just do this for more optimized Chips.
"easily" in sense of that wouldn't require that much investment. Nvidia knows how to invest and has done this for a long time. Their Ominiverse or robots platform isaac are all epxensive. Nvidia has 10x more software engineers than AMD
Also certain companies normally don't like to do things themselves if they don't have to.
Nonetheless nvidia is were it is because it has cude and an ecoysystem. Everyone uses this ecosystem and then you just run that stuff on the bigger version of the same ecosystem.
1. there had be fixed function hardware for certain graphics stages
2. Programmable massively parallel hardware took over. Nvidia was at the forefront of this.
TPUs seem to me similar to fixed function hardware. For Nvidia it's a step backwards and even though they go into this direction recently I can't see them go all the way.
Otherwise you don't need cuda, but hardware guy's that write verilog or vhdl. They don't have that much of an edge there.
There's a lot of misleading information in what they publish, plagiarism, and I believe some information that wouldn't be possible to get without breaking NDAs
…why would I care about this in the slightest?
I was trying to make the point that SemiAnalysis is semi-famous.
That's my reply. I assume everyone who wants to know my point has access to a LLM that can summarize videos.
Is this how internet communication is supposed to be now?