/s
For 2022 and 2023, Microsoft bought a significant portion of NVIDIA's available hardware. They spent quite a bit of 2023 trying to figure out how to even power the multiple fleets of GPUs. Just now with the mild to expected wild adoption of Azure OpenAI are they getting around to servicing all their (potential) customers.
Seriously, this is am outlandish claim just from looking at Microsoft and Nvidias market cap.
I am sure that Microsoft is gonna be one of Nvidias largest customers, but I sincerely doubt it's even a double digit percentage of their revenue.
To reach double digit revenue of NVIDIA's 2023 at $26.97 billion, you'd only need to hit ~$2.7B in sales.
H100's are priced anywhere between $20k - $35k, so required to purchase ~77k - ~135k units.
That is singularly H100s, Microsoft also offers lower compute, and they have the rest of Azure to service with a variety of solutions.
Being at #1 or #2 market cap worldwide is not a farfetched position to be a significant controller of chips, especially since they directly work in the space as a platform.
.. but is that true?
MSR has been putting out research in all derivatives of modern large neural network architectures (NLP, CV, etc.) for the same amount of time that Google has. If there was a drift between timelines, its not large IMO.
What you could argue is that Google historically was more successful in their research outputs.
However, historical consumption of resources may not compare to current resources consumption.
> I doubt we have the visibility to know how they compare in terms of available flops and the unit costs
Completely agreed, unfortunately, this is all guesswork at best
> Using Pierce-Arrow Motor Car Company as an example of such success is historically inaccurate. Pierce-Arrow was an American automobile manufacturer based in Buffalo, New York, which was known for producing luxury cars. It was indeed a dominant and prestigious brand in the early 20th century. However, the company did not manage to maintain its success and ultimately failed to adapt to changing market conditions. It faced financial difficulties during the Great Depression and eventually went bankrupt in 1938. Pierce-Arrow's inability to forecast and adapt to the economic changes and shifts in consumer preferences of the time led to its decline.
Training and designing LLMs doesn't mean you understand the semiconductors business.
Perhaps they could design a core and license it out? I'm trying to come up with a way they can do something significant without 100 people. Just the memory and serial connections are complex enough ignoring the GPU or heat/power issues.
What is essentially happening in my opinion is technical innovation has slowed so silicon valley is seeking money to prop up a house of cards that doesn't make much new that is useful or needed.
Can anyone specifically say what trillions of dollars invested in "AI" would buy for society?
It seems to me there are so many higher priorities.
A typical case of engineer's disease.
Perhaps they'll pull off an Apple (for ARM) and do their own architecture (either for training/tuning or inference) that will have a significant effect on the industry, but it seems unlikely. They haven't hired the right people.
The real advantage they might have is insight into how the algorithms can be adapted to reduce power consumption/latency while improving performance. It would seem odd to me, if there weren't more than an order of magnitude in new algorithms for LLMs. You're not going to get 10x the transistors or speed from silicon, but you might get an efficient architecture for a significant algorithmic improvement (that might not just be CUDA).
Is that true? I can't find anything suggesting it is. In fact, the little I can find suggests you are incorrect. I'll link them for the sake of referencing sources but they're both pretty awful ad-ridden sites...
A 2016 Tech Radar interview [0] with Norm Jouppi has him quoted as saying:
> [The] Tensor Processing Unit (TPU) is our first custom accelerator ASIC [application-specific integrated circuit] for machine learning [ML], and it fits in the same footprint as a hard drive.
And a 2023 Tom's hardware post [1] begins:
> Google has made significant progress in its endeavor to develop its own data center chips, according to a new report. The Information says that a key milestone has just been reached, which means that Google can plan to roll out server systems powered by the new chips starting from 2025.This is not the first processor that Google has successfully put through R&D - the company has previously made an ASIC for servers and an SoC for mobile devices. The search giant started using its internally developed Tensor Processing Unit (TPU) as far back as 2015.
[0]: https://www.techradar.com/news/computing-components/processo...
[1]: https://www.tomshardware.com/news/google-reaches-self-develo...
1/ https://www.wired.com/2012/03/google-microsoft-network-gear/
2/ I believe they had a few custom chips designed for the youtube workloads that predate the TPU.
I remember in 2010 there was a building in MV that focused on custom chips.
Even granted that OpenAI are not able to build a chip that is competitive with NVidia's latest GPUs for running LLMs right away (which is an opinion - not backed by any direct evidence, but I agree that it is plausible as they are going up against a lot of prior R&D) is it not possible that:
a) The unit economics could be so much better that the result is still a major win, e.g. 50% of the performance at 20% of the price.
b) OpenAI is decoupled from existing supply constraints and is able to grow faster and deliver more value. A "worse" chip that you can actually get (in insane volume) may be strategically better than a "superior" chip that is limiting your growth.
c) That the plan might include some elements you are not expecting - at the $trillions investment level they might be looking at doing some surprising things e.g. (I am just making this up but there are a lot of possibilities) buy a memory manufacturer and work directly on increasing memory bandwidth.
The idea that even with expertise, the wins would be so much over what other companies that have hired/bought these companies have been designing for the last 10 years based on very similar requirements (the ones that wrote so much of the foundational research) also seems implausible.
c) It's not actually possible to plan investments at that level with anything more than a very vague direction you're aiming. If it is long term, then everything is changing in unpredictable ways before you get even 25% there, but if you throw so much money at the problem in order to try to solve it much more quickly you are disrupting global economic and geopolitical forces in ways that also can't be planned for.
It seems more likely to me they'd get 20% of the performance at 50% of the price, and that might still work out for them if it allows them to scale faster without being bottlenecked on supply of existing GPUs. But there's no magic bullet here.
They also still need to source a bunch of other stuff, like RAM, even if they can source their own processors.
What it tells me is that Altman seems to believe that OpenAI can only make the next step if they can throw even more compute at the problem but that that isn't feasible at today's prices.