Sam Altman Seeks Trillions of Dollars to Reshape Business of Chips and AI
wsj.com
wsj.com
Nobody in the AI chip hypespace seems to understand this, it’s just stupid money running around trying to eat Nvidia’s margins. Sam Altman understands this less than plenty of people.
It’s becoming harder for me to see him as anything besides someone who is very talented at growing power, but not much else. Perhaps he will succeed in misallocating a trillion dollars along the way.
ROCm 6 was out Dec 16 (2023), 5.5 was May (2023). 5 was Feb 10 (2022). 4 was Dec 19 (2020)
And if you fuck that part up in any one of a dozen places, no one will use it, because the adoption cost is too high, or your implementation was 20% slower and so everything costs 20% more to use and no one uses it.
This is why you see things like TPUs never really damage NVIDIA, but why basically everyone is focused on open standards and open software. Basically the entire tech industry is using this approach as a way to slowly peel away the layers of this software until enough has been removed that NVIDIA can no longer use it as a moat.
PyTorch, TF, and JAX work great on TPUs. Adoption is low bc they are not really available outside the Google cloud.
That's gotta be way easier, no?
OpenAI could just build their own framework for internal use that works well on their silicon (see Jax+tpu)
Their starting point? Triton plus some triton libs. Jax chipped away at TF like this, and no reason why Triton can’t do the same to PyTorch.
Answer is - it is really works, but slow (about 0.5..1 tokens per second, with near 100% CPU usage).
i7-7700 is good weighted machine, but before I few times achieved memory speed bounds with highly optimized software. And it looks very different. When use all cores, I got somewhere about 50% of CPU usage.
BTW Llama.CPU is very good software.
Some of the things being done to improve quality of 6-8 bit inference use extra compute and push it a little in the other direction but it’s still pretty memory intense until the batch size gets quite large
Also, if you have just a single model you want to optimise (and not the training), you could build an array of asics that do specific matrix computations - then you don’t need to read weights from memory at all.
/s
For 2022 and 2023, Microsoft bought a significant portion of NVIDIA's available hardware. They spent quite a bit of 2023 trying to figure out how to even power the multiple fleets of GPUs. Just now with the mild to expected wild adoption of Azure OpenAI are they getting around to servicing all their (potential) customers.
Seriously, this is am outlandish claim just from looking at Microsoft and Nvidias market cap.
I am sure that Microsoft is gonna be one of Nvidias largest customers, but I sincerely doubt it's even a double digit percentage of their revenue.
To reach double digit revenue of NVIDIA's 2023 at $26.97 billion, you'd only need to hit ~$2.7B in sales.
H100's are priced anywhere between $20k - $35k, so required to purchase ~77k - ~135k units.
That is singularly H100s, Microsoft also offers lower compute, and they have the rest of Azure to service with a variety of solutions.
Being at #1 or #2 market cap worldwide is not a farfetched position to be a significant controller of chips, especially since they directly work in the space as a platform.
.. but is that true?
MSR has been putting out research in all derivatives of modern large neural network architectures (NLP, CV, etc.) for the same amount of time that Google has. If there was a drift between timelines, its not large IMO.
What you could argue is that Google historically was more successful in their research outputs.
However, historical consumption of resources may not compare to current resources consumption.
> I doubt we have the visibility to know how they compare in terms of available flops and the unit costs
Completely agreed, unfortunately, this is all guesswork at best
> Using Pierce-Arrow Motor Car Company as an example of such success is historically inaccurate. Pierce-Arrow was an American automobile manufacturer based in Buffalo, New York, which was known for producing luxury cars. It was indeed a dominant and prestigious brand in the early 20th century. However, the company did not manage to maintain its success and ultimately failed to adapt to changing market conditions. It faced financial difficulties during the Great Depression and eventually went bankrupt in 1938. Pierce-Arrow's inability to forecast and adapt to the economic changes and shifts in consumer preferences of the time led to its decline.
Even granted that OpenAI are not able to build a chip that is competitive with NVidia's latest GPUs for running LLMs right away (which is an opinion - not backed by any direct evidence, but I agree that it is plausible as they are going up against a lot of prior R&D) is it not possible that:
a) The unit economics could be so much better that the result is still a major win, e.g. 50% of the performance at 20% of the price.
b) OpenAI is decoupled from existing supply constraints and is able to grow faster and deliver more value. A "worse" chip that you can actually get (in insane volume) may be strategically better than a "superior" chip that is limiting your growth.
c) That the plan might include some elements you are not expecting - at the $trillions investment level they might be looking at doing some surprising things e.g. (I am just making this up but there are a lot of possibilities) buy a memory manufacturer and work directly on increasing memory bandwidth.
The idea that even with expertise, the wins would be so much over what other companies that have hired/bought these companies have been designing for the last 10 years based on very similar requirements (the ones that wrote so much of the foundational research) also seems implausible.
c) It's not actually possible to plan investments at that level with anything more than a very vague direction you're aiming. If it is long term, then everything is changing in unpredictable ways before you get even 25% there, but if you throw so much money at the problem in order to try to solve it much more quickly you are disrupting global economic and geopolitical forces in ways that also can't be planned for.
It seems more likely to me they'd get 20% of the performance at 50% of the price, and that might still work out for them if it allows them to scale faster without being bottlenecked on supply of existing GPUs. But there's no magic bullet here.
They also still need to source a bunch of other stuff, like RAM, even if they can source their own processors.
What it tells me is that Altman seems to believe that OpenAI can only make the next step if they can throw even more compute at the problem but that that isn't feasible at today's prices.
Perhaps they could design a core and license it out? I'm trying to come up with a way they can do something significant without 100 people. Just the memory and serial connections are complex enough ignoring the GPU or heat/power issues.
What is essentially happening in my opinion is technical innovation has slowed so silicon valley is seeking money to prop up a house of cards that doesn't make much new that is useful or needed.
Can anyone specifically say what trillions of dollars invested in "AI" would buy for society?
It seems to me there are so many higher priorities.
A typical case of engineer's disease.
Perhaps they'll pull off an Apple (for ARM) and do their own architecture (either for training/tuning or inference) that will have a significant effect on the industry, but it seems unlikely. They haven't hired the right people.
The real advantage they might have is insight into how the algorithms can be adapted to reduce power consumption/latency while improving performance. It would seem odd to me, if there weren't more than an order of magnitude in new algorithms for LLMs. You're not going to get 10x the transistors or speed from silicon, but you might get an efficient architecture for a significant algorithmic improvement (that might not just be CUDA).
Is that true? I can't find anything suggesting it is. In fact, the little I can find suggests you are incorrect. I'll link them for the sake of referencing sources but they're both pretty awful ad-ridden sites...
A 2016 Tech Radar interview [0] with Norm Jouppi has him quoted as saying:
> [The] Tensor Processing Unit (TPU) is our first custom accelerator ASIC [application-specific integrated circuit] for machine learning [ML], and it fits in the same footprint as a hard drive.
And a 2023 Tom's hardware post [1] begins:
> Google has made significant progress in its endeavor to develop its own data center chips, according to a new report. The Information says that a key milestone has just been reached, which means that Google can plan to roll out server systems powered by the new chips starting from 2025.This is not the first processor that Google has successfully put through R&D - the company has previously made an ASIC for servers and an SoC for mobile devices. The search giant started using its internally developed Tensor Processing Unit (TPU) as far back as 2015.
[0]: https://www.techradar.com/news/computing-components/processo...
[1]: https://www.tomshardware.com/news/google-reaches-self-develo...
1/ https://www.wired.com/2012/03/google-microsoft-network-gear/
2/ I believe they had a few custom chips designed for the youtube workloads that predate the TPU.
I remember in 2010 there was a building in MV that focused on custom chips.
Training and designing LLMs doesn't mean you understand the semiconductors business.
But perhaps it turns out that subnets can be trained independently or swapped with semantically equivalent but qualitatively different ones. The routing network would effectively "Standardize" and could in principle be well enough understood to "hand optimize" the routing network into hardware. Or maybe back propagation has some novel physical analogue that can be exploited in scales we can access. The real question is if Altman is capable of finding the right path in the notoriously dead end filled field of chip design. His backing of helion [1] didn't bode well in my view. But with enough R&D maybe he will flail into something useful trillions is enough for a lot of flailing.
[1] https://youtu.be/3vUPhsFoniw Edit: more derisive link
Must be really hard, being only a half-billionaire and trying to keep up with Elon's "success"...
I find it hard to argue that this mode supports a 1.7T valuation. I find it hard to believe that for a couple of billions + TSMC credits no one would be able to recreate the CUDA ecosystem + hardware in the medium term.
$7 trillion is like adding TSMC, Intel and AMD together, and multiplying that combination by seven.
This is about sheer capacity, not just circumventing CUDA.
Google came up with the TPU (2015) for GEMM. Nvidia just took the idea and ran with it (Turing 2018). So it wasn't that Nvidia had a head start on this.
Now Nvidia Hopper is ahead of everybody else by far. They have things like async memory management for the tensor cores (Tensor Memory Accelerator), mixed precission, and even FP8 support.
Most of the software out there has not yet caught up with that. And even Nvidia's own Tensor Engine software is not making the best use of it (Microsoft Research October 2023, backward pass and cross-device communication).
Last year FlashAttention was a game changer for performance by doing memory load optimizations. Nobody was optimizing properly for Nvidia in Transformer models.
Obviously I think my company is doing this in an unique and "correct" way, but I know of half a dozen other companies founded in the past ~18 months that are focused on the memory capacity and bandwidth bottlenecks that exist... the massive failures of the previous decade do not mean that they are going to be repeated.
Before appear Tensor cores, GPUs was about 4 times worse (speed, power consumption).
With Tensor cores, GPUs become better, but they still need to carry video hardware (ramdac, video connectors, 3D processing units, network to connect all this stuff), so they still late.
Really GPUs are interest just because current AI applications are not achieve enough revenue to pay for large scale production of special chips.
I don't know, if Altman have something Big to get revenue to pay for special chips.
Exists speculations that GPT-5 will be enough to replace human at work. If this is real, AI chips will be worth it.
Classics of management, to ask people more then they could, and they will do most possible, so I don't bother much on such claims.
And also this is teambuilding bs, to motivate people claiming impossible targets.
Will see, how Jensen Huang will use all his diplomatic skills and rhetoric art, to round corners, when become clear, that claimed things impossible.
And this is not first time, such things happen, there are near infinite number of examples. I just few days ago read about IBM 7030 fail, which delivered ~1/10 of claimed, and yesterday people remembered me about Itanium and i960.
What I see, NVIDIA is good, strong team, they bet very high stakes, when made great acquisitions in 2000s and they won. But NVIDIA made wide targeted product, they cannot made very narrow focus on just neural net. So it is possible to make NN product better then NVIDIA.
Real question is to predict, if Altman team could achieve so good economy, to pay expenses for hardware development.
What really main bottlenecks of NN hardware are neither number crunching, nor memory.
Real bottleneck is that GPT-2 is may be last LLM for which was possible train on one machine (even on one card).
About GPT-3 usually people said about 32-GPUs installations (possible to install into one machine), for GPT-4 scale said about clouds.
And modern clouds are NUMA beasts. I could say, modern clouds networking is slow, but it is not right words, as they are slow as hell.
What all these mean, NN are good target for parallel processing in clouds, but not good enough. Real benchmarks said, mentioned 32-cards machine is about 10 times faster than 1 card with such amount of memory, and when on GPT-4 things scaled, benchmarks become much worse. So, just improve network to move bottleneck to something else and will got additional 50-100x improve.
And with good team of AI scientists, it is more real to make special hardware network for NN processing, or to tune algorithms, than with team of GPU video processing specialized team.
This is not true. You have tones of models those are even better than GPT-3.5 and really close in performance to GPT-4 and you still can train them on a single GPU with 24GB video memory. There is a hint at yet better models published last year which you can train on a single GPU and have a model comparable in performance to LLaMA2 34B. The horizontal scaling which you appeal here, may fit into 10^6 performance increase, but in general I expect single node to be at least 1000 times faster than now. And it is totally feasible that you can't scale with 0.99 vertically and of course not horizontally, but I honestly expect the scaling per GPU get better than 0.75 in next 5 years.
It depends, on what target. For pure science (or for enjoy), I could train GPT-4 class model on C64, but this method will not fit on concurrent market, where need fast check hypotheses and fast deliver tuned models.
- Concurrent market is very sensitive for speed - for example, if MS present something on December 10, Google after New Year should present not equal, but significantly better, to just appear equal for customers.
So, horizontal scale is a must, not just my wish, even when speed increase is far from linear.
> I honestly expect the scaling per GPU get better than 0.75 in next 5 years
Could you give explanation, or even speculations, how this is possible, when we already hit Silicone limits (about 5GHz core, 1nm, etc)?
Nope. But i'm so desperate to give you a hint right now, it is almost impossible to hold myself... Stop looking into horizontal scalability. The vertical one is not exhausted yet. Btw that was not the hint.
Sure. B-747 officially need about 700 man-years so assemble, lets make them with small but highly motivated teams, with classics 3 pizza rule, world will wait :)
E.g. FB saying they want to buy 350k H100. That's just a whopping $14B price tag. With a >85% profit margin. While a fab is $20B.
Trillion? Sounds like anchoring to me. Nvidia has a market cap of $1.7T. You could literally buy NVIDIA for that. I read that as "a billion won't cut it, we need quite a few billions".
But it's not unreasonable that those hyperscalers throw in a few billion each.
Usually it's horrible business not to be best (see Intel/AMD). Because the margins are at the top. In this case though they want a whole range of products to go down in margin. Even a slightly worse chip might be worth it if it comes at a significant cost reduction. Especially if the optimal design is known!
In a sense the whole thing can fail at reaching the top or making lots of money and still succeed in bringing total cost down, potentially by 50% or more.
Further, on the hardware side, you have Nvidia, AMD, Intel that are competing fiercely.
OpenAI is already under attack from Meta and Google and who know how many LLM companies, while it's the first runner for now, that might change fast.
In the worst scenario, its LLM will become just one of a few in 2024, and its chip design will be a few years away if anything and might just turns out to be nothing, just look at Musk's Dojo AI chip, which started years ago and still he is buying Nvidia.
Keep acing at your LLM, and leverage Azure's platform for now, then also use some AWS to be safer. Forget about making your chip, it's too far away, and it's too crowded there and too late, you don't need to do everything to be successful, at least, no rush for that now.
The current AI wave is creating a lot of financial incentive to find ways to speed this up though. Interested to see if there ends up being a way around the TSMC / ASML bottleneck.
Beyond all the existing scaling up started based on pandemic shortages, that's the only thing I've heard of that has a chance of making a dent within 5-10 years. Of course, all that already started scaling up might already be more than enough if you're not a true believer.
This is a wild claim, can you explain? You are suggesting that a chip who's design is manufacturing ready today cant be built until 2039?
Is this true for FB? They just spent billions on 500,000 H100s (disclosed in their earnings report a week or two ago).
Joking aside, I do think the US needs massive semiconductor investments in places that arent Taiwan. Ideally in the US itself, but anywhere further away from that geopolitical time bomb would be great.
World GDP is now $100T.
If the $5T is spent over 5 years, that's 1% of GDP — incredibly high, but within the realm of possibility if we become singularly focused on building out AI.
Given gpt-4 is already ridiculously useful and we’ve barely scratched the surface, it makes complete sense to me. More capacity + faster gpt responses unlocks massive amounts of more potential/use cases.
No one is even talking about these kind of figures being used for climate change, which is a far more pressing problem.
Geometric growth doesn't feel like much until you slam into the wall.
Sure ... and we've already let Koch and Co. piss 50 years of lead time up against the wall since the first global recognition of the problem in the 1970s.
Now that it's starting to bite and properly ramp up there's far too many that are stretched out lizard like on Titanic deckchairs asking AI Jeeves for another drink.
The Green parties in Europe successfully stopped the expansion of every nuclear energy program in the EU. Greenpeace engaged in a number of terrorist attacks to sabotage nuclear energy.
It's a factor, sure, but "biggest" .. not so much.
The transition from ICEs to electric battery cars is largely orthogonal to base load power, but even electric battery cars depend on a base load source, so the extent that CO2 emitting energy sources have been replaced by non-emitting ones is highly dependent on the base load sources.
At the rate AI is accelerating we may not get even 3. It's the final force multiplier. If we end create runaway automation feedback loops, there may not be much of a recognizable planet left to have a climate in ten years. A spot of hull rust quickly consumes the entire ship once it takes root.
I grasp the looming disaster of climate-driven global collapse. We simply found a way to speedrun disaster even more efficiently.
Where's the killer app? The only one I can think of off hand is co-pilot and the reception I've seen is that it's pretty mid. Most of the proposed applications require human checking to get right which is a huge limitation to the adoption of these systems unless you accept a 3-5% error rate which is terrible. I've not met anyone who is interested in something like a book written using this thing and the main use case I've seen basically amounts to denial of service attacks with believable bullshit.
Frankly the only people I've seen who are super excited about this stuff are people in the field or the uninformed.
In terms of error rate: gpt 3.5 had a high hallucination rate that made use cases fairly narrow. It then got faster which opened up some more use cases. Then gpt 4 came out that had a significantly smaller hallucination rate which opened up a gigantic number of additional possibilities. And had a larger context window and output size that made it significantly more useful. Then it got faster with an even larger context size… each of these iterative improvements just continue to add more and more possibility in a gigantic range of cases that have literally never existed before.
But even more importantly, it seems like the ultimate goal could be reasonably expected to drive large scale unemployment, with no clear replacement for such. Necessity is the mother of invention, but in this sort of scenario you're expecting kings of industry and politicians to put aside their greed and self interest for the sake of social good. The odds there are going to be overwhelmingly in favor of dystopia, which we substantially collectively sacrificed to achieve?
If any discussion of AI build-out collapses to "Will AI lead to dystopia???", then we're not going to make much progress in discussing AI build-out.
I somehow doubt you can convince even half the liberal voterbase that that would be a good idea. let alone any of the conservative base. But I guess the US did manage to do that with the space race, so maybe the key is another cold war.
Don't conservatives like capitalism and religion? Just get some polyamorous nerd to give a TED talk about how they're building God for real to serve the Titans of Industry. That should get the conservative voter base in line, right?
But 7 T usd is about 1/3 of the US GDP. That is another way of saying the value of 100 million Americans working for a full year.
It is really ridicolous.
https://techcrunch.com/2019/05/18/sam-altmans-leap-of-faith/
What I find amazing is that this stuff sells even though it is way beyond common sense and I am sure in couple of years will be subject to ridicule.
7 Trillion? The U.A.E's GDP is only 509 Billion [1] ... Something seems off with those fundraising numbers
[1] https://en.wikipedia.org/wiki/Economy_of_the_United_Arab_Emi...
Moonshot R&D has a long history of leaving unlucky investors bankrupt.
This is just to show that GDP is not necessarily reflective of investment power.
Wrong currency. Saudi Aramco's market cap is around 2T USD.
But in terms of investment power Saudi Arabia's Sovereign Wealth Fund seems more relevant & that is only 776 Billion
EDIT: Typo
On second thought... go for it, Sam! I'm looking for some powerful AI hardware at great prices...
the thing is you can't know if it's the right time, it's indeterminate. all you can do is go for it. usually the world will benefit, whether you win or lose
It's clearly a plot to put more control of world finances into the hands of the US. Nations should not expose themselves to such risks.
though even trillions is an exaggeration here. If they all collectively had 10m employees making 100k on average for 10 years, we're an order of magnitude off.
Of course someone will make breakthroughs. Everyone stands on the shoulders of giants.
As evidence - if I learn a new game, like chess, my rate of mastery per game is going to be orders of magnitude higher than an AI (which will take millions or billions of training sessions). Granted my brain has more synapses than GPT4, but it's unclear how many of those I use playing chess (I wouldn't be surprised if it's less). And regardless I don't think that anybody's arguing that adding more layers accelerates training but rather increases the peak.
So perhaps a place to start is studying cases where AIs learn much slower than humans and trying to understand better algorithms of training/reinforcement that may come closer to what the human brain does.
We may eventually build machines that are able to do that, but it may well take us enormous amounts of "brute force" training in order to produce something that then no longer uses brute force to learn further.
As an analogy, look at the history of CPU design. The first CPUs have been designed manually, even "drawn" manually. Each generation of CPUs empowered engineers of the future generation to build more complex tasks because they could tap the computation power to assist them in the design and realization of the circuits.
Development is incremental.
Yes, sometimes you can come up with some new groundbreaking idea that will invalidate what's currently being worked on, but usually these ground breaking ideas come up later when the ground is fertile because you already live in the next generation.
I mean this should be obvious but the "ability to raise trillions" is because those people think they will get X returns on the money. Your "ability to raise trillions" will quickly evaporate if you offer to throw the money into a pit and set it on fire.
If national governments offered a $10T bounty for solving global warming, then sure, you could raise $1T to moonshot it... but there needs to be an incentive.
I suppose a big part of this is that the problem doesn't seem urgent enough that it would inspire this sort of ability to raise money.
But it's a collective action problem and you're not going to accomplish anything by guilting investors into burning cash singlehandedly solving a tragedy of the commons.
While I'm all for trying to lower your individual consumption, asking everyone nicely to use less WILL NOT solve the problem, the same way asking companies to use less won't either.
This is a coordination problem and can only be solved by governmental action.
Of course it doesn't help that governments tend to also be annoyingly inept at this, but at least they have more of a chance to get something done for this.
He’s got guts, but does he really have enough relevant experience in hardware (the hardest parts of it too) to think this will be successful? Where is all the confidence coming from?
We already know it is possible. You just need to hire the right people.
Elon launched reusable rockets into space using the same philosophy. Jobs built a personal computer likewise.
Before appear Tensor cores, GPUs was about 4 times worse (speed, power consumption).
With Tensor cores, GPUs become better, but they still need to carry video hardware (ramdac, video connectors, 3D processing units, network to connect all this stuff), so they still late.
Really GPUs are interest just because current AI applications are not achieve enough revenue to pay for large scale production of special chips.
I don't know, if Altman have something Big to get revenue to pay for special chips.
Exists speculations that GPT-5 will be enough to replace human at work. If this is real, AI chips will be worth it.
Can he compete against Nvidia?
The scale and scope of what's being described seems outlandish, I think $150B would probably be enough to at least get started. that should get you two plants in a country more amenable to building and running a chip facility than the US, plus a hardware design team (to make AltmanTPU v1), a software design team (to make ATPU work with pytorch), and the business side of things.
I don't see these things being sold commercial, it's more like airplanes- a few big players put up the capital to fund the construction of large batches and then get dibs on the first good volume runs.
I think the real question here is what will ASML's involvement be, what process and design will the TPU have, what will the computational infrastructure surrounding the TPU look like, and how we are going to plumb all the necessary power to the TPU farms.
The scale and scope of what's being described seems outlandish, I think $150B would probably be enough to at least get started. that should get you two plants in a country more amenable to building and running a chip facility than the US, plus a hardware design team (to make AltmanTPU v1), a software design team (to make ATPU work with pytorch), and the business side of things.
Ask Intel if they can just throw money at advanced chip node manufacturing. The only company that is able to do 3nm at scale with good yield is TSMC. Even at 5nm, Intel and Samsung aren't competitive.It is known very long time, that specialized AI chips are more efficient than GPUs, even when we talk about Tensor cores, because GPU is not only Tensor cores, they MUST include ramdac, all other video stuff, and network for fast communications of all this stuff, which is not of need for AI chip.
Before Tensor cores, benchmarks was about 4 times against GPUs, now better but still significant lag.
A billion isn't cool. You know what's cool? A trillion.
They should focus on fixing that problem instead of creating more expensive chips. There should be at least a dozen start-ups working on just that
This is journalistic malpractice.
They had an edge, they lost it, this is now desperate scrambling for cash
If the promise of AGI is true, then wouldn't we be able to ask it how to become more efficient? While it has taken millennia, mother nature has created an incredibly efficient mobile intelligence system. I struggle to understand how one could stake so much on the chance to replicate it.
Humanity can do better. Maybe not forever, but certainly for the moment.
Quit it with the speciesism. How could you question the wisdom of aloof techie galaxy brains?