Nearly half of Nvidia's revenue comes from four mystery whales each buying $3B+
fortune.com
fortune.com
Other big buyers area: Oracle, CoreWeave, Lambda, Tencent, Baidu, Alibaba, ByteDance, Tesla, xAI.
https://observer.com/2024/06/nvidia-largest-ai-chip-customer...
(Even without a report on this it would be obvious)
That was a major payoff for Apple - I wonder if any of the other fangs will actually be able to follow suit.
AFM-server was trained on 8192 TPUv4 chips
Someone more versed can say if that is huge or not.
Oh the irony for Apple to dislike others being secretive...
macOS is 1000 times better for talking UNIX systems than Windows and is POSIX compliant.
Lastly, they are not hindering the development of Asahi Linux, and did nothing when their devices were reverse engineered. On the contrary, they left a couple of ways open for Asahi guys to boot their distribution directly.
They are not the band of saints, but they are not the underhanded evils like a couple of others.
It drove apple crazy both with high failure rate of MacBooks where the GPU was desoldering itself and general problem of a hot as fuck bottom. Nvidia refused to pay out for damages to Apple as well from what I recall.
Then a few years later they made up but NVIDIA didn’t want to partner on drivers, so they had another rift
Some of us use them at double precision mode.
There’s a whole world using GPUs to accelerate things.
I'm saying that Meta and Amazon and Microsoft are all buying these chips in insane numbers for AI—their usage for all other types of GPU activity is at least an order of magnitude less. That's why Nvidia skyrocketed to become the most valuable company over just a few years.
For Google to be on that short list of whales would either mean that they for some reason have a much larger demand for GPUs for non-AI purposes than any of the others have for AI purposes (doubtful) or that they're using GPUs for AI.
They’d rather offer their clients what they need than push them on to their own products.
[0]and supposed software powerhouse.
Amazon doesn’t force everyone onto graviton cores, and offers competing cores.
It’s just the nature of business.
The public in-house projects that I'm aware of (but as far as I know haven't fully replaced demand for Nvidia GPUs) include:
- Google's "TPU" (in production, publicly rentable)
- Amazon/AWS's "Trainium" (in production, publicly rentable)
- Meta's "MTIA" (in production)
- Microsoft's "Maia 100" (I'm unclear on their status)
- Tesla's "D1" (I'm unclear on their status)
Is this trailing year NVIDIA sales, or order book ?
https://www.forbes.com/sites/craigsmith/2024/08/27/cerebras-...
Maybe it is working for Meta or Tesla where things can be vertically integrated, but for the public clouds, they have to buy NVIDIA for their customers.
According to Wikipedia:
In September 2023, Amazon announced an investment of up to $4 billion, followed by a $2 billion commitment from Google in the following month.
They probably use both. Almost certainly most of this is credits in their cloud platforms.
They might buy some, but I think that Google, Meta, Microsoft and Amazon and will be the ones buying in large batches to enable companies like Anthropic (and themselves) to scale up to world wide inferencing demands, as well as generally offering the most efficient GPUs to their customers.
Very plausible. I'm not sure at which point it makes economic sense to buy the GPUs and build out the infrastructure to continually be training something like Claude.
It's a major reason why they raised with Amazon [0]
There are actually a LOT of other large companies that participated in Anthropic's round but haven't announced it publicly.
[0] - https://www.aboutamazon.com/news/company-news/amazon-anthrop...
https://www.wsj.com/tech/ai/google-commits-2-billion-in-fund...
And while yeah, that's probably them, I'd put a non-trivial chance of some government intelligence organization to make it into the top 3.
(They absolutely use it as a holder for Meta)
Not at that size. That is VERY on the nose sanctions evasion.
Such sanctions evasions tend to use multiple smaller parties doing purchases and then reselling.
Most likely DoE. TLAs purchase indirectly (or use other federal agencies in the DoD as a front)
If MS drops OpenAI, it's not like they can just seamlessly pivot to running their own data centers with no downtime, even with pretty high investment.
I’d note that the supplier of GPUs is Nvidia, who also offers cloud GPU services and doesn’t have a stake in the GCP, Azure, AWS behemoth battle. I’d actually see that as a more natural less middle man relationship.
The real value azure brings is enterprise compliance chops. However IMO aws bedrock seems to be a more successful enterprise integration point. But they’re all commodity products and don’t provide the value OpenAI provides to the relationships.
The arrangement is mutually beneficial, but the owner of the IP holds the cards.
https://blogs.nvidia.com/blog/meta-llama3-inference-accelera...
1. Are chatbots going to get much more effective than they already are? It seems like all the major players are plateauing and the different models are becoming commoditized. That doesn't bode well for sustainable GPU sales. Also if the hallucination problem can't be solved, it's not clear that this generation of AI will ever be deployable at scale.
2. Are there genuine at scale use cases for AI outside of LLM's? Autonomous navigation seems like a major one, but I'm not sure how close that is to production ready. I know drug discovery and other applications are talked about, but not sure how much GPU consumption they can realistically generate. As we leave the novelty phase of the adoption curve, it's clear that a lot of the use of the image generators was unsustainable experimentation. My personal experience has been, a year ago my friends were creating tons of images but now we hardly do at all.
Assuming that "outside of LLMs" means "outside of text processing".
Yes. Robotics. Imitation learning with LLMs is working surprisingly well. It will require a lot of investments in data and training to get to a practical state, but all the early signs point that new revenue streams will be unlocked in Robotics.
One limitation that still stands on the way is the inference speed. My estimate is that we need ~10k tokens/sec prompt processing speed to get these smart robots working reasonably fast. We're getting there for 8B models (Groq & Cerebras silicon), but these 8B models are too dumb (especially, after being finetuned on robotics data), and 70B models are still 20x slower than practical.
The dot com bubble popped, but it's not like the Internet technologies that were launched then (and companies like Amazon and Google) weren't hugely impactful on all of society since then.
I think the AI bubble will pop, and while I think there is a lot of nonsense hype about AI I still think AI's societal impact will only grow.
Nvidia made 18 billion in profit last quarter, and expects to make 20 next quarter. That isn't speculation.
How much money is OpenAI or Anthropic making? Because that's what people are thinking is speculative value.
My position has always been that Gemini/ChatGPT/Claude are all pretty great at a cost of Free, and grow increasingly questionable past that. My work is already limiting how many ChatGPT users we can afford with their price increases, and I'm pretty sure OpenAI is still not profitable. If ChatGPT is $50/month as a breakeven cost for them, how many people/companies will buy it then? Most jobs I've been at won't pay for JetBrains licenses that cost way less per head.
I feel like the best comparison is something like Uber or AirBnB where it's easy to be excited about it when all the services are crazy discounted by free VC money, but when they have to start turning a profit, they're back to actually competing with other tools.
But the big deal isn't OAI being a profitable company. The big deal is that GPT6 will be 100x more useful in doing productive work.
Tulips did not have cash flow like this. It was only people selling to speculators who hoped to sell again to another speculator.
Citation extremely needed. There's a lot of people and companies downstream of OpenAI speculating on that 100x that are gonna be in a lot of trouble if it's even just 10x, let along 5x.
Again, not saying that none of this has any value, just that the value may well never live up to the cost. Uber's not a worthless company or service, but they're far from the values or profits they were pitching 10 years ago.
1. Drive a new Tesla with the latest Supervised FSD and measure how often you have to intervene to stop a crash.
2. Go back and look at your own expectations around AI two years ago. Did thing progress the way you expected or did they progress further?
2. If anything, GPT4 has turned out to be less of an advancement over 3.5 than either OpenAI was claiming and what I'd expected. 2 years ago, people were all but promising AGI by now. Even the folks I know working in the GenAI space are telling me they're using Copilot/ChatGPT less now than a year or so ago. My work has actively cut back on spending in the area and investors have been asking our board questions to make sure we're not overinvesting in it.
I want to be clear, I'm not a doomer at all about this. I use these tools a fair bit and find value in them. But the value that GPT3 and 3.5 brought to me versus what GPT 4 has brought certainly isn't 100x. GPT4 isn't even 100x better than me using Google Search most of the time.
2. I have no expectations for 'AI' because the term is a nonsense label. I have followed and been excited by machine learning for a good number of years, and my expectations of progress were pretty much on par. The progress with LLMs has taken me a little by surprise, but I am also cognisant that their progress is being massively over-hyped presently, not least by ppl who call them 'AI' and then, even more foolishly, go on to talk about 'AGI' (a nonsense upon a nonsense).
I'm intentionally including FSD and LLMs under the same category of technologies that will have a huge impact. The point of this thread is that the demand for inference is going to skyrocket because AI is going to get a lot more useful.
We both appear to agree that Machine Learning is a very powerful technology that will have huge impacts. Machine Learning requires (and will continue to require) a lot of compute and thus large costs but will also, almost certainly, produce great profits in some domains (FSD being one).
It's a lot less clear to me that LLMs will 1) continue to require lots of compute beyond the short term (languages can get close to being 'solved') or 2) that LLMs will generate substantial profits because a) the model can escape capture from a monopoly player far more easily and b) while useful for translation, pulling summarised data from a corpus, recognition of voice commands, etc, none of these applications actually make for the kind of profound impacts that ML is capable of, because none of them transcend human ability like ML has the power to do.
Same'll happen here.
Oh. It's been tried: https://www.pcgamer.com/nvidias-ultra-expensive-h100-hopper-...
Seems the pros of an h100 over a 4090 are: much higher vram, much faster vram, technologies like nvlink available, and a focus on lower precision performance more useful for ML (as opposed to 4090s focus on fp32).
If the AI bubble bursts, people will use the available GPUs for something else.
Yes, of course, but that just means that this bubble would be basically identical to previous capital intensive bubbles. For example, there was a railroad bubble in the 1800s, and a massive telecom bubble in the late 90s. These bubbles popped, resulting in massive corporate bankruptcies and failed companies. But the infrastructure they built (miles and miles of railroad and dark fiber, which has since been lit up) laid the foundation for huge economic development shortly thereafter.
If the US had maintained and kept the rail it built, it wouldn’t have the poor infrastructure it has right now.
Nvidia is not the train company on that scenario.
AI/LLMs are radically expanding my abilities, and as I adapt to this new power, I'm using it more frequently in everyday life.
Sure, Nvidia stock may be overpriced, but AI is empowering. I can't imagine not continuing to expand its use. As its abilities expand, I'll use it even more. I will have much further use even as a few bugs are fixed and integrations become more frictionless.
Maybe it is not for you. Maybe it is for people asking AI questions about you. (or chemistry or gold prospecting or legal documents or ...)
the eric schmidt talk made it seem like better hardware led to better results and there was a race.
Whereas what it could be reinforcing instead is that some people are better at "using AI" than others.
When I was young, I always saw how my parents never really "got" new technology that I was using all the time, like the internet. Many young people think about it and are sure it won't happen to them. I'm sure many on this technophile site think so.
And then a new technology like AI comes along, some people find ways to be incredibly productive with it, but a very widespread sentiment is that they're... lying? Mistaken? Not very good at their job so it helps them more? The number of excuses people have for "keep this new technology that I don't know how to use away for me" is pretty crazy.
(And I say this as someone who is probably not on the "cutting edge" of AI usage, compared to others I see.)
When businesses stop accepting dollars and your employer starts compensating you in crypto will it stop being a “scam” or will the goalposts move again?
(I don’t expect it to see it in my lifetime.)
The scam is comparing some ATMs to what is happening in AI. Trillions of dollars are going into AI and actually useful things like self driving cars are coming out.
Most cryptocurrencies are just straight up scams. Some people are getting some usage of Bitcoin as a store of value and for cross border transactions. This is increasing slowly. Stablecoins also have some usage for store of value and cross border transaction. They are also used for trading and arbitrage, which you can argue about whether that brings value to the world. The rest of the crypto market is struggling to find an enduring use cage. I'm saying this as someone who is marveling about Ethereum and Solana, but I don't value trading immaterial NFTs. Ethereum and the like are struggling to do anything that reaches into the real world. None of the cryptocurrencies outside the top 10 have found any real world use case that people care about.
So AI != crypto
https://youtu.be/NC5NZPrxbHk?si=8uQ4zdMU02f4X1Hc (at 1:41)
Hyperscalers & Meta.
(Corp speak 101: Hyperscalers = AWS, GCP, Azure)
- Nassim Nicholas Taleb, Twitter, 2021-09-11
There's no one stealing market share from Nvidia at the moment. Groq and Tenstorrent are extremely promising, but both are still private companies. Once Groq goes public, Nvidia will tank a bit for a while while all the "experts" announce the end of Nvidia. I wouldn't be surprised if then Nvidia would then also sell specialized AI accelerators, if they find that segment attractive enough due to losses in general GPU demand created by those companies.
To quote Steve Jobs when talking to Dropbox: you guys don’t have a product, you have a feature.
Or if closed models will dominate. For example, by the largest companies leveraging their existing distribution channels and/or acquiring promising startups.
They do this because proprietary AI models are not flexible enough and are lacking a lot of API.
For example, one app I wrote was to analyze scans of old maps and use generative AI to extrapolate and create animations.
I don't know where the market will go. But my feeling is that large proprietary models are very good at a very limited type of work and that open source will provide diversity.
Competition when ? Are amd, Intel or other companies in the situation to be able to eat some of nvidia insane margin ?
AMD and Intel would probably be better off researching entirely different approaches that they can leverage their existing expertise for - i.e., some architecture that relies heavily on efficient OoO processing pipelines and free (if predicted correctly) control flow changes. Techniques that are antagonistic to GPU processing could represent a competitive moat.
Joining an existing rabbit chase right in the middle can quickly evolve into a catastrophic strategic choice when the cost of entry is billions of dollars.
What if companies still keep buying insane amount of graphic card for the next 20 years ? At some point, other companies will want to eat some of the cake too.
What about Chinese? They always want to have Chinese made hardware because of the fear of spying and trade war, they too have extreme need for graphical power and nvidia is a Delaware company.
It seems there are multiple reason for competitors to step up.
(I’m only partially kidding, sigh…)
Office 365: several options
OneDrive: several options
Check out https://github.com/awesome-selfhosted/awesome-selfhosted
At least the overloads I have worked for don’t allow us access unless the machine is locked down so hard that it’s borderline unusable.
One benefit of non-securities underlying assets is that you can play with their pricing a lot more. like, you can have your friends vesting on some shoes you control the issuance of - or GPUs in this case - at a 99% discount and there’s no reporting or regulation to a government over this. Big problem to do that with shares.
There is also an alarming (as shareholder) rise in custom silicon. Groq sambanova cerebus etc.
N _ _
_ B _
_ _ A
_ H _
Right now it’s “buy GPUs at any cost”. If things slow, there will be a chance for these customers to consider how to optimize this cost. NVIDIA can’t sit on its laurels like Intel did with x86.
Compared to the problem of developing a new cutting-edge GPU, building a CUDA compatibility layer is a much smaller problem. Hire the author of ZLUDA, throw a small team at it, and have a legal department on standby. And separately, there'd also be value in some source-translation projects to help people migrate to some better native framework.
Also as for AMD, as I understand, they are unwilling to make ML libraries for consumer-grade GPUs and GPUs built into CPUs.
The reverse approach, of trying to entice people to move from CUDA to a different library and switch GPUs at the same time, has been tried repeatedly and has not yet succeeded. Trying something different seems warranted.
So most probable end result is that we end up with multiple competing alternatives all with their own vendor lock ins. And general public might be lucky to get one or two options.
First, it's the classic chicken and egg problem. Why would you invest in a CUDA alternative when you're going to be using nvidia hardware anyway?
Second, something can be not impossible but still quite difficult. As AMD and Intel have shown, creating a GPGPU API for your hardware that people want to use is not a trivial task and to date have not managed to do it.
Lastly this must just be differences in our experiences with cooperate management, because mine has been that in general they would always prefer to spend on stuff over headcount if said stuff reduces the headcount required.
People do not scale.
I hope they're not saving their best chips for the likes of Tesla/Grok. That'd be a PR nightmare if and when it leaks.
https://arstechnica.com/tech-policy/2023/09/musks-unpaid-bil...
If you asked 1000 random people to say what they know about Elon musk, what percent do you think will say "oh. You mean the guy that doesn't pay vendors!"
Also. What is the base rate for contract disputes with vendors among large companies? He runs 3 large companies, surely with tens of thousands of contracts for services. There will always be disputes there - is his rate higher than average? Does he lose very dispute in court?