AMD's MI300X Outperforms Nvidia's H100 for LLM Inference
blog.tensorwave.com
blog.tensorwave.com
I suggest taking the report with a grain of salt.
Fun weekend project for anybody.
AWS doesn't let you use p5 instances (not getting a quota as a private person), lambda cloud is sold out.
Also, stuff like this is hard to take the results seriously:
* To make an accurate comparison between the systems with different settings of tensor parallelism, we extrapolate throughput for the MI300X by 2.
* All inference frameworks are configured to use FP16 compute paths. Enabling FP8 compute is left for future work.
They did everything they can to make sure AMD is faster.They should probably show separately the throughput per completion as the tensor parallelism is often used for that purpose in addition to the doubling the VRAM.
I think that'd give us a better idea of perf/cost and whether multiplying MI300X results by 2 is justified.
The do the standard AMD comparison:
8x AMD MI300X (192GB, 750W) GPU
8x H100 SXM5 (80GB, 700W) GPU
The fair comparison would be against 8x H100 NVL (188GB, <800W) GPU
Price tells a story. If AMD performance would be in par with Nvidia they would not sell their cards for 1/4 price.Explain why the performance difference does not matter?
AMD does only 33% better with a chip that has 2X transistors and 2X memory.
nearly 95% of deeplearning github repos are "tested using cuda gpu - others, not so sure"
the only way to out-run nvidia is to have 3~10x better bang-for-buck.
Or AMD can just provide a "DIY unlimited gpu RAM upgrade" kit -- a lot of people are buying macstudio 128gb ram because of its "bigger ram-for-buck" than nvidia gpus
MTr
------------------
H100 SXM5 80,000
MI300X 153,000
H100 NVL 160,000
H100 SXM4 has 52% of the transistors MI300X has, half of the RAM and MI300X achieves *ONLY* 33% higher throughput compared to the H100. MI300X was launched 6 months ago, H100 20 months ago.AMD has work to do.
At best, Apple has Metal API which iOS video games use. I guess there's a level of SIMD-compute expertise here, but it'd take a lot of investment to turn that into a full scale GPU that tangos with Supercomputers. Software is a bit piece of the puzzle for sure, but Metal isn't ready for prime time.
I'd say Apple is ahead of Intel (Intel keeps wasting their time and collapsing their own progress from Xeon Phi / Battlemage / etc. etc. Intel cannot keep investing in its own stuff to reach critical mass). Intel does have OneAPI but given how many times Intel collapses everything and starts over again, I'm not sure how long OneAPI will last.
But Apple vs AMD? AMD 100% understands SIMD compute and has decades worth of investments in it. The only problem with AMD is that they don't have the raw cash to build out their expertise to cover software, so AMD has to rely upon Microsoft (DirectX), Vulkan, or whatever. ROCm may have its warts, but it does represent over a decade of software development too (especially when we consider that ROCm was "Boltzmann", which had several years of use before it came out as ROCm).
-------
AMD ain't perfect. They had a little diversion with C++Amp with Microsoft (and this served as the API for Boltzmann / early ROCm). But the overall path AMD is making at least makes sense, if a bit suboptimal compared to NVidia's huge efforts into CUDA.
I'd definitely rate AMD's efforts above Apple's Metal.
M3 Max's GPU is significantly more efficient in perf/watt than RDNA3, already has better ray tracing performance, and is even faster than a 7900XT desktop GPU in Blender.[0]
[0]https://opendata.blender.org/benchmarks/query/?compute_type=...
The M3 Max is also in a sense a generation ahead in terms of perf/watt of the 7900 XT as it uses a newer manufacturing node.
I suppose it's also worth highlighting that if you enable Optix in the comparison above, you can see Nvidia parts stomping all over both AMD and Apple parts alike.
M3 Max GPU uses at most 60-70w. Meanwhile, the 7900XT uses up to 412w in burst mode.[0] TSMC N3 (M3 Max) uses 25-30% less power than TSMC N5 (7900XT). [1] In other words, if 7900XT used N3 and optimizes for the same performance, it would burst to 300w instead which is still 5-6x more than M3 Max. In other words, the perf/watt advantage of the M3 Max is mostly not related to the node used. It's the design.
[0]https://www.techpowerup.com/review/amd-radeon-rx-7900-xt/37....
[1]https://www.anandtech.com/show/18833/tsmc-details-3nm-evolut...
The article is MI300X, which is beating NVidia's H100.
> Do you have benchmarks for when AMD doesn't nerf Blender performance?
Go read the article above.
> Notably, our results show that MI300X running MK1 Flywheel outperforms H100 running vLLM for every batch size, with an increase in performance ranging from 1.22x to 2.94x.
-------
> Why does AMD nerf RDNA3 when they're so far behind Nvidia and Apple in Blender performance?
Nerf is a weird word.
AMD has focused on 32-bit FLOPs and 64-bit FLOPs until now. AMD never put much effort into raytracing. They reach acceptable levels on XBox / PS5 but NVidia always was pushing Raytracing (not AMD).
Similarly: Blender is a raytracer that uses those Raytracing cores. So any chip with substantial on-chip ray-tracing / ray-matching / ray-intersection routines will perform faster.
Blender isn't what people do with GPUs. The #1 thing they do is video games like Baldur's gate 3.
-------
It'd be like me asking why Apple's M3 can't run Baldur's gate 3. Its not a "nerf", its a purposeful engineering decision.
I haven't done a head to head and I suppose it depends on whether tensor parallelism actually scales linearly or not, but my understanding is since the NVL's are just PCIe/NVLink paired H100s, you're not really getting much if any benefit on something like vLLM.
I think the more interesting thing critique might be the slightly odd choice of Mixtral 8x7B vs say a more standard Llama2/3 70B (or just test multiple models including some big ones like 8x22B or DBRX.
Also, while I don't have a problem w/ vLLM, as TensorRT gets easier to set up, it might become a factor in comparisons (since they punted on FP8/AMP in this tests). Inferless published a shootoff a couple months ago comparing a few different inference engines: https://www.inferless.com/learn/exploring-llms-speed-benchma...
Price/perf does tell a story, but I think it's one that's mostly about Nvidia's platform dominance and profit margins more than intrinsic hardware advantages. On the spec sheet MI300X has a memory bandwidth and even raw FLOPS advantage but so far it has lacked proper software optimization/support and wide availability (has anyone besides hyperscalers and select partners been able to get them?)
Profit margins and dominance are result from performance, not the other way around.
It does not matter if Nvidia tools are better when you deploy large number of chips for inference and it does more flops per watt or second. It's seller market and if AMD can't ask high price, their chip do not perform.
----
Question:
People here seem to think that Nvidia has absolutely no advantage in their microarchitecture design skills. It's all in software or monopoly.
Is this right?
That's an extrapolation. Microarchitecture design skills are not theoretical numbers you manage to put on a spec sheet. You cannot decouple the software driving the hardware - that's not a trivial problem.
not only can you measure this, not only do they measure this, but it's literally the first component of the Rayleigh resolution equation and everyone is constantly optimizing for it all the time.
https://youtu.be/HxyM2Chu9Vc?t=196
https://www.lithoguru.com/scientist/CHE323/Lecture48.pdf
in the abstract, why does it surprise you that the semiconductor industry would have a way to quantify that?
like, realize that NVIDIA being on a tear with their design has specifically coincided with the point in time when they decided to go all-in on AI (2014-2015 era). Maxwell was the first architecture that showed what a stripped-down architecture could do with neural nets, and it is pretty clear that NVIDIA has been working on this ML-assisted computational lithography and computational design stuff for a while. Since then, I would say - but they've been public about it for several years now (and might be longer, I'd have to look back).
https://www.newyorker.com/magazine/2023/12/04/how-jensen-hua...
https://www.youtube.com/watch?v=JXb1n0OrdeI&t=1383s
Since that "mid 2010s" moment, it's been Pascal vs Vega, Turing (significant redesign and explicit focus on AI/tensor) vs RDNA1 (significant focus on crashing to desktop), Ampere vs RDNA2, etc. Since then, NVIDIA has almost continuously done more with less: beaten custom advanced tech like HBM with commodity products and small evolutions thereupon (like GDDR5X/6X), matched or beaten the efficiency of extremely expensive TSMC nodes with junk samsung crap they got for a song, etc. Quantitatively by any metric they have done much better than AMD. Like Vega is your example of AMD design? Or RDNA1, the architecture that never quite ran stable? RDNA3, the architecture that still doesn't idle right, and whose MCM still uses so much silicon it raises costs instead of lowering them? Literally the sole generation that's not been a total disaster from The Competition has been RDNA2, so yeah, solid wins and iteration is all it takes to say they are doing quantitatively better, especially considering NVIDIA was overcoming a node disadvantage for most of that. They were focused on bringing costs down, and frankly they were so successful despite that that AMD kinda gave up on trying to outprice them.
Contrast to the POSCAP/MLCC problem in 2020: despite a lot of hype from tech media that it was gonna be a huge scandal/cost performance, NVIDIA patched it dead in a week with basically no perf cost etc. Gosh do you think they might have done some GPGPU accelerated simulations to help them figure that out so quickly, how the chip was going to boost and what the transient surges were going to be etc?
literally they do have better design skills, and some of it is their systems thinking, and some of it is their engineers (they pay better/have better QOL and working conditions, and get the cream of the crop), and some of it is their better design+computational lithography techniques that they have been dogfooding for 3-4 generations now.
people don't get it: startup mentality, founder-led, with a $3t market cap. Jensen is built different. Why wouldn’t they have been using this stuff internally? That’s an extremely Jensen move.
AMD's PE is ~55. Nvidia's PE is above 70.
What were your thoughts on Zen (1) vs Intel's offerings then? AMD offered more back for the buck then too.
I don't think it should be ignored, especially when the power consumption is similar.
if so ok it's fair to compare 1 mi300x with 1 h100 NVL but then price ( and tco ) should be added to the some metrics conclusion , also the NVL is a 2xpci5.0 quad slot , so not the same thing..
I am not sure about system compatibility and if and how you can stack 8 of those in one system ( like you can do with non NVL and mi300x.. ) so it's a bit a diffent ( and more niche ) beast
What would be a suitable input length in your oppinion?
And why isnt this a good one: Are real-life queries shorter? Or longer?
If i count one word as a token, then in my case most of the queries are less than 128 words.
If I understood that correctly, context length is something like session storage or short term memory. If it's too small the AI starts to forget what it's talking about.
It's not just the query (if you're running a chatbot, which many of us are not). It's the entire context window. It's not uncommon to have a system prompt that is > 512 tokens alone.
I would like to see benchmarks for 512, 1024, 4096 and 8192 token inputs.
I tried to look for some service provider to publish this kind of metrics, but haven't found any.
But there's a long list of German companies not on the DAX
(though Germany DAX really deserves to be worth less than NVidia)
Not to be too nitpicky here but these are only the publicly traded companies. You have a number of pretty large German companies that are still entirely private such as Aldi, Schwarz Group, Boehringer or Bosch.
https://www.famcap.com/top-500-german-family-businesses-the-...
Not all of those are listed, or listed in Frankfurt
While I would enjoy a US tech salary, I'm not sure we want a world where all manufacturing is set aside to focus on the attention economy.
Nvidia value deserves to be much higher than any company on the DAX (maybe all of them together, as it currently is) - but how much of that current value is real rather than an AI speculation bubble?
Nvidia sells chips ...
The reason Nvidia's value has been so inflated is the software stack and the lock-in they offer. CUDA, CuDNN, that's where Nvidia's value lies.
And obviously, now that all relevant ML frameworks are designed for Nvidia's software stack, Nvidia has a monopoly on the supply. That's why their value is being inflated so much.
And Nvidia doesn't have produce the chips themselves, that's all contracted out as well.
This, as the kids say, is just cope. American big tech makes real products. Google is not just ads. Apple is not. Amazon is not. Tesla is not. NVidia is not. Netflix is not.
NVidia might be overvalued because of the current AI hype but that does not diminish their real accomplishments!
Europe has almost no real tech companies. There is one exception, founded in 1984. Not exactly a spring chicken. How can a wealthy continent with 750 million people produce no big tech companies? It's a big problem.
Much more difficult to scale a product across 26 different countries and nearly as many languages and regulatory jurisdictions. US is one country, not a collection of countries fighting each other, meaning your product is instantly available to 300M people speaking the same language under (nearly) the same regulations.
It's a single market on paper as the eu only mandates a small subset of common rules and regulations such as removing tarrifs or freedom of movement, but have you ever tried in practice to launch your company from Belgium to France or from Netherlands to Belgium or from Austria to Germany, or from Romania to Italy?
It's much more difficult when the rubber hits the road as every country has various extra laws and protectionist measures in place to protect it's domestic players from outsiders even if they came from within the EU. And that's besides the language barrier which means added costs. This is much less efficient than the US market.
EU countries and voters still value their national sovereignty and culture (both with the upsides and downsides) above a united EU under the same laws and language for everyone, ruled from outside their country's borders. See what happened with Brexit and the constant internal squabbling and sabotaging over critical EU issues that affect us all like the war in Ukraine or illegal mass migration. An US style unification just won't work here since every little country wants to be it's own king while having its cake and eating it too.
California has more burdensome regulations and higher taxes than other states and yet it's home to silicon valley.
We're talking about scaling internal companies across EU, not about imports and exports. And scaling local start-up across the EU is a regulatory and legal nightmare for small companies.
Shipping and selling imports and exports of commodities are a solved problem for decades, but scaling a on-line notary service for instance, that works both in Germany and in Italy, isn't. The EU doesn't help much with that as they only say you should have no tariffs between each other, not that you shouldn't have various legal, cultural and bureaucratic protectionism idiosyncrasies in place. The EU won't and can't force countries to improve that to make doing business easier for cross-country start-ups.
EU countries have a lot more roadblocks between each others than US states do when ti comes to scaling businesses.
Obviously there is a lot to criticize about the EU and I can offer you a gigantic list there too. However, I do not see any clear failure of the EU’s approach as a single market so far. Additionally part of the philosophy was establishing peace in a region that was torn up by wars for a lot longer than Christianity exists. I would argue the EU was quite successful there too.
The US is a republic of 50 states. Each state has a huge amount of sovereignty and autonomy. There are 50 state-level regulatory jurisdictions. Not to mention the local-level of government.
But in spite of this, the US does not over-regulate. This is the big difference to Europe (I say this as an American expat living in Europe).
Currently EU welfare systems are under massive strain and huge waiting lists due to ageing population and economy that hasn't kept up to fund it.
There's no free lunch here. You need big companies with scale that pay huge wages as those mean a lot more tax revenue. Saying no to that kind money out of some made up idealism is just silly copium.
The EU income taxes paid by a single FANG salary employee would be the equivalent of the taxes paid by ~10 average workers. Pretty sure Germany and every other EU country would like to have such taxpayers contributing into the welfare system and not say no to it.
Europe's share of global GDP gas been on a constant decline at the expense of US and Chinese growth. Yeah it's nice to have a better welfare system than China or the US, but how will you fund it in the future if you keep having less money? Political idealism doesn't pay your food and rent.
European economic production is nowhere near high enough and now Europe is struggling to provide for its aging population and doesn't have enough good jobs for younger people. I support redistribution generally, but the wealth has to be created first or there won't be anything to redistribute.
Also in terms of tech innovation: What part of the US-based tech innovation couldn't have been (and actually were) achieved with open-source solutions many many years earlier for a fraction of the cost, if we didn't have copyright?
Honestly, a significant chunk of the "innovation" seems to relate directly to maximizing advertisement opportunities and inducing increased consumption. Who cares if a website takes a second to load rather than 0.1 seconds? If it has content I want, 1 second isn't a big deal. If I don't care about the content, I lose nothing by being distracted by something else in that 1 second.
---
More importantly:
European Economic production isn't high enough... by what standard?
https://data.oecd.org/lprdty/gdp-per-hour-worked.htm
GDP per hour worked is 74 in the US vs 69 in Germany and 54 in the EU. And the EU includes many large countries that emerged from communist dictatorship only 35 years ago, and are very much still in the process of catching up. Incidentally, the German economy is the result of the West German economy with 63 million people absorbing a failing economy hosting 16 million people in 1990.
The idea that the US is some promised land of economic prosperity while Europe is falling is entirely absurd. It's a narrative built on small relative differences and a US system that pressures people into working a lot more than Europeans do.
More importantly, even GDP per Capita wise:
https://data.oecd.org/gdp/gross-domestic-product-gdp.htm
EU per Capita GDP in 2022 is the same as USA 2016. Was the USA in 2016 struggling but now isn't?
This is all bullshit. Economic output is more than high enough and rising steadily. The problem remains solely in the distribution of
Absolutely the US would be struggling if the GDP was still 2016 values with today’s costs.
So it boggles my mind that your parent tried to make a point by equating USA 2016 with EU 2022 GDP/capita as if nothing's wrong with that. Are some people that oblivious?
Nominal: https://data.oecd.org/gdp/gross-domestic-product-gdp.htm
PPP/Inflation adjusted not so much: https://ourworldindata.org/grapher/gdp-per-capita-worldbank?...
But pretty much all countries are still well ahead of where we were in 2017. If you feel poorer than in 2017 it's because you're getting less of a larger pie, not because the economy is producing less than it did then.
> More importantly, even GDP per Capita wise: > > https://data.oecd.org/gdp/gross-domestic-product-gdp.htm > > EU per Capita GDP in 2022 is the same as USA 2016. Was the USA in 2016 struggling but now isn't? > > This is all bullshit. Economic output is more than high enough and rising steadily. The problem remains solely in the distribution of the ouptut.
The conclusion that economic output has been rising steadily, even in the last couple of years, is true. But the 2016 vs 2022 numbers are nominal, thus useless. There is a much more significant difference over time when working in PPP/Inflation adjusted numbers:
https://ourworldindata.org/grapher/gdp-per-capita-worldbank?...
The overall point holds though: The economic output of the EU is at 45K per capita today, the level of the US in 1997. The US was not a poor country in 1997. Germany is at the economic output per capita of 2009.
Did the US in 1997 suffer from the problem that it didn't produce enough economic output? Of course not.
And given that, adjusting for inflation, GDP per capita is at an all-time high, the conclusion that you're poorer because economic production is distributed to others is necessarily true. And it tracks, too. Corporate profits and the Dow Jones are not down. The already extremely wealthy have accumulated nearly two thirds of the new wealth being created since 2020:
https://www.oxfam.org/en/press-releases/richest-1-bag-nearly...
> Billionaire wealth surged in 2022 with rapidly rising food and energy profits. The report shows that 95 food and energy corporations have more than doubled their profits in 2022. They made $306 billion in windfall profits, and paid out $257 billion (84 percent) of that to rich shareholders. The Walton dynasty, which owns half of Walmart, received $8.5 billion over the last year. Indian billionaire Gautam Adani, owner of major energy corporations, has seen this wealth soar by $42 billion (46 percent) in 2022 alone.
Given these facts, if we have to accept lower economic production in the name of a fairer distribution of economic production, that seems more than acceptable to me.
But those are all low-marin chips. Qualcomm, Nvidia, Intel, AMD and Apple have much higher margins on their chips. They don't bother competing with the EU chips companies.
https://pbs.twimg.com/media/F3PGpsrWEAEiplB?format=jpg&name=...
Unless you want to say that the US was much poorer in 2017 than it was in 2022 that's a fairly ridiculous statement.
Also, the highest productivity places in the EU have much lower hours worked per capita than the US, with Germans on average working 25% less than Americans and the EU as a whole working 13% less than the US.
European Union gdp per capita for 2022 was $37,433, a 3.33% decline from 2021.
U.S. gdp per capita for 2022 was $76,330, a 8.7% increase from 2021.
It's not even close?
For example Ireland has by a long margin the highest GDP/capita in the whole EU, and it would make you think the average Irish worker earns more that any other worker in the EU and drives a Lambo, but that's not what's happening. It's because most US corporations funnel their EU money through their Irish holding companies skewing the statistic.
No, Luxembourg does.
World bank data in PPP dollars is reported as 64,600 vs 45,900 here:
https://ourworldindata.org/grapher/gdp-per-capita-worldbank?...
Germany is at 53,900 there, but a good chunk of the difference is simply that US works more per capita. GDP per hour worked is 74$ in the US vs 69$ in Germany, 53$ in Canada. Sweden is ahead of the US. And the EU also includes countries like Bulgaria, which at 29$ is barely ahead of Russias 28$.
https://data.oecd.org/lprdty/gdp-per-hour-worked.htm
France is at 65$ per hour worked, but Germany and France also have significantly lower poverty and inequality rates by any measure you chose, with France more equal than Germany.
The US, of course, remains the dominant economy of the world by any measure. There is no question of that. But the exponential nature of economics, and the structural differences between these different economies, means that GDP numbers compared directly are fairly meaningless.
Edit: That last sentence is too strong as stated. GDP obviously matters a big deal in the grand scheme of things, especially as you jump from lower or middle income to high income countries. But it's all logscale. A factor of 2 is a big deal, a factor of 1.2 might not be.
Can't edit anymore, but: That were nominal numbers, and thus useless. See here:
The only real answer is protectionism, and there's a good chance that'll hurt more than it helps.
Bad example given how aggressively they terminate products which don't generate the same revenue as ads.
> Apple is not.
Best example, they have done a fantastic job of being both a tech company and pseudo-fashion company.
> Amazon is not.
They don't make anything (at least nothing people want to buy) and have ad revenue as an increase slice of their pie.
> Tesla is not.
Even bigger hype/speculation vehicle than Nvidia.
> NVidia is not.
Nvidia of 5 years ago would not have appeared on this list, being too much of a niche tech company. Good at what they do, but hugely hype-fuelled.
> Netflix is not.
Running out of growth potential with their current business model, starting to introduce ads!
For all his insanity, the one thing I respect Musk for, is that he actually started successful companies that make stuff. Creating a new car manufacturer of the scale of BMW out of nothing was widely considered impossible before.
Of course he did this from a position of extreme wealth, but none of his peers managed to do that. Everyone else is just seeking rent by trying to be first to implement some tech transition that is coming anyway. And that might be a lot more valuable to society if it was managed differently...
Not to be snarky, but if AWS counts as “nothing” I’d sure like a slice of nothing please.
If I pay for a database server in Virginia, how is that not real?
Cloud services don’t just exist on their own accord. Datacenters are physical and real!
How do you define "tech"? Europe's domestic markets are jam-packed full of local tech companies.
Europe has many small and not very profitable tech companies. Almost no large and profitable ones. https://pbs.twimg.com/media/GNDtCtTXcAAiwFk?format=jpg&name=...
Most of them are just payment middlemen not some innovative product nobody else can do, and Spotify survives on monopolizing and squeezing artists, not some innovative product. Kind of like Netflix except Netflix has some cutting edge streaming tech as a product not just IP licenses.
ASML is the only product innovator there except their innovative EUV lightsources are licensed from Sandia labs in the US and made by Cymer in the US which ASML bought and licensed to not seel to China. So an US invention at the end of the day.
Adyen is very underrated, and Spotify is definitely tech.
Stripe should be on the list. DeepMind at one point.
It's just much cheaper and easier for start-ups if you're developing a SW product to sell it in the US market first and only when you've made money there, slowly bring it in the EU.
Starting off SW products in the EU is suicide (unless you're targeting some niche in the local market that's safe from competitors from abroad because it ties into some local idiosyncrasies on language, culture and law).
I mean, by definition given that it trades freely their market cap is real. Your market cap today is what the market thinks your future cash flows are worth. The bubble and the bubble popping should in theory both be priced into Nvidia's market cap.
What isnt' is events the market doesn't anticipate, AMD coming out with a current generation chip that can do inference as well as the H100 is something the market hasn't priced in.
Andy our manufacturing example is very poor as NVidia is certainly part of the manufacturing pipe line by designing physical products that people buy.
I think probability of that would still be priced in. Not sure what the exact probability is, though.
But if say it was clear that AMD can come up with a competitive option, then NVDA stock would drop. But if it was clear the other way that AMD can't do it, NVDA price would increase.
This is a bit of a tired viewpoint, and is evidently proven not true time after time. The collective despair/euphoria of market participants is extremely powerful and well documented, at least as far back as dutch tulips.
Stock valuations are relative, and they are relatively misvalued most of the time. That's why there are (albeit rare) funds that are capable of outperforming the market for decades - Berkshire, and Medallion for example.
It's certainly possible that AMD is valued (almost) fairly. It's just as likely that it's relatively misvalued for no reason other than emotions (lack of hype).
Nvidia sells physical things, and they are bigger than 40 companies because the companies are selling physical things?
I am not arguing hardware scales better than software but this is a strange argument in this context.
Nah, then ill get my very good wagie pennies here and have plenty jobs available, plus good health insurrance and whatnot.
https://www.bls.gov/news.release/empsit.nr0.htm https://www.destatis.de/EN/Press/2024/06/PE24_217_132.html
Please stop breaking HN rules. I never said that. HN rules state you need to reply to the strongest interpretation of someone's argument, not the weakest that's easiest to criticize.
I just pointed out once country's economics performance for comparison, if you're want to extrapolate from that that you should move there, that's your issues to deal with, but not my argument.
Nvidia problem will sort itself out naturally in the coming months/years.
Jensen isn't stupid. He's making accelerators for anything so that they'll be ready to catch the next bubble that depends on crazy compute power that can't be done efficiently on CPUs. They're so far the only semi company beating Moore's law by a large margin due to their clever scaling tech while everyone else is like "hey look our new product is 15% more efficient and 15% more IPC than the one we launched 3 years ago".
They may be overvalued now but they definitely won't crash back to their "just gaming GPUs" days.
I'm tired of hearing about Nvidia's "luck". There was no luck involved. Nvidia shiped Cuda on consumer GPUs since 2006. That's almost 20 years time researchers had to find used cases for that compute and Nvidia made it possible. In other words the AI bubble happened because Nvidia made the necessary ground work for it to happen, they didn't just fall into it by luck.
Sure, they didn't get lucky with the tech they had to offer - that was well developed for years. They just got lucky that the next big thing was compute-based. If the next thing is memory/storage-based, they're screwed and the compute market is saturated for years - they have only gamers left.
If all this extra computing power is available, smart people will find a way to use it somehow.
Of course the equivalent can happen to Nvidia. Seems almost certain.
Nvidia could just as easily triple in short order as get cut in half from here imho.
[0]https://www.macrumors.com/2024/04/23/apple-cuts-vision-pro-s... [1] https://www.macrumors.com/2024/04/22/apple-vision-pro-custom...
I predict this is yet another domain rich with opportunity for AI.
Or how about building the world virtually? https://www.nvidia.com/en-us/high-performance-computing/eart...
Or, leaning more towards your examples, a Grand Theft Auto style environment containing millions of them, except life like.
Also curious how many companies are dropping that much money on those kind of accelerators just to run 8x 7B param models in parallel... You're also talking about being able to train a 14B model on a single accelerator. I'd be curious to see how "full-accelerator train and inferrence" workloads would look ie: Training a 14B param model then inferrence throughput on a 4x14B workload.
AMD (and almost every other inferrence claim maker so far... intel and apple specifically) have consistently cherry picked the benchmarks to claim a win over, and ignored the remainder which all show nvidia in the lead - and they've used mid-gen comparison models as many commenters here pointed out in this article.
in a single system ( 8x accelerators ) LLMs, mi300x has very competitive inference TCO vs h100 .
also :
AMD Instinct MI300X Offers The Best Price To Performance on GPT-4 According To Microsoft, Red Team On-Track For 100x Perf/Watt By 2027
https://wccftech.com/amd-instinct-mi300x-best-price-performa...
and with a growing but certainly less mature product ( expecially software ), it requires suitable pricing and allocation strategies
1. https://www.techspot.com/news/102056-nvidia-allegedly-punish...
amd is successfully attacking the inference sector, increasing its advantage with mi325 and aiming for training from 2025 with mi350 (and Infinity Fabric interconnect and other types of interconnection that are arriving for the various topologies), which will probably have an advantage over blackwell, and then fall back against rubin and come back ahead against mi400,
at least, this is what it seems, and as long as the rocm continues to improve.
Personally I am happy to see some competition in the sector and especially on open source software
boo boo, a GTX 670 that cost you $399 in 2012 now costs $599 - grow up, do the inflation calculation, and realize you’re being a child. gamers get the best deal on bulk silicon on the planet, R&D subsidized by enterprise, fantastic blue-sky research that takes years for competitors to (not even) match, and it’s still never enough. ”Gamers” have justified every single cliche and stereotype over the last 5 years, absolutely inveterate manbabies.
(Hardware Unboxed put out a video today with the headline+caption combo “are gamers entitled”/“are GeForce gpus gross”, and that’s what passes for reasoned discourse among the most popular channels. They’ve been trading segments back and forth with GN that are just absolute “how bad is nvidia” “real bad, but what do you guys think???” tier shit, lmao.
https://i.imgur.com/98x0F1H.png
this stuff is real shit, nvidia has been leaning on partners to maintain their segmentation, micromanaging shipment release to maintain price levels (cartel behavior), punishing customers and suppliers with “you know what will happen if you cross us”, literally putting it in writing with GPP (big mistake), playing fuck fuck games with not letting the drivers be run in a datacenter, etc. You see how that’s a little different than a gpu going from an inflation-adjusted $570 to $599 over 10 years?
(And what’s worse the competition can’t even keep that much, they’re falling off even harder now that Moores law has really kicked the bucket and they have to do architectural work every gen just to make progress, instead of getting free shrinks etc… let alone having to develop software! /gasp)
In entirely unrelated news… gigabyte suddenly has a 4070 ti super with a blower cooler. Oh, and it’s single-slot with end-fire power connector. All three forbidden features at once - very subtle, extremely law-abiding.
https://videocardz.com/newz/gigabyte-unveils-geforce-rtx-407...
and literally gamers can’t help but think this whole ftc case is all about themselves anyway…
large orders for those accelerators are placed months ahead
meanwhile mi300x on microsoft are fully booked...
https://techcommunity.microsoft.com/t5/azure-high-performanc...
"Scalable AI infrastructure running the capable OpenAI models These VMs, and the software that powers them, were purpose-built for our own Azure AI services production workloads. We have already optimized the most capable natural language model in the world, GPT-4 Turbo, for these VMs. ND MI300X v5 VMs offer leading cost performance for popular OpenAI and open-source models."
According to the article: """ AMD Configuration: Tensor parallelism set to 1 (tp=1), since we can fit the entire model Mixtral 8x7B in a single MI300X’s 192GB of VRAM.
NVIDIA Configuration: Tensor parallelism set to 2 (tp=2), which is required to fit Mixtral 8x7B in two H100’s 80GB VRAM. """
Everybody thinks it’s CUDA that makes Nvidia the dominant player. It’s not - almost 40% of their revenue this year comes from mega corporations that use their own custom stack to interact with GPUs. It’s only a matter of time before competition catches up and gives us cheaper GPUs.
thats just a question of negotiating with tsmc or their few competitors
(also didn't tsmc start production of some factories in the US and/or EU?)
I mean, nvidia use tsmc, so does amd.
But now that there’s a larger incentive to produce GPUs, their moat will eventually fall.
TSMC runs at 100% capacity for top tier processes - their bottleneck is more foundries. These take time to build. So the question becomes - how long can Nvidia remain dominant? It could be quarters or it could be years before any real competitor convinces large customers to switch over.
Microsoft and Google are producing their own AI hardware too - nobody wants to depend solely on Nvidia, but they’re currently forced to if they want to keep up.
alright fine it's the codegen and the runtime and the driver and the library ecosystem...
> If everyone started emitting IR and didn't rely on NVidia-owned libs then the space would become unrecognizable.
I have no clue what this means - which libs are you talking about here? the libs that contain the implementations of their runtime? or the libs that contain the user space components of their driver? or the libs that contain their driver and firmware code? And exactly which of these will "everyone emitting IR" save us from?
lol completely made up.
are you conflating CUDA the platform with the C/C++ like language that people write into files that end with .cu? because while some people are indeed not writing .cu files, absolutely no one is skipping the rest of the "stack" (nvcc/ptx/sass/runtime/driver/etc).
source: i work at one of these "mega corps". hell if you don't believe me go look at how many CUDA kernels pytorch has https://github.com/pytorch/pytorch/tree/main/aten/src/ATen/n....
> Everybody thinks it’s CUDA that makes Nvidia the dominant player.
it 100% does
NVidia relies on TMSC for manufacturing. Samsung is building competing manufacturing infrastructure which is also a good thing, so Taiwan is not a single point of failure.
95% would be nice too
https://www.reddit.com/r/AMD_MI300/comments/1dgimxt/benchmar...
> MI300X Accelerator: 192GB VRAM, 5.3 TB/s, ~1300 TFLOPS for FP16
> Hardware: Baremetal node with 8 H100 SXM5 accelerators with NVLink, 160 CPU cores, and 1.2 TB of DDR5 RAM.
> H100 SXM5 Accelerator: 80GB VRAM, 3.35 TB/s, ~986 TFLOPS for FP16
I really wonder about the pricing. In theory the MI300X is supposed to be cheaper, but whether is that is really the case in practice remains to be seen.
So, probably around the same price?
The tests look promising, though!
The weird thing on Runpod is the virtual CPUs, you can't run MI300x in virtual machines yet. It is a missing feature that AMD is working on.
probably this.
If they need a ChatBot that uses a model with same accuracy and performance as on non-CUDA hardware, would they still want CUDA based hardware?
https://www.amd.com/en/newsroom/press-releases/2024-5-21-amd...
It looks like the price to performance of inference tasks gives providers a big incentive to move away from Nvidia.
1. They're only comparing against VLLM, which isn't SOTA for latency-focused inference. For example, their vllm benchmark on 2 GPUs sees 102 tokens/s for BS=1, gpt-fast gets around 190 tok/s. https://github.com/pytorch-labs/gpt-fast 2. As others have pointed out, they're comparing H100 running with TP=2 vs. 2 AMD GPUs running independently.
Specifically,
> To make an accurate comparison between the systems with different settings of tensor parallelism, we extrapolate throughput for the MI300X by 2.
This is uhh.... very misleading, for a number of reasons. For one, at BS=1, what does running with 2 GPUs even mean? Do they mean that they're getting the results for one AMD GPUs at BS=1 and then... doubling that? Isn't that just... running at BS=2?
3. It's very strange to me that their throughput nearly doubles going from BS=1 to BS=2. MoE models have an interesting property that low amounts of batching doesn't actually significantly improve their throughput, and so on their Nvidia vllm benchmark they just go from 102 => 105 tokens/s throughput when going from BS=1 to BS=2. But on AMD GPUs they go from 142 to 280? That doesn't make any sense to me.
That info is conspicuously absent from the article.
The H100 rents for about $4.5/hr consuming 0.7kWh in that hour which will likely cost them less than 7 cents.
That just says you don't run a cloud for profit :)
https://www.reddit.com/r/AMD_MI300/comments/1dgimxt/benchmar...
Maybe the benchmark should be performance per $... though I suspect power consumption will eclipse the cost of purchasing the chips from NVDA or AMD (and costs of chips will vary over time and with discounts). EDIT: was wrong on eclipsing; still am looking for a more durable benchmark (performance per billion transistors?) given it's suspected NVDA's chips are over-priced due to demand outstripping supply for now, and AMD's are under- to get a foothold in this market.
Making AMD work effortlessly with pytorch et al should make the switch transparent.
Also, the price difference is not quantified.
Additionally, CUDA is a known and tangible software stack - can I try out this "MK1 FLywheel" on my local (AMD) hardware?
For consumer grade inference, there's already many options available.
They also used Flywheel for AMD while not bothering to turn on Flywheel for Nvidia, which is crazy since Flywheel improves Nvidia performance by 70%. https://mk1.ai/blog/flywheel-launch
In this context the 33% performance lead by AMD looks terrible, and straight up looks slower.
This is a new AMD vs last generation nvidia benchmark.
https://www.theregister.com/2024/03/21/nvidia_dgx_gb200_nvk7...
MI300X launched 3 months earlier at the end of December.
H100 launched March 2023,
Also Blackwell's lead will be short lived, because mi350x is coming out next year, and it will have a node and architecture advantage. So AMD will be ahead again.
(Otherwise it's apples and oranges)