Discussion: https://news.ycombinator.com/item?id=39344815
Much more sensible to work on getting rock solid support for their own standards into all the major ML platforms/libraries.
Nvidia and standard-maker is limited in what breaking changes they introduce - these can harm their customers as much as they harm the competition. Intel failed to force all their changes on AMD as the xxx86 market expanded (notably, the current iteration of CPUs standards was set by AMD after Intel was unable to sell their completely new standard).
Still, I'd acknowledge that "business sense" today follows the approach of only aiming for markets the company can completely control and by that measure, CUDA compatibility isn't desirable.
The key is that while there were many clones of x86, there never really was an attempt at a company built around "run MS Windows programs natively" because maintaining software compatability is an order of magnitude harder than doing it for hardware.
Moreover, companies aren't buying GPUs to keep their huge stable of legacy applications running. They want to create new AI applications and CUDA is a simple API for doing that (at a certain level).
Microsoft's entire history is around building a moat of APIs because the PC software industries has a wide variety. Nvidia has, so far, been focused on building actually useful things for developers. Basically, where all the other manufacturers viewed their chips as special purpose devices, Nvidia allowed developers to treat their chips as generic parallel processors and this facilitated the current AI revolution/bubble. Now that Nvidia has created this market, it can charge by the use rather than charging by processing power. The thing is that Nvidia's large potential competitors simply don't want to create clones even if they could - because clones would have to be sold by processing power rather than with a markup for their usefulness. It's worth looking at the list of x86 compatible makers [1]. Making an x86 wasn't quite something you could do in your garage but clearly the barriers to entry weren't huge. But any Nvidia compatible is going to cost a large amount of capital but can only sell by processor power and so AMD, Intel and similar sized entities don't have an interest in doing this.
https://ir.amd.com/news-events/press-releases/detail/1206/am...
“Silo AI has been a pioneer in scaling large language model training on LUMI, Europe’s fastest supercomputer powered by over 12,000 AMD Instinct MI250X GPUs,”
If they’ve trained LLMs with lumi which has a lot of instinct GPUs there is a high chance they’ve had to work through and solve a lot of the gaps in software support from AMD.
They may have already figured out a lot of stuff and kept it all proprietary and AMD buying them out is a quick way to get access to all the solutions.
I suspect AMD is trying to fast track their software stack and this acquisition allows them to do just that.
But isn't getting a software stack the exact kind of thing they need? Is there no overlap in the skills at the purchased company and the skills needed to make the AMD software stack not suck?
I said that I interpreted the previous comment as sarcastic so I could be called out if it wasn't. The author hasn't yet disagreed. And I think sarcasm is warranted in a space that has witnessed so many bad acquisitions.
On software at AMD; if my world is so simple, please explain where I am wrong. I never said this was a simple solution, I implied there was some overlap needed skills.
ROCm sucks, it has licensing and apparently use issues. It has had performance issues, and that is getting better. It isn't in a lot of the places it needs to be where it could be considered a default choice.
Apparently, Silo uses AMD stuff to do ML work. Apparently, they have domain experts in this space. It seems likely that getting input from such people could positively influence the ML and hardware.
Of course there will be complexity in this process. This is a 600 million dollar deal involving thousands of people (not just Silo employee, but AMD people, regulators, stakeholders, etc). I don't think anyone is implying this is simple.
I only wanted to say, "This isn't obviously dumb".
600 million dollars is a lot, and in order for that 12 billion increase to stick around this team up needs to present a lot of value. I'm optimistic but I'm also an outsider.
Think of it as a reverse McDonnell-Douglas.
It's certainly not Lisa Su's fault that the clowns over at Intel got stuck on variations of 14nm (with clever marketing names like 14nm+++++) for nearly a decade, but credit certainly is hers for introducing Zen and putting AMD back on top of the x86 market.
With the new x870(e) motherboards and Granite Ridge chips right around the corner, effortlessly destroying the pyrotechnic processing units known as Raptor Lake, it's honestly a miracle to me that Intel's stock price is still as high as it is.
Guess wall street still loves those billions of forcefully confiscated taxpayer dollars being doled out by Uncle Sam to a graying dinosaur like Intel who couldn't even compete without those handouts... the quality of their marketplace offerings certainly isn't what's keeping that valuation up!
> x86 is never going to reclaim the crown of most important architecture
To be clear, I assume you are including 32-bit and 64-bit, e.g., x86-64. I am surprised by this comment. To me, x86 won the architecture battle because of Linux (and less Microsoft Windows). Nothing is so cheap to deploy and maintain as a Linux server that runs x86-64 procs. Yes, I know you can buy single board computers, but x86 wins in the triangulation of dollars-watts-performance. If you disagree, what do you think is the most important architecture today?As time goes on, more and more ARM chips will take roles traditionally taken by x86-64. Hell, ARM is already the best-selling architecture. Laws of scale will dictate that investments in x86-64 will fail to keep pace with ARM. Apple Silicon is already showing a small fragment of that effect. The chips are incredibly competitive, and for what Apple has chosen to focus on (perf-per-watt), unbeatable. ARM investments by other companies are catching up to Apple, and x86-64 does not make enough money to reverse that trend.
Look at it this way; Apple can design their own cores whether they use ARM or RISC-V, and control the software from top to bottom either way. Nvidia's already shipping RISC-V microcontrollers to cut down on manufacturing margins, and without better options the rest of the world might follow. ARM's dominance is only possible if better RISC options don't exist; and for anyone that's not ARM the idea of IP serfdom sounds awful.
In my mind it's not about AI per se, but about using the hot use case for GPU to drive meaningful change in your software stack. There are tons and tons and tons of GPGPU users out there who aren't training LLMs but who need a high-quality compute stack.
There's everything to fix. AMD is sitting on a gold mine and is squandering massive amounts of money every month that they don't just get their shitty software stack in order.
AMD could be as rich as NVIDIA. Instead, Lisa Su for some insane reason refuses to build even the most mediocre ML-capable libraries for their GPUs.
If I could ask anyone in the ML world at the moment what the heck they're thinking, it would be her. Nothing makes sense about AMDs actions for years on this topic. If I was the board, I'd be talking about her exit for wasting such an opportunity.
Spending $665m on a company that builds AI tooling, is a refusal?
They'd get more value offering $500k+ comp to a few people from https://handmadecities.com/
It is a wonder why you aren't the CEO.