AMD's Next GPU Is a 3D-Integrated Superchip
spectrum.ieee.org
spectrum.ieee.org
It's 24 x64 cores and 228 CDNA compute units on a common 128GB of memory. Personally I want to run constraint solvers on one. The general approach of doing lots of integer/float work on the GPU and branchy work on the CPU, both hitting the same memory, feels like an order of magnitude capability improvement over the current systems.
AMD is coming out with a "strix halo" APU somewhat appropriate for compute.
You would save money on improved yields of chiplets and on using cheaper nodes where appropriate. I imagine that would help offset the increased packaging costs.
However, they are moving to compete in a market where the profit margins are so obscene that a small increase in the cost of materials really wouldn't be an important factor.
> Nvidia Makes Nearly 1,000% Profit on H100 GPUs: Report
https://www.tomshardware.com/news/nvidia-makes-1000-profit-o...
It seems unfair to criticize AMD or Intel for placing losing bets on other market segments when things could have just as easily turned out differently.
Five years ago when AMD was recovering from near bankruptcy, it was battling Intel (100K+engineers) in CPUs and Nvidia (20K+ engineers) in GPUs
It was doing that with just 10K engineers and a fraction of the R&D budget of the giants.
Things have changed now. AMD has more than doubled their headcount and R&D budgets. Within the last year the company has pivoted to AI as their main focus.
OpenCL sucked for other reasons too, but it likely wouldn't have succeeded even if it didn't suck.
Other than locking out the technologies they see threatening, they’re completely a fair company, if you want to believe in that.
Also know as cost of the chip.
As a shorthand for expected retail cost, that is terribly misapplied. Or even for internal cost.
The chip would be one line item on the BOM for the entire thing you plan on shipping (but not packaged yet). Even in the case you are "just" selling a chip, the BOM is likely more complicated,and in this context primarily the manufacturing side cares about that. The COGS (cost of goods sold) is something the company as a whole will care more about - this is what it actually costs you to get it out the door. You will hear "BOM cost" referring to the elemental cost of one item on the BOM, but that's not the BOM itself.
None of these are related to the retail (or wholesale) cost in a simple way, either than forming a floors on long term sustainable price.
The GGG-whatever post is using this sloppily to suggest that the chips are going to be very expensive to produce, therefore the product is going to be expensive.
You'll also see BOM in a materials and labor type invoice, (like when you get your car serviced) but that's not relevant here.
What I really meant was total production cost + reasonable amortization cost for the development/tape out, which is of course is huge even if this chip was mass produced.
The point kinda stands though, this thing is collectively way too hard to fab + assemble to sell to consumers at any reasonable cost, especially on top of the massive R&D AMD, TSMC and everyone along the huge chain put into it.
Its not like a 4090 or Apple Silicon where the production is reasonably cheap, but margins are super high because they can be super high.
There are some existing APU systems - I use low wattage ones as thin clients. Currently thinking it should be possible to write code that runs slowly on a cheap APU and dramatically faster on a MI300A system. Debug locally, batch compute on AWS (or wherever ends up hosting these things).
(Or maybe that's TSMC's secret)
https://3dfabric.tsmc.com/english/dedicatedFoundry/technolog...
Older explanation, but lots of this stuff is just now shipping: https://www.anandtech.com/show/16051/3dfabric-the-home-for-t...
Not that AMD doesn't deserve any credit, they have considerable multi chip experience under their belt and undoubtedly served as a guinea pig/pipe cleaner for TSMC's advanced package.
The contraints all depend on the scale of the alignment required.
I wonder what is the wavelength range used for their interferometers, and what kind of mecanical engines they use (probably piezo electric based engines).
You can check if you correctly patterned the wafer almost immediately. You won't know if the layer is any good until many process steps later. Maybe not for sure until EDS. Tuning the interactions between manufacturing processes is the actual secret sauce that all the manufacturers are trying to protect. How much dose on the EUV machine depends a lot on how you intend to etch the wafer. Imagine iteration cycles measured in months for changing individual floating point variables.
The chiplets need to stay aligned as temperatures change. Much more difficult.
And there wasn't a parts shortage (modulo some cryptocurrency mining, but that impacted both GPU vendors)
And ML models weren't so large as to make 8GB of vram sound meagre.
And there weren't a bunch of venture capitalists throwing money at the work, because the state of the art models were doing uninspiring things. Like trying to tag your holiday photos, but doing it wrong because they couldn't tell a bicycle helmet and a bicycle apart.
Meanwhile, priests like Intel with Itanium, Microsoft with WinRT, FOSS nerds, or AMD with GPUs will continue failing because most people ain't got no time to be preached to about how they achieve something.
From a consumer perspective, I agree. From a datacenter, edge and industrial application perspective though, I think those crowds are content funding an effective monopoly. Hell, even after CUDA gets dethroned for AI, it wouldn't surprise me if the demand continued for supporting older CUDA codebases. AI is just one facet of HPC application.
We'll see where things go in the long-run, but unless someone resurrects OpenCL it feels unlikely that we'll be digging CUDA's grave anytime soon. In the world where GPGPU libraries are splintered and proprietary, the largest stack is king.
The SW story has been bad for a long time, but it is perhaps right now better than you think
Or do you mean the profiler tooling?
I hear everyone say that AMD doesn't have the software, but I'm a little confused --- have you tried HIP? And have you tried the automatic CUDA - HIP translation? What's missing?
CUDA runs on pretty much every NVIDIA GPU, this year they dropped support for GPUs released 9 years ago, and older binaries are very likely to be forward compatible.
Meanwhile my Radeon VII is already unsupported despite still being pretty capable (especially for FP64), and my 5700XT was never supported at all (I may be mixing this up with support for their math libraries), everyone was just led on with promises of upcoming support for 2 years. So "AMD has the software now" is not really convincing.
But if we're talking datacenter GPUs, the software is there. Data centers is where most GPGPU computing happens after all.
It's not ideal when it comes to hobby development, but if you're working in a professional capacity I'm assuming you're working with a modern HPC or AI cluster.
Of course the end goal was to run on a large HPC cluster we had access to, but for efficient development, support on personal machines was necessary. My personal dual 3090 setup has been invaluable for getting through debugging and testing before dealing with the queueing system on the cluster (Side note: it also ended up revealing another important benefit of consumer side support for GPGPU, a 3090 was easily matching the performance of a single node of the CPU-only version of the cluster, thus massively bringing down the cost of entry to an otherwise computationally restrictive topic).
Why don’t they position the SOC to have the HBM in the center with the CCDs and XCDs on the perimeter?
Seems to me that would yield lower “wire length” through the interconnect for each CCD/XCD to all of the memory.
The core systems of interconnect are at the base most center-most (well, there's a passive interposer too, but it's just wires): the IOD. These intermediate connections across the chips on top of them, the IOD next to them, and the memory.
Also, the physical interconnect between the XCDs is different than the HBM.
The die-to-die links, on the other hand, have extremely short wire length limits (less than a centimeter iirc for UCIe). At the physical layer, you just waste a bunch of power and area driving a high voltage to make the signal go further. Do that enough and you basically wind up with a PCIe PHY.
This is just for AI? So not for gaming PCs?
That is exactly what did happen. The year was 2006 and it was called the AMD Fusion project. AMD launched the chips, that they call "APUs", in 2011.
Nowadays, this configuration is common in both AMD and Intel "CPUs".
After you transfer 17 terabytes of data, it is worn out and you can't use it any more
The literal interpretation of the article is funnier though.
I guess those chips are for the movies/video industry, online or not. Because for the "consumer", the CPU is already very efficient, and I don't think we would save interesting about of battery usage in a "real usage" perspective. I may be very wrong, but I don't watch hours and hours in a row of ultra high quality videos on a small screen, that off the AC plug, the battery is unusable in a matter of a few years anyway... it not less.
Compare to switching from x86 to ARM.
Currently they seem to have a particular focus on AI frameworks and tools like PyTorch/Tensorflow/ONNX. They have sponsored and helped with a lot of PyTorch development for example, so PyTorch support for AMD is much better than it was this time last year².
(this is literally what the El Capitan supercomputer mentioned in the article is for)
ROCm, which is AMD's equivalent of CUDA. The thing is you don't have to directly interface with CUDA or ROCm. Once the framework you want to use supports these, you're done.
AMD is consistently getting used on the TOP500 machines, and this gives them insane amounts of money to improve ROCm. CUDA has an ecosystem and hardware moat, but not because they're vastly superior, but because AMD both prioritized processors first, and NVIDIA played dirty (OpenCL support and performance shenanigans, anyone?).
This moat is bound to be eaten away by both Intel and AMD, and compute will be commoditized. NVIDIA foresaw this and bought Mellanox and wanted ARM to be a complete black box, but it didn't work out.
Ethernet consortium woke up and got agitated by the fact that only ultra-low latency fabric provider is not independent anymore, so they're started to build alternatives to Infiniband ecosystem.
Interesting times are ahead.
Also there's OneAPI, but I'm not very knowledgeable about it. It's a superAPI which can target many compute platforms, like OpenCL, but takes a different approach. It compiles to platform native artifacts (CUDA/ROCm/Intel/x86/custom, etc.) IIRC.
Cool, that sounds interesting. Anything you can point at? :)
When you look to latest list [0], #1 system, Frontier, is a Nukeputer, but it's not only a Nukeputer. #5, LUMI is definitely not a Nukeputer. They are very close to us, we work together under a project with them (we're equals. They operate under a different consortium, we have stakes on another computer which is also very famous and also in top 10 in TOP500, and is not a Nukeputer). We also have our smaller systems in our own datacenter.
We operate mostly in the long tail of science, and this means we see heavy use of popular software packages, and many special software packages optimized for these long tail problems. ROCm was invisible in this niche before, but with Frontier and LUMI, and with AMD's announcements, this started to change very quickly.
ROCm libraries are more open w.r.t. their CUDA counterparts, and they started to land to mainstream distribution repositories directly. This is important for our niche. AMD is sponsoring integration of their libraries to popular packages and it started to payoff already.
Also, small-guy-programmers start to optimize LLM training routines for AMD cards, getting 99% of the performance of NVIDIA counterparts with way less power consumption.
As a result, AMD is already much more visible and more capable position when compared to last year.
https://owehrens.com/whisper-nvidia-rtx-4090-vs-m1pro-with-m...
Summary: they cherrypicked legacy nvidia sdk's and used llama batch sizes that are not used often in production...
https://twitter.com/karlfreund/status/1735078641631998271
https://developer.nvidia.com/blog/achieving-top-inference-pe...
What AMD did was a true comparison, while nvidia is applying their transformer engine which modifies & optimizes some of the computation to FP8 & they claim no measurable change in output. So yes, nvidia has some software tricks left up on their sleeve and that makes comparisons hard, but the fact remains that their best hardware can't match the mi300x in raw power. Given some time, AMD can apply the same software optimizations, or one of their partners will.
I think AMD will likely hold the hardware advantage for a while, nVidia doesn't have any product that uses chiplets while AMD has been developing this technology for years. If the trend continues to have these huge AI chips, AMD has a better hand to economically scale their AI chips.
The entire industry is motivated to break the nvidia monopoly. The cloud providers, various startups & established players like intel are building their own AI solutions. Simultaneously, CUDA is rarely used directly, typically a higher level (Python) API that can target any low-level API like cuda, PTX or rocm.
What AMD is lacking right now is decent support for rocm on their customer cards on all platforms. Right now if you don't have one of these MI cards or a rx7900 & you're not running linux you're not going to have a nice time. I believe the reason for this is that they have 2 different architectures, CDNA (the MI cards) and RDNA (the customer hardware).
Are you saying that having rx7900 + linux = happy path for ML? This is news to me, can you tell more?
I would love to escape cuda & high prices for nvidia gpus.
Except they have been given time, lots of it, and yet AMD is not anywhere close to parity with CUDA. It's almost like you can't just snap your figures and willy-nilly replicate the billions of dollars and decades investment that went into CUDA.
To get a picture of the current state which has changed a lot, this MS Ignite presentation from three weeks ago may be of interest. The slides show the drop in compatibility they have for higher levels of the stack and the tools for translation at the lower levels. Finally there's a live demo at the end.
Rename it to batch_size/sec if you don't see the issue.