Hot Chips 34: AMD’s Instinct MI200 Architecture
chipsandcheese.com
chipsandcheese.com
From what I’ve seen of Biren’s new chip it’s nice but won’t be competitive with any of Nvidia’s chips. Maybe in a few generations it’ll be more compelling, and they can build out their software stack. AMD seems to be aiming directly for HPC. Maybe Intel and their OneAPI strategy will get there? I want to root for Intel, but it’s hard to have a ton of confidence after watching their performance through the late 2010s.
The future users of your highest-end HPC accelerators start with affordable hardware they can develop and test against on their desktop. No sane developer jumps headfirst into your most expensive product (with the highest power/cooling infrastructure requirements) just to test compatibility and get familiar with your APIs.
Nvidia understands this, and you are able to run the same basic algorithms across everything from mid-range desktop GPUs to high-end HPC accelerators (performance varies, of course). Intel somewhat understands this, although has had major missteps in this area (e.g. limited AVX-512 support on desktop processors, poor developer story for Xeon Phi).
You can tell HSA to pretend your GPU is Navi 21 by setting an environment variable:
export HSA_OVERRIDE_GFX_VERSION=10.3.0
This is not a configuration that has gone through any QA testing, so I couldn't in good conscience recommend buying a GPU to use in that way. However, if you already have a 6000 series desktop GPU and you always wanted to play around with PyTorch... maybe set that variable and give it a try.But you see the catch right? People buy hardware to have support from the manufacturer. The no QA part is very very bad. :/
Nobody wants to be the one troubleshooting issues all the time, and that can alone make an NVIDIA GPU worthwhile to buy over an AMD one.
Hopefully this gets fixed in the future.
and maybe some very big past mistakes too. See the G4ad instance on AWS. That runs on the Navi12 ASIC, which never got (proper) ROCm support. Wouldn't it be awesome if an AWS instance was available widely for people to test their software with ROCm? The hardware is already there...
HIP/ROCm support should absolutely be better supported on all AMD hardware for more adoption, instead it seems to barely register like OpenACC or Vulkan compute. Intel might have better luck with OpenAPI.
For sm_70 onwards (which is the arch that comes with tensor cores) NVIDIA made the task significantly harder.
Those newer architectures use a separate instruction pointer per thread/lane, for notably supporting C++ atomics across threads in the same warp without deadlocks.
This doesn't match the semantics present on AMD GPUs.
For HIP/ROCm:
I think that they need an abstraction layer that can make a single slice of binary code that is usable across multiple gens.
Compounded by the fact that different dies have different binary slices on the AMD side, so that 6800 XT and 6700 XT run different code slices. ROCm only supports Navi21 cards for RDNA2, not the other ones...
For oneAPI, OpenCL SPIR-V fulfills that role.
https://www.amd.com/en/products/professional-graphics/amd-ra...
The goal is a card that is primarily capable of gaming (VR), but can also pull double duty for training and running moderately large networks. I think AMD systematically underestimate the importance of that niche.
For training, I’m certainly interested but not at all convinced that it will dominate. I work on multiple projects right now where the reduced dynamic range of analogue signals would be a complete non-starter given the problem domain. I’m not sure how they get around that.
We already have mixed signal accelerators (e.g. Mythic) which are not much faster than digital competition (for many reasons).
From this point of view there’s no difference between inference and training, especially as FP8 format is being adopted by Nvidia and others.
If they were to release a design with the same architecture, but with all compute larger than f16 completely gone, they could probably get identical performance from a chip that is significantly smaller. I think both AMD and Nvidia will be releasing this kind of design (with more tweaks I'd guess) in the future.
And I definitely see some future for FP16/BF16 or INT8 devices for inference, but I don’t know how widespread such devices will become for training. For many types of models, they simply won’t converge if they don’t at least use a mixed precision scheme. And there are certain problems where even for inference you get a significant benefit from using FP32 - for instance, I work on building DL surrogate models for physical simulations. We get better results with the increased dynamic range and precision.
On the other hand, work being done in the big three DL packages for making distributed training easier has been quite nice. I know the least about TF, but their dtensor looks promising. JAX’s entire distributed paradigm using pmap/xmap etc make certain classes of models very easy to distribute. The one I’m following most closely though is Pytorch and their sharded tensor, and it looks like they’re planning on implementing native distributed ops powered by their RPC framework, which should make full tensor- and model-parallelism significantly easier.
Question; Does the innovation need to happen on the hardware side?
I ask because there is a concerted effort, at least for inference, to reduce memory requirements and provide more access to giant LLMs.[0] It seems like everyone for the past few years has just thrown more and more compute at the wall to see what sticks. Where does it stop? 3 or 4 trillion parameter models that cost a $Billion to train? Is there a law of diminishing returns? Have we reached the apogee or will it be the 15 trillion parameter that ate $100Billion?
Accessibility to all manner of models, for knuckle heads like myself, is relatively new. Maybe I should rephrase, accessibility to interesting and FUN models that mere mortals can use and hack on is relatively new. Stable Diffusion is nothing short of spectacular for just having some fun with ML/DL.
My naïve thought is that, similar to most natural systems, smaller and more specialized models that can be "glued" together as a way forward. In my mind this makes more sense than these ridiculously large language models. That is to say, purpose built or fine--tuned models that can communicate seems like a good idea. GPT-NeoX "communicates" with to Bloom and makes a gaggle of small models that can be fine-tuned further. If regular people with a consumer GPU can train/fine-tune a model at home, this happens a lot faster.
So there is clearly still ways to go.
Since you brought up the brain, I will roll with some thoughts. Instead of thinking about a hunk of wet meat that contains a quadrillion synapses... wouldn't it make sense to break those into discrete, specialized and smaller hunks of wet meat of say 10 billion? Then we link those discrete units under a unified command which directs information to the correct specialized core and then the next until some arbitrary stopping point was reached. AKA: An ensemble of models.
My point; It seems that throwing more compute at individual problems has probably reached a kind of apogee in time and cost. PaLM, Bloom, GPT-NeoX and GPT-3 all live in silos. Cool story, but you can't use GPT-3 and PaLM at the same time without an absolute ton of overhead right now. At some point, I think there will be a unifying middleware that can utilize these models in there current form. That is, you will not need to train yet another model to combine the networks of one or more discrete models. Just a hunch.... could be wrong.