Radeon Open Compute 3.0
github.com
github.com
While ROCm works well, I think it is yet another example of what AMD gets wrong software-wise. I like to think I'm fairly familiar with GPGPU, in that I have used OpenCL and CUDA for years, and even wrote my own optimising compiler that generates GPU code. I even use ROCm every day on my home system. Yet, I would be hard pressed to succinctly state what ROCm is. The GitHub page states the following:
> ROCm is designed to be a universal platform for gpu-accelerated computing. This modular design allows hardware vendors to build drivers that support the ROCm framework. ROCm is also designed to integrate multiple programming languages and makes it easy to add support for other languages.
Okay? It then lists some other stuff:
> • Drivers > • Tools > • Libraries > • Source Code
The AMD GPU kernel driver, amdgpu, is upstreamed in Linux, so what exactly does ROCm contain? The OpenCL ICD? (I'm pretty sure it does.) As far as I can see, I don't have any ROCm-specific kernel drivers loaded, but OpenCL on my AMD GPU works fine.
The tools are also confusing. There's an "HCC compiler", but are we even still supposed to use HCC? As I recall, it was/is some heterogeneous compute C++ dialect. There's also HIP, which converts CUDA to something that can run on AMD GPUs. Trying to refresh my memory, I actually can't figure out in two minutes what language the output is. HCC? OpenCL?
Among the other tools, there is inconsistency between ROCM and ROCm (pedantry, I know, I assume it's the same), but there is also the "ROC Profiler" and "ROCr Debug Agent". Are "ROC" and "ROCr" significant terms? And when should I use the "Radeon Compute Profiler" (rocm-profiler) and when should I use the "ROC Profiler" (rocprofiler-dev)?
I do so dearly want to support AMD. Their hardware is good and their fully open source drivers are a wonderful accomplishment. Seriously, I want to underline just how amazed I am that I have a fully free high-performance GPGPU stack running. However, if you compare AMDs software efforts to NVIDIA, it is clear that NVIDIA deserves their success. It's not just about breadth (NVIDIA is bigger, I can understand they have more resources), but the fact that AMD seems to spread itself too thin over too many effort, and it makes everything feel kind of shoddy and poorly documented.
Kinda making your point, HIP is actually an (open source) one-to-one replacement for the CUDA API. While there are tools (https://github.com/ROCm-Developer-Tools/HIP/tree/master/hipi... and https://github.com/ROCm-Developer-Tools/HIP/blob/master/bin/...) that convert CUDA to HIP, these are meant to be run once and perhaps touched up manually. HIP source can then be compiled (no human interaction) for AMD (ROCm) devices and NVIDIA (CUDA) devices.
The CUDA brand refers to both the API and the device. AMD certainly could have done a better job, but wanted to distinguish these, presumably in hopes that existing users of the CUDA API would consider porting their code to HIP even when targeting CUDA devices.
[0]: https://gpuopen.com/wp-content/uploads/2016/01/7637_HIP_Data...
I couldn't figure out what ROCM is, what it does, what parts of it I was meant to install (if anything?) and what was a brand name for coordinated upstream efforts in the linux kernel. Does it have an API? Who is meant to be using it for what? I dunno.
I honestly don't understand. I don't know anything about CUDA because, as mentioned, no Nvidia GPU in a decade. My theory is if I pour all my effort into learning about Vulkan that will be good enough. I'll run tensorflow on my CPU or just do without and suffer for it. Whatever ROCMs use case and target audience is I hope it isn't me.
If I wanted to be serious about getting in to AI research I'd probably need to buckle and buy Nvidia.
Eg, there is an "Important features include the following" section where a typical example is "User-mode queues and DMA". We've had user mode queues since Sedgewick's Algorithms in C. Or "Large memory allocations" which we've had since the introduction of the 64 bit CPU. Obviously these must be features in the context of a GPU, but they don't really have a lot of explanatory power for anyone not already intimately familiar with the field.
I suppose the obvious conclusion is ROCM is professionals only; but presumably those professionals are already perfectly happy with CUDA. I'd find it helpful to know what these professionals are expected to implement with all this so I can go figure that out.
Specifically for me I know Vulkan (aka OpenGL++) is a general purpose API for running things on a GPU. I don't understand if that supersedes ROCM, compliments ROCM or what. I'm assuming that Vulkan is all anyone needs and then ROCM is going to ... provide something, maybe libraries? ... as middleware to Tensorflow and related tools. But blow me down if I can find anything that fills me with confidence that I've got the right take. ROCM seems to be very low level which makes little sense to me given that Vulkan (and the slowly obsoleting OpenCL) seem to be where the compute work is being done.
If I wanted to summarise PyTorch it is "a Python API & library that provides convenience functions and data models for machine learning." Or CUDA, at a guess, is "a C++ API for highly parallel compute on Nvidia GPU". ROCm claims to be "the first open-source HPC/Hyperscale-class platform for GPU computing that’s also programming-language independent" and none of that means much to me. Doesn't hint what programming language I'm expecting to be using, and I don't think I even want Hyperscale computing. But I get the feeling I want ROCm because it is associated with PyTorch on AMD GPUs. Befuddling.
I dunno, the ROCm version of PyTorch looked complicated and convoluted so I'm just going to sit it out and wait for someone to explain who the circus is performing for.
While Khronos was busy pushing the C only API they go from Apple, Nvidia created a GPGPU platform, with support for C, C++, Fortran and any language with PTX capable backend.
When Khronos woke up to the fact that most researchers don't want to use plain old C, it was too late.
It remains to be seen how much uptake SYSCL or SPIR-V will ever get, across GPGPU vendors.
Intel is trying to get back into the game via their extended SYSCL, so lets see.
Also note that on mobile devices only Apple actually has supported OpenCL, while Google pushed their Renderscript instead.
> Going to 11: Amping Up the Programming-Language Run-Time Foundation
I'm pretty sure it needs amdkfd, which is not really "specific" but doesn't have any other consumers.
> but OpenCL on my AMD GPU works fine
Well, if instead of ROCm (or proprietary "pro" stuff) you have Mesa's Clover… that's the opposite of fine. I mean, fine for simple things, but Darktable won't work (no image support), Blender won't work (just too complex lol)…
---
my casual user opinion: I HATE ALL the special compute-specific GPU APIs. I don't want a special fancy kernel driver that loads everything differently, I don't like OpenCL and other unusual special garbage. Please please please just use Vulkan compute shaders for everything.
(And thankfully, people are starting to do that https://github.com/nihui/waifu2x-ncnn-vulkan https://github.com/hanatos/vkdt …)
https://towardsdatascience.com/on-the-state-of-deep-learning...
These things take time. CUDA took what, a decade to become what it is today? And it was developed by a company that wasn't on its deathbed.
Arguably AMD still doesn't have much weight to throw around. Their profits are meager. Their P/E is absolutely insane for a hardware company (240!), suggesting a dramatic correction in the foreseeable future. Stocks on an upward trajectory help attract top talent. On a downward trajectory they actively repel it.
So they probably have like half a dozen engineers working on this, and not those magical $500K/yr engineers that can actually turn shit into gold, because a hardware company wouldn't be able to hire them (or afford them, for that matter).
Yes, in a market today that is trading relatively low in PE~20 - 25ish.
AMD would have to grow 12 times its profits to normalise that price, as much as I am pro AMD in the Intel/AMD case, 12 times profit is a little over the top.
I'm sorry, but this makes absolutely no sense at all. P/E or corrections have absolutely nothing to do with it. Salaries are paid from revenue, not from profits, and their revenues have been growing for years. They don't need to pay engineers in stock options.
This is about priorities and execution. Getting 10 engineers to focus on an obvious multi-billion dollar opportunity that leverages existing investments is not "throwing your weight around". It's basic management competence.
And it's interesting work. I don't believe that engineers need low P/Es as an incentive to work on a key part of AI infrastructure that only very few people get to work on.
In the time it took the other company to contact me, do 6 interviews, negotiate contract, and get me signed, I haven’t even been contacted by AMD. Their position is still open, I’m a perfect fit and I haven’t been rejected either. I suppose I’ll get an automatic rejection letter at some point.
The online application system for AMD is just quite bad. Looking at my application in their system, everything is full with “Unknown” and I can’t event put in my skills. It wouldn’t surprise me if on their database I’m a horrible candidate.
It's also a hard slog, and something very few people _really_ know how to do. Ideally you have to be both deep learning / HPC expert _and_ low level programming expert, and on an "exotic" platform to boot, not on the CPU. Such people are scarcer than hen's teeth, and they have no problem whatsoever getting mid-six-figure employment at companies that don't treat them as an afterthought, and where they can hope for a strong upward trajectory career-wise.
I agree with the rest of your point, and in particular that this is an existential threat to AMD as a whole. I don't get why they don't treat it as one.
Be that as it may, I think you're severely overestimating the present level of "management competence" allocated to something that's not currently perceived as a do-or-die situation at a company that was dying just a year ago.
Just for completeness: Tensorflow has (upstream) ROCm support. When training some of our transformer models, the performance difference was not large compared to AMD's Tensorflow fork.
I am curious on what's the motivation given that nvcc is fairly competent.
Not just NixOS. Our work AMD GPU server uses Ubuntu 18.04, but I use Anthony's package set for easy use of PyTorch and Tensorflow in various environments (using nix-shell + direnv).
I spent some time wondering why I couldn't get access to the device until some kind soul pointed me to the different group name.
I don't know if there's any relationship between that and the free software infrastructure, but it's a relief to see an alternative to the CUDA-ish stuff. In the absence of a shared library dummy interface, we can't distribute OS packages of performance engineering tools with NVIDIA support, for instance.
Also I don't think they've got their driver usage right. Metal on the GPU gets different results than OpenCL. OpenCL is supposed to be the 'beta' and the results are more stable with it - and again different radically between the PlaidML devices.
We’re dealing with single precision floats for everything. FWIW, every implementation of IEE 754 floats has its own set of minor oddities and quirks.
However, I suspect Apple has no intention of doing this, as they are more focus on pushing Metal and reliance on macOS.
I wouldn't mind seeing PCI or GPU passthrough capabilities this in Windows 10 Pro (not just Server). With their Linux Subsystem progressing the way it is, they really have an opportunity here. Though since NVidia and CUDA are supported in Windows, maybe not as necessary as the situation on macOS.
Granted that I haven't really tried properly, as the card does show up as a compatible device, but trying to run Tensorflow on it fails due to associated GPU code not being available. Perhaps it doesn't have official support yet due to needing more compiler work.
It’ll be interesting to see how these cards and ROCm perform in comparison to NVIDIA’s offering and CUDA.
https://github.com/RadeonOpenCompute/ROCm/tree/roc-3.0.0#roc...
It's a collection of GPU development tools and libraries.