Run CUDA, unmodified, on AMD GPUs
docs.scale-lang.com
docs.scale-lang.com
Chasing bug-for-bug compatibility is a fool's errand. The important users of CUDA are open source. AMD can implement support directly in the upstream projects like pytorch or llama.cpp. And once support is there it can be maintained by the community.
I disagree. AMD can simply not implement those APIs, similar to how game emulators implement the most used APIs first and sometimes never bother implementing obscure ones. It would only matter that NVIDIA added eg. patented APIs to CUDA if those APIs were useful. In which case AMD should have a way to do them anyway. Unless NVIDIA comes up with a new patented API which is both useful and impossible to implement in any other way, which would be bad for AMD in any event. On the other hand, if AMD start supporting CUDA and people start using AMD cards, then developers will be hesitant to use APIs that only work on NVIDIA cards. Right now they are losing billions of dollars on this. Then again they barely seem capable of supporting RocM on their cards, much less CUDA.
You have a fair point in terms of cuDNN and cuBLAS but I don't know that that kind of ToS is actually binding.
h.265 is one way, av1 is another way.
Support this, reimplement that, support upstream efforts, dont really care. Any of those would cost a couple of million and be worth a trillion dollars to AMD shareholders.
AMD only claims support for a select few GPUs, but in my testing I find all the GPUs work fine if the architecture is supported. I've tested rx6600, rx6700xt for example and even though they aren't officially supported, they work fine on ROCm.
AMD had a big architecture switchover exactly 5 years ago, and the full launch wasn't over until 4.5 years ago. I think that generation should have full support. Especially because it's not like they're cutting support now. They didn't support it at launch, and they didn't support it after 1, 2, 3, 4 years either.
The other way to look at things, I'd say that for a mid to high tier GPU to be obsolete based on performance, the replacement model needs to be over twice as fast. 7700XT is just over 50% faster than 5700XT.
It's not ideal but I'm pretty sure CUDA didn't support everything from day 1. And ROCm is part of AMD's vendor part of the Windows AI stack so from upcoming gen on out basically anything that outputs video should support ROCm.
[1]https://www.gamesindustry.biz/nvidia-unveils-cuda-the-gpu-co...
[2]https://www.tomshardware.com/news/amd-rocm-comes-to-windows-...
https://rocm.docs.amd.com/projects/install-on-linux/en/lates...
https://en.wikipedia.org/wiki/ROCm#:~:text=GCN%205%20%2D%20V...
Stable Diffusion ran fine for me on RX 570 and RX 6600XT with nothing but distro packages.
I agree that the official support list is very conservative, but I wouldn't recommend pre-Vega GPUs for use with ROCm. Stick to gfx900 and newer, if you can.
Who’s the root cause? The company with the dominant platform that refuses to open it up, or the competitor who can’t catch up because they’re running so far behind? Even if AMD made their own version of CUDA that was better in every way, it still wouldn’t gain adoption because CUDA has become the standard. No matter what they do, they’ll need to have a compatibility layer. And in that case maybe it makes sense for them to invest in the best one that emerges from the community.
Nvidia has put in the legwork and are reaping the rewards. They've worked closely with the people who are actually using their stuff, funding development and giving loads of support to researchers, teachers and so on, for probably a decade now. Why should they give all that away?
> But there are counterexamples that suggest otherwise (Android).
How is Android a counterexample? Google makes no money off of it, nor does anyone else. Google keeps Android open so that Apple can't move everyone onto their ad platform, so it's worth it for them as a strategic move, but Nvidia has no such motive.
> Even if AMD made their own version of CUDA that was better in every way, it still wouldn’t gain adoption because CUDA has become the standard.
Maybe. But again, that's because NVidia has been putting in the work to make something better for a decade or more. The best time for AMD to start actually trying was 10 years ago; the second-best time is today.
Google makes no money off of Android? That seems like a really weird claim to make. Do you really think Google would be anywhere near as valuable of a company if iOS had all of the market share that the data vacuum that is Android has? I can't imagine that being the case.
Google makes a boatload off of Android, just like AMD would if they supported open GPGPU efforts aggressively.
nvidia could give away the software platform - CUDA - to hardware vendors for free, making the hardware into cheap, low-margin commodity items. But how would they make boatloads of money when there's nowhere to put ads, tracking or an app store?
On the other hand, they could always wait for the most viable threat to emerge and then pay a few billion dollars to acquire it and own its direction. Google didn’t invent Android, after all…
> Google is a multi-sided platform that does lots of things for free for some people… That isn't the chip business whatsoever.
This is a reductionist differentiation that overlooks the similarities between the platforms of “mobile” and “GPU” (and also mischaracterizes the business model of Google, who does in fact make money directly from Android sales, and even moved all the way down the stack to selling hardware). In fact there is even a potentially direct analogy between the two platforms: LLM is the top of the stack with GPU on the bottom, just like Advertising is the top of the stack with Mobile on the bottom.
Yes, Google’s top level money printer is advertising, and everything they do (including Android) is about controlling the maximum number of layers below that money printer. But that doesn’t mean there is no benefit to Nvidia doing the same. They might approach it differently, since they currently own the bottom layer whereas Google started from the top layer. But the end result of controlling the whole stack will lead to the same benefits.
And you even admit in your comment that Nvidia is investing in these higher levels. My argument is that they are jeopardizing the longevity of these high-level investments due to their reluctance to invest in an open platform at the bottom layer (not even the bottom, but one level above their hardware). This will leave them vulnerable to encroachment by a player that comes from a higher level, like OpenAI for example, who gets to define the open platform before Nvidia ever has a chance to own it.
30 years ago people were making the same argument that MS should have kept DirectX open or else they were going to lose to OpenGL. Look how that's worked out for them.
> Google, who does in fact make money directly from Android sales
They don't though. They have some amount of revenue from it, but it's a loss-making operation.
> In fact there is even a potentially direct analogy between the two platforms: LLM is the top of the stack with GPU on the bottom, just like Advertising is the top of the stack with Mobile on the bottom.
But which layer is the differentiator, and which layer is just commodity? Google gives away Android because it isn't better than iOS and isn't trying to be; "good enough" is fine for their business (if anything, being open is a way to stay relevant where they would otherwise fall behind). They don't give away the ad-tech, nor would they open up e.g. Maps data where they have a competitive advantage.
NVidia has no reason to open up CUDA; they have nothing to gain and a lot to lose by doing so. They make a lot of their money from hardware sales which they would open up to cannibalisation, and CUDA is already the industry standard that everyone builds on and stays compatible with. If there was ever a real competitive threat then that might change, but AMD has a long way to go to get there.
Not even a little bit. It simply isn't Nvidia's job to provide competitive alternatives to Nvidia. Competing is something AMD must take responsibility for.
The only reason CUDA is such a big talking point is because AMD tripped over their own feet supporting accelerated BLAS on AMD GPUs. Realistically it probably is hard to implement (AMD have a lot of competent people on staff) but Nvidia hasn't done anything unfair apart from execute so well that they make all the alternatives look bad.
Also, was there any followup to this story? It seems a bit unnecessary because nVidia has already neutered consumer cards for many/most data center purposes by not using ECC and by providing so few FP64 units that double precision FLOPS is barely better than CPU SIMD.
of course people continued to melt down about that for some reason too, in the customary “nothing is ever libre enough!” circular firing squad. Just like streamline etc.
There’s a really shitty strain of fanboy thought that wants libre software to be actively worsened (even stonewalled by the kernel team if necessary) so that they can continue to argue against nvidia as a bad actor that doesn’t play nicely with open source. You saw it with all these things but especially with the open kernel driver, people were really happy it didn’t get upstreamed. Shitty behavior all around.
You see it every time someone quotes Linus Torvalds on the issue. Some slight from 2006 is more important than users having good, open drivers upstreamed. Some petty brand preferences are legitimately far important than working with and bringing that vendor into the fold long-term, for a large number of people. Most of whom don’t even consider themselves fanboys! They just say all the things a fanboy would say, and act all the ways a fanboy would act…
Also: look into why the Nouveau driver performance is limited.
and you don't have the final say on what NVIDIA is allowed to do with their software either
1) they choose (chose?) not to supprt standard display protocols that Wayland compositors target with their drivers (annoying, but not the end of the world)
2) they cryptographically lock users out of writing their own drivers for their own graphics cards (which should be illegal and is exactly contradictory to "that's not actually a thing").
Again: look into why the Nouveau driver performance is limited.
It's not. Even as it is, I do not trust HIP or RocM to be a viable alternative to Cuda. George Hotz did plenty of work trying to port various ML architectures to AMD and was met with countless driver bugs. The problem isn't nvidia won't build an open platform - the problem is AMD won't invest in a competitive platform. 99% of ML engineers do not write CUDA. For the vast majority of workloads, there are probably 20 engineers at Meta who write the Cuda backend for Pytorch that every other engineer uses. Meta could hire another 20 engineers to support whatever AMD has (they did, and it's not as robust as CUDA).
Even if CUDA was open - do you expect nvidia to also write drivers for AMD? I don't believe 3rd parties will get anywhere writing "compatibility layers" because AMD's own GPU aren't optimized or tested for CUDA-like workloads.
Instead they managed 15 years of disappointment, in a standard stuck in C99, that adopted C++ and a polyglot bytecode too late to matter, never produced an ecosystem of IDE tooling and GPU libraries.
Naturally CUDA became the standard, when NVIDIA provided what the GPU community cared about.
Because it IS AMD/Apple/etcs fault for the position they're in right now. CUDA showed where the world was heading and where the gains in compute would be made well over a decade ago now.
They even had OpenCL, didn't put the right amount of effort into it, all the talent found CUDA easier to work with so built there. Then what did AMD, Apple do? Double down and try and make something better and compete? Nah they fragmented and went their own way, AMD with what feels like a fraction of the effort even Apple put in.
From the actions of the other teams in the game it's not hard to imagine a world without CUDA being a world where this tech is running at a fraction of it's potential.
With the number of devs that use Apple silicon now-a-days, I have to think that their support for khronos initiatives like SYCL and OpenCL would have significantly accelerated progress and adoption in both.
We need an open standard that isn't just AMD specific to be successful in toppling CUDA.
Pretty sure APIs are not copyrightable, e.g. https://www.law.cornell.edu/supremecourt/text/18-956
> against the license agreement of cuDNN or cuBLAS to run them on this
They don’t run either of them, they instead implement an equivalent API on top of something else. Here’s a quote: “Open-source wrapper libraries providing the "CUDA-X" APIs by delegating to the corresponding ROCm libraries. This is how libraries such as cuBLAS and cuSOLVER are handled.”
Constitutional crises involve fundamental breaks in the working of government that bring two or more of its elements into direct conflict that can't be reconciled through the normal means. The last of these by my accounting was over desegregation, which was resolved with the President ordering the Army to force the recalcitrant states to comply. Before that was a showdown between the New Deal Congress and the Supreme Court, which the former won by credibly threatening to pack the latter (which is IMO a much less severe crisis but still more substantial than anything happening today). However, that was almost a century ago, and Congress has not been that coherent lately.
If "I think this is a very bad decision" was cause for a constitutional crisis, any state with more than three digit population would be in constitutional crisis perpetually.
This happened as recently as 2021-01-06; strong evidence that the military subverted the president to call the National Guard into Washington DC and secure the electoral count.
I'd say it was narrowly averted though.
And generally courts/judges just choose the scope of their legal opinions based on how far reaching they want the legal principles to apply.
IMHO, copyright-ability of APIs is so far away from their political agenda that they probably just decided to leave the issue on a cliffhanger...
Ever heard of Intel?
C compilers didn't offer an "AMD" CPU target* until AMD came out with the "AMD64" instruction set. Today we call this "x86_64" or "x64".
* Feel free to point out some custom multimedia vector extensions for Athlons or something, but the point remains.
The 1982 deal you speak of was actually pretty interesting: as a condition of the x86's use in the IBM PC, IBM requested a second source for x86 chips. AMD was that source, and so they cross-licensed the x86 in 1982 to allow the IBM PC project to proceed forward. This makes the Intel/AMD deal even more important for both companies: the PC market would never have developed without the cross-licensing, which would've been bad for all companies involved. This gave Intel an ongoing stake in AMD's success at least until the PC market consolidated on the x86 standard.
Don't believe me? Include this at the top of your CUDA code, build with hipcc, and see what happens:
https://gitlab.com/StanfordLegion/legion/-/blob/master/runti...
It's incomplete because I'm lazy but you can see most things are just a single #ifdef away in the implementation.
you have to be able to pip install something and just have it work, reasonably fast, without crashing, and also it has to not interfere with 100 other weird poorly maintained ML library dependencies.
Of course becoming "really good" is a lot different and like anything else, it presumably takes a lot of callused fingertips (from typing) to get there.
What's your opinion on the future of writing optimized GPU code/kernels - how long before compilers are as good or better than (most) humans writing hand-optimized PTX?
You'll also note that M1/2 macs with large amounts of system memory are good at inference because of the fact that the gpu has a very high speed interconnect between the soldiered on ram modules and the on die gpu. It's all about avoiding bottlenecks whereever possible.
So yeah it’s not ideal if you’re on a budget, but it seems like there are some solutions that don’t involve massive capex.
A new 4090 costs around $1800 (https://www.centralcomputer.com/asus-tuf-rtx4090-o24g-gaming...) and that's probably affordable to AWS users. I see a 2080Ti on Craigslist for $300 (https://sfbay.craigslist.org/scz/sop/d/aptos-nvidia-geforce-...) though used GPU's are possibly thrashed by bitcoin mining. I don't have a suitable host machine, unfortunately.
When people were mining Ethereum (which was the last craze that GPUs were capable of playing in -- BTC has been off the GPU radar for a long time), profitable mining was fairly kind to cards compared to gaming.
Folks wanted their hardware to produce as much as possible, for as little as possible, before it became outdated.
The load was constant, so heat cycles weren't really a thing.
That heat was minimized; cards were clocked (and voltages tweaked) to optimize the ratio of crypto output to Watts input. For Ethereum, this meant undervolting and underclocking the GPU -- which are kind to it.
Fan speeds were kept both moderate and tightly controlled; too fast, and it would cost more (the fans themselves cost money to run, and money to replace). Too slow, and potential output was left on the table.
For Ethereum, RAM got hit hard. But RAM doesn't necessarily care about that; DRAM in general is more or less just an array of solid-state capacitors. And people needed that RAM to work reliably -- it's NFG to spend money producing bad blocks.
Power supplies tended to be stable, because good, cheap, stable, high-current, and stupidly-efficient are qualities that go hand-in-hand thanks to HP server PSUs being cheap like chips.
There were exceptions, of course: Some people did not mine smartly.
---
But this is broadly very different from how gamers treat hardware, wherein: Heat cycles are real, over clocking everything to eek out an extra few FPS is real, pushing things a bit too far and producing glitches can be tolerated sometimes, fan speeds are whatever, and power supplies are picked based on what they look like instead of an actual price/performance comparison.
A card that was used for mining is not implicitly worse in any way than one that was used for gaming. Purchasing either thing involves non-zero risk.
> Fan speeds were kept both moderate and tightly controlled; too fast, and it would cost more (the fans themselves cost money to run, and money to replace). Too slow, and potential output was left on the table.
In the ideal case, this is spot on. Annoyingly however, this hinges on the assumption of an awful lot of competence from top to bottom.
If I've learned anything in my considerable career, it's that reality is typically one of the first things tossed when situations and goals become complex.
The few successful crypto miners maybe did some of the optimizations you mention. The odds aren't good enough for me to want to purchase a Craigslist or FB marketplace card for only a 30% discount.
I do genuinely admire your idealism, though.
A used card is it sale. It was previously used for mining, or it was previously used for gaming.
We can't tell, and caveat emptor.
Which one is worse? Neither.
Gaming cards/rigs, which many of the early miners were based on, rarely run at 100% all the time, the workload is burst-y (and distributed amongst different areas of the system). In comparison, a miner runs at 100% all the time.
On top of that, for silicon there is an effect called electromigration [1], where the literal movement of electrons erodes the material over time - made worse by ever shrinking feature sizes as well as, again, the chips being used in exactly the same way all the time.
Optimizing something for SIMD execution isn't often straightforward and it isn't something a lot of developers encounter outside a few small areas. There are also a lot of hardware architecture considerations you have to work with (memory transfer speed is a big one) to even come close to saturating the compute units.
5320814 / 180 / 16 = ~1847.5
Per https://www.apmex.com/gold-price and https://goldprice.org/, current value is north of $2400 / oz. It was around $1800 in 2020. That growth for _gold_ of all things (up 71% in the last 5 years) is crazy to me.
It's worth noting that anyone with a ski house that expensive probably has a net worth well over twice the price of that ski house. I guess it's time to start learning CUDA!
https://en.wikipedia.org/wiki/Troy_weight
180 avoirdupois pounds is 2,625 ounces troy. The gold price is around $2470/ounce troy today, so $2470*2625 ~= $6.483 million
For comparison: S&P500 grew about the same during that period (more than 100% from Jan 2019, about 70 from Dec 2019), so the higher price of gold did not outperform the growth of the general (financial) economy.
Anyway, there isn't a lot of evidence that the value of gold is going up. It seems to just be keeping pace with the M2. Both doubled-and-a-bit since 2010 (working in USD).
Nvidia literally wrote most of the textbooks in this field and you’d probably be taught using one of these anyway:
https://developer.nvidia.com/cuda-books-archive
“GPGPU Gems” is another “cookbook” sort of textbook that might be helpful starting out but you’ll want a good understanding of the SIMT model etc.
GP says it is just some #ifdefs in most cases, so an LLM should be able to do it, right?
This seems to be fairly common problem with software. The people who create software regularly deal with complex tool chains, dependency management, configuration files, and so on. As a result they think that if a solutions "exists" everything is fine. Need to edit a config file for your particular setup? No problem. The thing is, I have been programming stuff for decades and I really hate having to do that stuff and will avoid tools that make me do it. I have my own problems to solve, and don't want to deal with figuring out tools no matter how "simple" the author thinks that is to do.
A huge part of the reason commercial software exists today is probably because open source projects don't take things to this extreme. I look at some things that qualify as products and think they're really simplistic, but they take care of some minutia that regular people are will to pay so they don't have to learn or deal with it. The same can be true for developers and ML researchers or whatever.
(and for one data point, I believe Blender is actively using HIP for AMD GPU support in Cycles.)
In the case of these abstraction layers, then it would be the responsibility of the abstraction maintainers (or AMD) to port them. Obviously, someone who does not even use CUDA would not use HIP either.
To be honest, I have a hard time believing that a truly zero-effort solution exists. Especially one that gets high performance. Once you start talking about the full stack, there are too many potholes and sharp edges to believe that it will really work. So I am highly skeptical of original article. Not that I wouldn't want to be proved wrong. But what they're claiming to do is a big lift, even taking HIP as a starting point.
The easiest, fastest (for end users), highest-performance solution for ML will come when the ecosystem integrates it natively. HIP would be a way to get there faster, but it will take nonzero effort from CUDA-proficient engineers to get there.
As other commenters have pointed out, this is probably a good solution for HPC jobs where everyone is using C++ or Fortran anyway and you frequently write your own CUDA kernels.
From time to time I run into a decision maker who understandably wants to believe that AMD cards are now "ready" to be used for deep learning, and points to things like the fact that HIP mostly works pretty well. I was kind of reacting against that.
I don't think so. I agree it is too hard for the ML researches at the companies which will have their rear ends handed to them by the other companies whose ML researchers can be bothered to follow a blog post and prompt ChatGPT to resolve error messages.
and rightly so, they have more complicated issues to tackle
It's on developers to provide better infrastructure and solve these challenges
I'm also speaking from personal experience: I once had to hand-write my own CUDA kernels (on official NVIDIA cards, not even this weird translation layer): it was useful and I figured it out, but everything was constantly breaking at first.
It was a drag on productivity and more importantly, it made it too difficult for other people to run my code (which means they are less likely to cite my work).
If you want to write very efficient CUDA kernel for modern datacenter NVIDIA GPU (read H100), you need to write it with having hardware in mind (and preferably in hands, H100 and RTX 4090 behave very differently in practice). So I don't think the difference between AMD and NVIDIA is as big as everyone perceives.
It does not mean that the compiler is able to generate code that has optimal performance, when that can be achieved by using certain instructions without a direct equivalent in a high-level language.
No compiler that supports the Intel-AMD ISA knows how to use all the instructions available in this ISA.
It's been a while since I looked at CUDA, but it used to be that NVIDIA were continually extending cuDNN to add support for kernels needed by SOTA models, and I assume these kernels were all hand optimized.
I'm curious what kind of models people are writing where not only is there is no optimized cuDNN support, but also solutions like Triton or torch.compile, and even hand optimized CUDA C kernels are too slow. Are hand written PTX kernels really that common ?
I'm more curious about what types of model people in research or industry are developing, where NVIDIA support such as this is not enough, and they are developing their own PTX kernels.
It almost make me wonder. Is there a shady trade somewhere to ask amd never release sdk for Windows to hike the price of nvidia card higher? Why they keep developing these without release it at all?
(Let's put the legal questions aside for a moment.)
nVidia changes GPU architectures every generation / few generations, right? How does CUDA work across those—and how can it have forwards compatibility in the future—if it's not designed to be technologically agnostic?
There are also versions of PyTorch, TensorFlow and JAX with AMD support.
PyTorch's torch.compile can generate Triton (OpenAI's GPU compiler) kernels, with Triton also supporting AMD.
We have seen this succeed multiple times: FreeSync vs GSync, DLSS vs FSR, (not AMD but) Vulkan vs DirectX & Metal.
All of the big tech companies are obsessed with ring-fencing developers behind the thin veil of "innovation" - where really it's just good for business (I swear it should be regulated because it's really bad for consumers).
A CUDA translation layer is okay for now but it does risk CUDA becoming the standard API. Personally, I am comfortable with waiting on an open standard to take over - ROCm has serviced my needs pretty well so far.
Just wish GPU sharing with VMs was as easy as CPU sharing.
That's all to say NVIDIA could pull a SGI and open their stuff, but they're going more sony style and trying to monopolize. Oh, and SGI also wrote another ancient lore library known as "STL" or the "SGI Template Library" which is like the original boost template metaprogramming granddaddy
Zero impact on Switch, Playstation, XBox, Windows, macOS, iOS, iPadOS, Vision OS.
dxvk-gplasync is a game changer for dx9-11 shader stutter.
Which Android Studios can't even be bothered to target with their NDK engines, based on GL ES, Vulkan.
Even if there's no shader stutter, Vulkan tends to use less juice than DX.
I'll definitely agree with you on Sync and Vulkan, but dlss and xess are both better than fsr.
Karol Herbst is working on Rusticl, which is mesa's latest OpenCL implementation and will pave the way for other things such as SYCL.
Realistically companies only have an obligation to make themselves profitable so really companies "should" only strive for profitability within the boundaries of the law above all else.
AMD have no obligation to drive an open standard, it's at their discretion to choose that approach - and it might actually come at the cost of profitability as it opens them up to competitors.
In this case - I believe that hardware & platform software companies that distribute a closed platform which cannot be genuinely justified as anything other than intending to prevent consumers from using competitor products "should" be moderated by regulator intervention as it results in a slower rate of innovation and poor outcomes for consumers.
That said, dreaming for the regulation of American tech giants is a pipe dream, haha.
AMD has always been notoriously bad at the software side, and they frequently abandon their projects when they're almost usable, so I won't hold my breath.
That said, easier said than done. You need very specialized developers to build a CUDA equivalent and have people start using it. AMD could do it with a more open development process leveraging the open source community. I believe this will happen at some point anyway by AMD or someone else. The market just gets more attractive by the day and at some point the high entry barrier will not matter much.
So why should AMD skimp on their ambitions here? This would be a most sensible investment, few risks and high gains if successful.
Presumably that stuff doesn't "just work" but they don't want to mention it?
A lot of our hw-aware bits are parameterized where we fill in constants based on the available hw . Doable to port, same as we do whenever new Nvidia architectures come out.
But yeah, we have tricky bits that inline PTX, and.. that will be more annoying to redo.
SCALE does not use any part of ZLUDA. We have modified the clang frontend to convert inline PTX asm block to LLVM IR.
To put in a less compiler-engineer-ey way: for any given block of PTX, there exists a hypothetical sequence of C++/CUDA code you could have written to achieve the same effect, but on AMD (perhaps using funky __builtin_... functions if the code includes shuffles/ballots/other-weird-gpu-stuff). Our compiler effectively converts the PTX into that hypothetical C++.
Regarding memory consistency etc.: NVIDIA document the "CUDA memory consistency model" extremely thoroughly, and likewise, the consistency guarantees for PTX. It is therefore sufficient to ensure that we use operations at least as synchronising as those called for in the documented semantics of the language (be it CUDA or PTX, for each operation).
Differing consistency _between architectures_ is the AMDGPU backend's problem.
wgmma.mma_async.sync.aligned.m64n256k16.f32.bf16.bf16
Do you reverse it back into C++ that does the corresponding FMAs manually instead of using tensor hardware? Or are you able to convert it into a series of __builtin_amdgcn_mfma_CDFmt_MxNxKABFmt instructions that emulate the same behavior?But in general the answer to your question is yes: we use AMD-specific builtins where available/efficient to make things work. Otherwise many things would be unrepresentble, not just slow!
If there's no instruction, either, you can write a C++ function to replicate the behaviour and codegen a call to it. Since the PTX blocks are expanded during initial IR generation, it all inlines nicely by the end. Of course, such software emulation is potentially suboptimal (depends on the situation).
I'm curious how something like this example would translate:
===
Mapping lower-level ptx patterns to higher-level AMD constructs like __ballot, and knowing it's safe
```
#ifdef INLINEPTX
inline uint ptx_thread_vote(float rSq, float rCritSq) {
uint result = 0;
asm("{\n\t"
".reg .pred cond, out;\n\t"
"setp.ge.f32 cond, %1, %2;\n\t"
"vote.sync.all.pred out, cond, 0xffffffff;\n\t"
"selp.u32 %0, 1, 0, out;\n\t"
"}\n\t"
: "=r"(result)
: "f"(rSq), "f"(rCritSq));
return result;
}
#endif
```===
Again, I'm guessing there might be an equiv simpler program involving AMD's __ballot, but I'm unsure of the true equivalence wrt safety, and it seems like a tricky rewrite as it needs to (afaict) decompile to recover the higher-level abstraction. Normally it's easier to compile down or sideways (translate), and it's not clear to me these primitives are 1:1 for safely doing so.
===
FWIW, this is all pretty cool. We stay away from PTX -- most of our app code is higher-level, whether RAPIDS (GPU dataframes, GPU ML, etc libs), minimal cuda, and minimal opencl, with only small traces of inline ptx. So more realistically, if we had the motivation, we'd likely explore just #ifdef'ing it with something predictable.
.p2align 2 ; -- Begin function _Z15ptx_thread_voteff
.type _Z15ptx_thread_voteff,@function
_Z15ptx_thread_voteff: ; @_Z15ptx_thread_voteff
; %bb.0: ; %entry
s_waitcnt vmcnt(0) expcnt(0) lgkmcnt(0)
s_waitcnt_vscnt null, 0x0
v_cmp_ge_f32_e32 vcc_lo, v0, v1
s_cmp_eq_u32 vcc_lo, -1
s_cselect_b32 s4, -1, 0
v_cndmask_b32_e64 v0, 0, 1, s4
s_setpc_b64 s[30:31]
.Lfunc_end1:
.size _Z15ptx_thread_voteff, .Lfunc_end1-_Z15ptx_thread_voteff
; -- End function
What were the safety concerns you had? This code seems to be something like `return __all_sync(rSq >= rCritSq) ? 1 : 0`, right?I'm not familiar with AMD enough to know if additional synchronization is needed. ChatGPT recommended adding barriers beyond what that gave, but again, I'm not familiar with AMD commands.
Even on NVIDIA, you could've written this without the asm a discussed above!
So in the AMD version, the compiler correctly realized the synchronization was on the comparison, so adds the AMD version right before it. That seems like a straightforward transform here.
It'd be interesting to understand the comparison of what Nvidia primitives map vs what doesn't. The above is a fairly simple barrier. We avoided PTX as much as we could and wrote it as simply as we could, I'd expect most of our PTX to port for similar reasons. The story is a bit diff for libraries we call. E.g., cudf probably has little compute-tier ptx directly, but will call nvidia libs, and use weird IO bits like cufile / gpu direct storage.
Edit: not sure why I just sort of expect projects to be open source or at least source available these days.
Yes, we're not open source, however our license is very permissive. It's both in the software distribution and viewable online at https://docs.scale-lang.com/licensing/
It's open source with a long delay, but paying users get the latest updates.
Make the git repo from "today - N years" open source, where N is something like 1 or 2.
That way, students can learn on old versions, and when they grow into professionals they can pay for access to the cutting Edge builds.
Win win win win
I'm curious, for what reasons are you interested in the source code yourself?
Now should you make business decisions based on that? Probably not. But while I don't claim to be a representative sample, I am pretty sure the number of people who share my beliefs in this regard is substantially "non zero". shrug
I am the founder/editor of PLDB. So I try to do my best to help people "build the next great programming language".
We clone the git repos of over 1,000 compilers and interpreters and use cloc to determine what languages the people who are building languages are using. The people who build languages obviously are the experts, so how they go so goes the world.
We call this measurement "Foundation Score". A Foundation Score of 100 means 100 other languages uses this language somehow in their primary implementation.
It is utterly dominated by open source languages, and the disparity is only getting more extreme.
You can see for yourself here:
https://pldb.io/lists/explorer.html#columns=rank~name~id~app...
Some that might have become irrelevant have gained a second wind after going open source.
But some keep falling further behind.
I look at Mathematica, a very powerful and amazing language, and it makes me sad to see so few other language designers using it, and the reason is because its closed source. So they are not doing so hot, and that's a language from one of our world's smartest and most prolific thinkers that's been around for decades.
I don't see a way for a new language to catch on nowadays that is not open source.
We do believe in open source software and we do want to move the GPGPU market away from fully closed languages. The future is open for discussion but regardless, the status-quo at the moment is a proprietary and dominant implementation which only supports a single vendor.
> I don't see a way for a new language to catch on nowadays that is not open source.
I do note that CUDA is itself closed source -- while there's an open source implementation in the LLVM project, it is not as bleeding edge as NVIDIA's own.
And this is a good point. However, it also has a 17 year head start, and many of those years were spent developing before people realized what a huge market there was.
All it will take is one committed genius to create an open source alternative to CUDA to dethrone it.
But they would have to have some Mojo (hint hint) to pull that off.
Imagine the shift of capital if for example, Intel GPUS suddenly had the same ML software compatibility as Nvidia
On the other hand, if they want better adoption (which would drive sales of their hardware) then Intel / AMD should make a deal to release it as opensource. Closed source will make some profit, but not that much. If this thing really means that everything can run on AMD GPU cards today, then this is a game changer and is worth a lot.
Why do we think all software should be free, and then think that those that don't give it away are the abnormal ones?
Why do people return Windows laptops when they have to pay for a Windows License Activation? Because every single OEM pays for it; you don't expect to buy Windows because it is a failed B2C business model. Nobody wants it. Same goes for proprietary UNIX, and people wish it was the case for Nvidia drivers. I own CUDA hardware and lament the fact that cross-industry GPGPU died so FAANG could sell licensed AI SDKs. The only thing stopping AI from being "free" is the limitations OEMs impose on their hardware.
> that those that don't give it away are the abnormal ones?
They are. Admit it; the internet is the new normal, if your software isn't as "free" as opening a website, you're weird. If I have to pay to access your little forum, I won't use it. If I have to buy your app to see what it's like, I'll never know what you're offering. Part of what makes Nvidia's business model so successful is that they do "give away" CUDA to anyone that owns their hardware. There is no developer fee or mandatory licensing cost, it is plug-and-play with the hardware. Same goes for OpenAI, they'd have never succeeded if you had to buy "the ChatGPT App" from your App Store.
The internet echo chamber strikes again. Exactly how many people are actually doing this? Not many, and those that are all hangout together. The rest of the world just blindly goes about their day using Windows while surfing the web using Chrome. Sometimes, it's a good thing to get outside your bubble. It's a big world out there, and not everybody sees the world as you do
Paying for Windows? I think you missed my point. If your computer doesn't ship with an OS, paid or otherwise, people think it's a glitch. The average consumer will sooner return their laptop before they buy a license of Windows, create an Install Media from their old device and flash the new hardware with a purchased license. They'll get a Chromebook instead, people don't buy Windows today.
The internet has conditioned the majority of modern technology users to reject and habitually avoid non-free experiences. Ad-enabled free platforms and their pervasive success is all the evidence you need. Commercial software as it existed 20 or 30 years ago is a dead business. Free reigns supreme.
What nonsense. Go into any business and you will find every single piece of software they use is bought and paid for with bells on. The 'Free World' you speak of is only there to get you, an individual, used to using the software so that businesses are made to purchase it. In the old days we called this 'demo' or 'shareware'. Now its 'free' or 'personal' tier subscription.
Go and ask any designer if their copy of Adobe Creative Cloud, 3d studio Max, or AutoCAD is free. Any office worker if Micsrosoft Office(including Teams and Sharedpoint etc) or even google docs for business. Majority of developers are running paid versions of Jetbrains. Running an online shop? Chances are you are paying for shopify software, or something like Zoho to manage your customers and orders.
'Free' as you put it is very much only in the online individual consumer world, a very small part of the software world.
The commercial software market is more alive and expensive than it has ever been.
These all have the property which is that they are scarce physical goods or services. Software is not scarce (though of course the labor to create it is), so this is a really bad comparison.
And again I did not say it should or should not be free, I said there are engineering benefits to open source software and more and more people recognize those benefits and choose to make things free because they see the value and are willing to recognize the tradeoffs. I never said what "should" be done. "Should" is kind of a nonsense term when used in this way as it hides a lot of assumptions, so I generally do not use it, and notably did not use it in my comment. I want to point out the peculiarity in your rather strong response to a word and concept I never used. I think you are having an argument with imagined people, not a discussion with me.
And for what it is worth, I am a robotics engineer and I am designing a completely open source solar powered farming robot designed to be made in a small shop in any city in the world (see my profile), funded by a wealthy robotics entrepreneur who recognizes the value in making this technology available to people all over the world.
So I am one of those engineers making this choice, and not someone just asking for things without doing the same of my work. Everything I produce is open source, including person projects and even my personal writing.
Free software, like open science, clearly has something going for it pragmatically. The developer hours put into it have paid for themselves magnitudes of times over. Megacorps hire people to work on free software. If you can't see the value, that's a you problem.
It will be interesting to see if this is the case in the long run, assuming "huge" has a positive connotation in your post, of course.
If AGI comes to pass and it winds up being a net negative for humanity, then the ethics of any practice which involves freely distributing information that can be endlessly copied for very little cost must be reevaluated.
Increasingly, I am not putting much weight in any predictions about whether this will happen in the way we think it will, or what it could possibly mean. We might as well be talking about the rapture.
For now that is fiction, but so is "if all software was free". I do think though that both would lead to a faster rate of innovation in society versus one where critical information is withheld from society to pay someone's rent and food bills.
This is somewhat analogous to music or books/literature. Most composers and performers and authors make no money from people copying and sharing their works. Some pay the bills working professionally for entities who want their product enough to pay for it; some do other things in life. Some indeed give up their work on music because they can't afford to not do more gainful work. And still, neither music nor books go away as copying them gets closer to being free.
AMD just bought company working with similar things for more than 600m.
https://www.amd.com/en/products/accelerators/instinct/mi300/...
Another big AMD fuckup in my opinion. Nobody is going to drop millions on these things without being able to test them out first.
First rule of sales: If you have something for sale, take my money.
"Buy now" buttons and online shopping carts are not generally how organizations looking to spend serious money on AI buy their hardware.
They have a long list of server hardware partners, and odds are you'd already have an existing relationship with one or more of them, and they'd provide a quote.
They even go one step further and show off some of their partners' solutions:
https://www.amd.com/en/graphics/servers-instinct-deep-learni...
FWIW I believe Supermicro and Exxact actually do have web-based shopping carts these days, so maybe you could skip the quotation and buy directly if you were so motivated? Seems kind of weird at this price point.
They could break the trend and offer a "buy now" button instead of offering quotes and coffee chats. It's very likely that will kickstart the software snowball with early adopters.
Nobody is going to drop millions on an unproven platform.
> Seems kind of weird at this price point.
Yeah that $234K server is too much for people to do a trial. It has 8xMI300X GPUs along with a bunch of other shit.
Give me a single MI300X GPU in PCIe form factor for $20K and I'd very seriously consider. I'm sure there are many people who would help adapt the ecosystem if they were truly available.
I'm not looking for entry level hardware.
As far as I'm aware you can't simply buy an Nvidia B200 PCIe card over the counter, either.
You can purchase H100 and A100 PCIe cards over the counter. They're great for compiling CUDA code, testing code before you launch a multi-node job into a cluster, and for running evaluations.
AMD has nothing of the sort, and it's hurting them.
I cannot blow 250K on an SMCI server, nor do I have the electricity setup for it. I can blow 20K on a PCIe GPU and start contributing to the ecosystem, or maybe prove out an idea on one GPU before trying to raise millions from a VC to build a more cost-effective datacenter that actually works.
What are you talking about? Have you looked?
https://www.dell.com/en-us/shop/amd-mi210-300w-pcie-64gb-pas...
https://www.bitworks.io/product/amd-instinct-mi210-64gb-hbm2...
I know this isn't what you're looking for entirely, but my business, Hot Aisle, is working on making MI300x available for rental. Our pricing isn't too crazy given that the GPU has 192GB and one week minimum isn't too bad. We will add on-demand hourly pricing as soon as we technically can.
I'm also pushing hard on Dell and AMD to pre-purchase developer credits on our hardware, that we can then give away to people who want to "kick the tires".
Maybe AMD fears antitrust action, or maybe there is something about its underlying hardware approach that would limit competitiveness, but the company seems to have left billions of dollars on the table during the crypto mining GPU demand spike and now during the AI boom demand spike.
* https://www.levels.fyi/companies/amd/salaries/software-engin...
* https://www.levels.fyi/companies/nvidia/salaries/software-en...
And it's probably better now. Nvidia was paying much more long before, also their stock growing attracts even more talent.
Rumor is that ML engineers (that AMD really needs) are expensive; and AMD doesn't want to give them more money than the rest of the SWEs they have (for pissing off the existing SWEs). So AMD is caught in a bind: can't pay to get top MLE talent and can't just sit by and watch NVDA eat its lunch.
Culturally, none of these companies were happy to pay anyone except the tip, top "distinguished" engineers more than 300K. AMD seems to be stuck in this mentality, just as IBM is.
And that's why creative destruction is essential for technological progress. It's common for organizations to get stuck in stable-but-suboptimal social equilibria: everyone knows there's a problem but nobody can fix it. The only way out is to make a new organization and let the old one die.
This isn't being caught in a bind. This is, if true, just making a poor decision. Nothing is really preventing them from paying more for specialized work.
When I think of AMD ignoring machine learning, I can't help imagine a future YouTuber's voiceover explaining how this caused their downfall.
There's a tendency sometimes to think "they know what they're doing, they must have good reasons". And sometimes that's right, and sometimes that's wrong. Perhaps there's some great technical, legal, or economic reason I'm just not aware of. But when you actually look into these things, it's surprising how often the answer is indeed just shortsightedness.
They could end up like BlackBerry, Blockbuster, Nokia, and Kodak. I guess it's not quite as severe, since they will still have a market in games and therefore may well continue to exist, but it will still be looked back on as a colossal mistake.
Same with Toyota ignoring electric cars.
I'm not an investor, but I still have stakes in the sense that Nvidia has no significant competition in the machine learning space, and that sucks. GPU prices are sky high and there's nobody else to turn to if there's something about Nvidia you just don't like or if they decide to screw us.
Also, ignoring is a strong word: I’m staring at a little << $1000, silent 53 watt mini-PC with an AMD SoC. It has an NPU comparable to an M1. In a few months, with the ryzen 9000 series, NPUs for devices of its class will bump from 16 tops to 50 tops.
I’m pretty sure the linux taint bit is off, and everything just worked out of the box.
https://www.tomshardware.com/news/jensen-huang-and-lisa-su-f...
If they are found colluding due to nepotism, both will get a very swift revocation of business licence and a huge prison term. Remember they are just one step of kinship away from presumed collusion.
At the time, not only did they target AMD (with less compatibility than they have now), but also outperformed the default LLVM ptx backend, and even NVCC, when compiling for Nvidia GPUs!
Nvidia were just ahead in this particular category due to CUDA, so AMD may have just let them run with it for now.
Now AMD is spinning up CUDA compatibility layer after CUDA compatibility layer. It's like trying to beat Windows by building another ReactOS/Wine. It's an approach doomed to fail unless AMD somehow manages to gain vastly more resources than the competition.
Apple's NPU may not be very powerful, but many models have been altered specifically to run on them, making their NPUs vastly more useful than most equivalently powerful iGPUs. AMD doesn't have that just yet, they're always catching up.
It'll be interesting to see what Qualcomm will do to get developers to make use of their NPUs on the new laptop chips.
A complete API comparison table is coming soon, I belive. :D
In a nutshell: - DPX: Yes. - Shuffles: Yes. Including the PTX versions, with all their weird/wacky/insane arguments. - Atomics: yes, except the 128-bit atomics nvidia added very recently. - MMA: in development, though of course we can't fix the fact that nvidia's hardware in this area is just better than AMD's, so don't expect performance to be as good in all cases. - TMA: On the same branch as MMA, though it'll just be using AMD's async copy instructions.
> mapping every PTX instruction to a direct RDNA counterpart or a list of instructions used to emulate it.
We plan to publish a compatibility table of which instructons are supported, but a list of the instructions used to produce each PTX instruction is not in general meaningful. The inline PTX handler works by converting the PTX block to LLVM IR at the start of compilation (at the same time the rest of your code gets turned into IR), so it then "compiles forward" with the rest of the program. As a result, the actual instructions chosen vary on a csae-by-case basis due to the whims of the optimiser. This design in principle produces better performance than a hypothetical solution that turned PTX asm into AMD asm, because it conveniently eliminates the optimisation barrier an asm block typically represents. Care, of course, is taken to handle the wacky memory consistency concerns that this implies!
We're documenting which ones are expected to perform worse than on NVIDIA, though!
This is true first and foremost for the host-side API. From my StackOverflow and NVIDIA forums experience - I'm often the first and only person to ask about any number of nooks and crannies of the CUDA Driver API, with issues which nobody seems to have stumbled onto before; or at least - not stumbled and wrote anything in public about it.
There's a bunch of pretty obscure functions in the device side apis too: some esoteric math functions, old simd "intrinsics" that are mostly irrelevant with modern compilers, etc.
But I can't help but think if something like this can be done to this extend, I wonder what went wrong/why it's a struggle for OpenCL to unify the two fragmentized communities. While this is very practical and has a significant impact for people who develop GPGPU/AI applications, for the heterogeneous computing community as a whole, relying on/promoting a proprietary interface/API/language to become THE interface to work with different GPUs sounds like bad news.
Can someone educate me on why OpenCL seems to be out of scene in the comments/any of the recent discussions related to this topic?
Or you can compile as freestanding c++ with clang extensions and it works much like a CPU does. Or you can compile as cuda or openmp and most stuff you write actually turns into code, not a semantic error.
Currently cuda holds lead position but it should lose that place because it's horrible to work in (and to a lesser extent because more than one company knows how to make a GPU). Openmp is an interesting alternative - need to be a little careful to get fast code out but lots of things work somewhat intuitively.
Personally, I think raw C++ is going to win out and the many heterogeneous languages will ultimately be dropped as basically a bad idea. But time will tell. Opencl looks very DoA.
If they plan to open it up, it can be something useful to add to options of breaking CUDA lock-in.
E.g., how does Cycles compare on AMD vs Nvidia?
Working from cuda source that doesn't use inline ptx to target amdgpu is roughly regex find and replace to get hip, which has implemented pretty much the same functionality.
Some of the details would be dubious, e.g. the atomic models probably don't match, and volta has a different instruction pointer model, but it could all be done correctly.
Amd won't do this. Cuda isn't a very nice thing in general and the legal team would have kittens. But other people totally could.
Mapping inline ptx to AMD machine code would indeed suck. Converting it to LLVM IR right at the start of compilation (when the initial IR is being generated) is much simpler, since it then gets "compiled forward" with the rest of the code. It's as if you wrote C++/intrinsics/whatever instead.
Note that nvcc accepts a different dialect of C++ from clang (and hence hipcc), so there is in fact more that separates CUDA from hip (at the language level) than just find/replace. We discuss this a little in [the manual](https://docs.scale-lang.com/manual/dialects/)
Handling differences between the atomic models is, indeed, "fun". But since CUDA is a programming language with documented semantics for its memory consistency (and so is PTX) it is entirely possible to arrange for the compiler to "play by NVIDIA's rules".
I believe nvcc is roughly an antique clang build hacked out of all recognition. I remember it rejecting templates with 'I' as the type name and working when changing to 'T', nonsense like that. The HIP language probably corresponds pretty closely to clang's cuda implementation in terms of semantics (a lot of the control flow in clang treats them identically), but I don't believe an exact match to nvcc was considered particularly necessary for the clang -x cuda work.
The ptx to llvm IR approach is clever. I think upstream would be game for that, feel free to tag me on reviews if you want to get that divergence out of your local codebase.
Nowadays the issues are more subtle and nasty. Subtle differences in overload resolution. Subtle differences in lambda handling. Enough to break code in "spicy" ways when you try to port it over.
I'm spawned a thread on the llvm board asking if anyone else wants that as a feature https://discourse.llvm.org/t/fexpand-inline-ptx-as-a-feature... in the upstream. That doesn't feel great - you've done something clever in a proprietary compiler and I'm suggesting upstream reimplement it - so I hope that doesn't cause you any distress. AMD is relatively unlikely to greenlight me writing it so it's probably just more marketing unless other people are keen to parse asm in string literals.
Our jerry-rigged solution for now is writing kernels that are the same source for both OpenCL and CUDA, with a few macros doing a bit of adaptation (e.g. the syntax for constructing a struct). This requires no special library or complicated runtime work - but it does have the downside of forcing our code to be C'ish rather than C++'ish, which is quite annoying if you want to write anything that's templated.
Note that all of this regards device-side, not host-side, code. For the host-side, I would like, at some point, to take the modern-C++ CUDA API wrappers (https://github.com/eyalroz/cuda-api-wrappers/) and derive from them something which supports CUDA, OpenCL and maybe HIP/ROCm. Unfortunately, I don't have the free time to do this on my own, so if anyone is interested in collaborating on something like that, please drop me a line.
-----
You can find the OpenCL-that-is-also-CUDA mechanism at:
https://github.com/eyalroz/gpu-kernel-runner/blob/main/kerne...
and
https://github.com/eyalroz/gpu-kernel-runner/blob/main/kerne...
(the files are provided alongside a tool for testing, profiling and debugging individual kernels outside of their respective applications.)
Use an interface over memory allocation/queue launch with implementations in cuda, hsa, opencl whatever.
All the rest of the GPU side stuff is syntax sugar/salt over slightly weird semantics, totally possible to opt out of all of that.
I've also trained the larger GPT2-XL model from scratch on bigger CDNA machines.
Works fine.
[1] https://github.com/anthonix/llm.c [2] https://x.com/zealandic1
That's not important if the goal is to run existing CUDA code on AMD GPUs. All you have to do is write portable CUDA code in the future regardless of what Nvidia does if you want to keep writing CUDA.
I don't know the economics here, but if the AMD provides a significant cost saving, companies are going to make it work.
> Nvidia can always add things to make it difficult
Sounds like Microsoft embedding the browser in the OS. It's hard to see how doing something like that wouldn't trigger an antitrust case.
How turn-key / happy an experience that is depends on how closely your system correlates with one of the documented/tested distro versions and what GPU you have. If it's one that doesn't have binary versions of rocblas etc in the binary blob, either build rocm from source or don't bother with rocblas.
But can this help me directly? Or would OpenAI have to use this tool for me to benefit?
It is not immediately clear to me (but I am a beginner in this space).
Cuda-fortran is not currently supported by scale since we haven't seen much use of it "in the wild" to push it up our priority list.
It wouldn't even be technically possible for SCALE to distribute and use cuBlas, since the source code is not available. I suppose maybe you could do distribute cuBlas and run it through ZLUDA, but that would likely become legally troublesome.
And this is the problem. I guarantee you NVIDIA has more engineers working on cuBLAS et al than AMD does.
The NVIDIA moat is not CUDA the language or CUDA the library. It's CUDA the ecosystem. That means things like all the high performance libraries; all the high performance libraries with clustering support (does AMD even have a clustering solution like NVLink -- everyone forgets that NVIDIA also does high speed networking); all the high perf appliances (everyone also forgets that NVIDIA sells entire systems, not GPUS); all the high perf servers (Triton inference server, etc). We can go on.
I commend the project volunteers for what they've done, but I would recommend getting VC money and competing directly with NVIDIA.
Not sure what is the situation with "CDNA", which is the compute-oriented evolution of "GCN", i.e. whether CDNA is 64-wavefront only or dual like RNDA.
See this part of the documentation for more details regarding warp sizes: https://docs.scale-lang.com/manual/language-extensions/#impr...
1) Kudos
2) Finally !And its linked against an old release of ROCm.
So unclear to me how it is supposed to be an improvement over something like hipify
It appears we implemented `--threads` but not `-t` for the compiler flag. Oeps. In either case, the flag has no effect at present, since fatbinary support is still in development, and that's the only part of the process that could conceivably be parallelised.
That said: clang (and hence the SCALE compiler) tends to compile CUDA much faster than nvcc does, so this lack of the parallelism feature is less problematic than it might at first seem.
NVTX support (if you want more than just "no-ops to make the code compile") requires cooperation with the authors of profilers etc., which has not so far been available
bfloat16 is not properly supported by AMD anyway: the hardware doesn't do it, and HIP's implementatoin just lies and does the math in `float`. For that reason we haven't prioritised putting together the API.
cublasLt is a fair cop. We've got a ticket :D.
For the hardware you are focussing on (gfx11), the reference manual [2] and the list of LLVM gfx11 instructions supported [1] describe the bfloat16 vdot & WMMA operations, and these are in fact implemented and working in various software such as composable kernels and rocBLAS, which I have used (and can guarantee they are not simply being run as float). I've also used these in the AMD fork of llm.c [3]
Outside of gfx11, I have also used bfloat16 in CDNA2 & 3 devices, and they are working and being supported.
Regarding cublasLt, what is your plan for support there? Pass everything through to hipblasLt (hipify style) or something else?
Cheers, -A
[1] https://llvm.org/docs/AMDGPU/AMDGPUAsmGFX11.html [2] https://www.amd.com/content/dam/amd/en/documents/radeon-tech... [3] http://github.com/anthonix/llm.c
Apologies, I appear to be talking nonsense. I conflated bfloat16 with nvidia's other wacky floating point formats. This is probably my cue to stop answering reddit/HN comments and go to bed. :D
So: ahem: bfloat16 support is basically just missing the fairly boring header.
> Regarding cublasLt, what is your plan for support there? Pass everything through to hipblasLt (hipify style) or something else?
Prettymuch that, yes. Not much point reimplementing all the math libraries when AMD is doing that part of the legwork already.
Nvidia actually has more and more capable matrix multiplication units, so even with a translation layer I wouldn't expect the same performance until AMD produces better ML cards.
Additionally, these kernels usually have high sensitivity to cache and smem sizes, so they might need to be retuned.
The online ptx implementation is notable for being even more annoying to deal with than the cuda, but it's just bytes in / different bytes out. No magic.
CUDA has a couple of extra problems beyond just any other programming language:
- CUDA is more than a language: it's a giant library (for both CPU and GPU) for interacting with the GPU, and for writing the GPU code. This needed reimplementing. At least for the device-side stuff we can implement it in CUDA, so when we add support for other GPU vendors the code can (mostly) just be recompiled and work there :D. - CUDA (the language) is not actually specified. It is, informally, "whatever nvcc does". This differs significantly from what Clang's CUDA support does (which is ultimately what the HIP compiler is derived from).
PTX is indeed vastly annoying.
You might also find raw c++ for device libraries saner to deal with than cuda. In particular you don't need to jury rig the thing to not spuriously embed the GPU code in x64 elf objects and/or pull the binaries apart. Though if you're feeding the same device libraries to nvcc with #ifdef around the divergence your hands are tied.
Actually, we just compile all the device libraries to LLVM bitcode and be done with it. Then we can write them using all the clang-dialect, not-nvcc-emulating, C++23 we feel like, and it'll still work when someone imports them into their c++98 CUDA project from hell. :D
Also nvidia might savage you with lawyers for threatening their revenue stream. Big companies can kill small ones by strangling them in the courts then paying the fine when they lose a decade later.
Does NCCL just work? If not, what would be involved in getting it to work?
How do I find out which do I have?
gfx1100 : https://www.techpowerup.com/gpu-specs/amd-navi-31.g998
gfx1030 : https://www.techpowerup.com/gpu-specs/amd-navi-21.g923
gfx1010 : https://www.techpowerup.com/gpu-specs/amd-navi-10.g861
gfx900 : https://www.techpowerup.com/gpu-specs/amd-vega-10.g800
Compile to DFA by repeatedly differentiating then unroll the machine? You'd still have back edges for the repeating sections.
At least grounds for suing and starting an extensive discovery process and possibly a costly injunction...
It was clean-room implemented purely from the API surface and by trial-and-error with open CUDA code.
Namely:
"4.1 License Scope. The SDK is licensed for you to develop applications only for use in systems with NVIDIA GPUs."
how big of a deal is this?
That changing the compiler is strongly equivalent to changing the source doesn't necessarily influence this pattern of thinking. Customer requests to keep the performance gains from a new compiler but not change the UB they were relying on with the old are definitely a thing.
There are some delightful AMD driver issues that make certain models of GPU intermittently freeze the kernel when used from docker. That was great fun when building SCALE's CI system :D.
https://rocm.docs.amd.com/projects/install-on-linux/en/lates...
This sounds like DirectX vs OpenGL debate when I was younger lol