GPU Survival Toolkit for the AI age
journal.hexmos.com
journal.hexmos.com
I'd also like to point out that 90 % of the time spent to "compute" the Mandelbrot set with the JIT-compiled code is spent on compiling the function, not on computation.
If you actually want to learn something about CUDA, implementing matrix multiplication is a great exercise. Here are two tutorials:
There is SAXPY (matrix math A*X+Y), purportedly ([1]) the hello world of parallel math code.
>SAXPY stands for “Single-Precision A·X Plus Y”. It is a function in the standard Basic Linear Algebra Subroutines (BLAS)library. SAXPY is a combination of scalar multiplication and vector addition, and it’s very simple: it takes as input two vectors of 32-bit floats X and Y with N elements each, and a scalar value A. It multiplies each element X[i] by A and adds the result to Y[i].
- You might want to use some ML in the project you are assigned next month
- It can help collaborating with someone who tackles that aspect of a project
- Fundamental knowledge helps you understand the "AI" stuff being marketed to your manager
The "I don't need this adjacent field" mentality feels familiar from schools I went to: first I did system administration where my classmates didn't care about programming because they felt like they didn't understand it anyway and they would never need it (scripting, anyone?); then I switched to a software development school where, guess what, the kids couldn't care about networking and they'd never need it anyway. I don't understand it, to me it's both interesting, but more practically: fast-forward five years and the term devops became popular in job ads.
The article is 1500 words at a rough count. Average reading speed is 250wpm, but for studying something, let's assume half of that: 1500/125 = 12 minutes of your time. Perhaps you toy around with it a little, run the code samples, and spend two hours learning. That's not a huge time investment. Assuming this is a good starting guide in the first place.
I've come to see this sort of clickbait headline as playing on the prevalence of imposter-syndrome insecurity among devs, and try to ignore them on general principle.
I remember how I joined a startup after working for a traditional embedded shop and a colleague made (friendly) fun of me for not knowing how to use curl to post a JSON request. I learned a lot since then about backend, frontend and infrastructure despite still being an embedded developer. It seems likely that people all around the industry will be in a similar position when it comes to AI in the next years.
SageMaker is JupyterLab with a GPU attached.
Cognito is just OAuth.
And of course networking is fucked… somehow AWS made it more complicated than the real thing, like they abstracted in the opposite direction.
Terrible article IMO.
I agree that most developers are not AI developers... OP seems to be a bit out of touch with the general population and otherwise is assuming the world around them based on their own perception.
Lost of programmers started with an understanding of what happens physically on the hardware when code runs and it is unfair advantage when debugging at times
I agree, but to say that all developers must know how AI benefits from GPUs is a different claim. One which is false. I would say most developers don't even understand how the CPU works, let alone modern CPU features like Data/Instruction Caching, SIMD instructions, and Branch prediction.
Most developers I encounter learned Javascript and make websites
AI is also just software that runs on hardware.
GPUs are just hardware
(I'm not touching Nvidia since they don't provide open source drivers.)
GPUs are extremely performant, and also very hard to code in, so people just use highly abstracted API calls like pytorch to command the GPU.
C is very performant, and hard to code in, so people just use python as a abstraction layer over C.
Its not clear if people need to understand GPUs that much (Unless you are deep in AI training/ops land). In time, since moore's law has ended and multithreading becomes the dominant mode of speed increases, there'll probably be brand new languages dedicated to this new paradigm of parallel programming. Mojo is a start.
As in, every instruction, from a simple loop of calculations onward, is designed behind the scenes so that it intelligently maximises usage of every available CPU core in parallel, and also farms everything possible out to the GPU?
Has this been done? Is it possible?
There are little bits of research on algorithm replacement. Like, have the compiler detect that you're trying to sort, and generate the code for quick sort or timsort. it works, kinda. There are a lot of ways to hide a sort in code, and the compiler can't readily find them all.
There is also https://futhark-lang.org/ , though I haven’t tried it, just heard about it.
It's a bit like "design a zip algorithm which can compress any file".
Wut? We hit the power wall back in 2004. There was a little bit of optimization around the memory wall and ilp wall afterwards, but really, cores haven't gotten faster since.
It's been all about being able to cram more cores in since then, which implies at least multi-threading, but multi-processing is basically required to get the most out of a cpu these days.
How do you operate in that world if "multithreading isn't the answer"?
Spreading an algorithm across multiple threads makes it more difficult for an optimizing compiler to find opportunities for SIMD optimization.
Similarly to how modern languages make it easier to safely use threads, runtimes also make it easier to take advantage of SIMD optimizations. For example, recently a SIMD-optimized sorting algorithm was included in OpenJDK. Apart from that, SIMD is way less brittle at runtime than GPUs and other accelerators.
* Overhead: the overhead to start and manage multiple threads is considerable in practice. Most multithreaded algorithms are in fact slower than optimized serial implementations when n_threads=1
* Communication: threads have to communicate with each other and synchronize access to shared resources. "Embarrassingly parallel" problems don't require synchronization, but many interesting problems are not of that kind.
* Amdahl's law: there is a point of diminishing returns on parallelizing an application since it quite likely contains parts that are not easily parallelized.
Edit: latency is difficult. But accellerating CPUs and using GPUs for compute was never about latency. Most I/O bottlenecks are because CPUs have sped up so much and left the rest of the platform in the dust. Much of it is also due to fundamental limitations due to the speed of light. Increasing throughput is always easier that reducing latency.
From a circuit complexity standpoint, you could evaluate a wide but shallow circuit to evaluate the function in a nanosecond, or a deep but narrow circuit that takes eons. Whether parts of that circuit are evaluated synchronously or asynchronously is immaterial, although synchronous computation does seem easier to reason about and the UX is more user-friendly from a programmability standpoint.
I agree with you the fundamental limitation is the speed of light if you are width-bounded (i.e., if physical space is the dominating constraint).
I'm sure CUDA is great, and if I had more free time and/or better reasons to improve the performance of my code it would probably be great for me. My point was mainly that a few lines of code which may be trivial for one person to write may not be for someone else with different experience. Depending on what the code is being used for even a vast increase in performance may not be worth the extra time it takes to implement it.
[0] https://docs.nvidia.com/cuda/cuda-c-programming-guide/index....
If you're already familiar with one of the languages that the nvidia compiler supports? Not that many. For people familiar with C or C++, it's a couple extra attributes, and a slightly different syntax for launching kernels vs a regular function call. I'm admittedly not experienced with Fortran, which is the other language they support, so I can't speak to that. There's c-style memory allocation functions, which might be annoying to C++ devs, but it's nothing that would confuse them.
Edit: There's also a couple weird magic globals you have access to in a kernel (blockIdx, blockDim, threadIdx), but those are generally covered in the intros.
But once that challenge is overcome, GPU truly rocks.
Finally debugging complex shaders (I do some specific case of computational fluid dynamics where equations are not that easy, full of "if/then" edge cases, etc) is not fun at all, tooling is sorely missed (unless I've missed something)
But we are in a comment chain spawned by:
> CUDA is fairly straightforward for many tasks and in many cases there is an easy 100x improvement in processing speed just sitting there to be had with <100 lines of code.
And a follow up comment about how easy it would be to write that "<100 lines of code", so I feel like we're definitely talking about the easy case of naturally parallel calculations, and sticking to that as an intro seems fair to me.
And there's also value in seeing how other people approached a problem.
C is a way of life. Those of us who code exclusively, or nearly so, in C cannot stomach python's notion of "significant white-space."
Complaining about significant white-space is like complaining that lisp has too many parentheses. It’s an aesthetic preference that just doesn’t matter in practice.
Syntax is mostly an irrelevance, they have surprisingly similar patterns in my opinion.
In a modern language I want a type system that both reduces risk and reduces typing — safety and metaprogramming. C obviously doesn't, python doesn't really either.
Python's approach to dynamic-ness is very similar to how I'd expect C to be as a dynamic language (if it had proper arrays/lists).
Why belly ache about it? Whitespace is significant to one’s fellow humans.
Well-known for supporting any formatting style you like ;)
Ha! I wish CPUs were still that simple.
Granted, it is legitimate for the article to focus on the programming model. But "CPUs execute instructions sequentially" is basically wrong if you talk about performance. (There are pipelines executing instructions in parallel, there is SIMD, and multiple cores can work on the same problem.)
A CPU has ~100 "cores" each running one (and-a-hyperthread). independent things, and it hides memory latency by branch prediction and pipelining.
A GPU has ~100 "compute units", each running ~80 independent things interleaved, and it hides memory latency by executing the next instruction from one of the other 80 things.
Terminology is a bit of a mess, and the CPU probably has a 256bit wide vector unit while the GPU probably has a 2048bit wide vector unit, but from a short distance the two architectures look rather similar.
And memory access order is much more important that on CPU. Truly random access has very bad performance.
Call it Xeon Chi.
If you use high-bandwidth high-latency GDDR memory, CPU cores will underperform due to high latency, like there: https://www.tomshardware.com/reviews/amd-4700s-desktop-kit-r...
If you use low-latency memory, GPU cores will underperform due to low bandwidth, see modern AMD APUs with many RDNA3 cores connected to DDR5 memory. On paper, Radeon 780M delivers up to 9 FP32 TFLOPS, the figure is close to desktop version of Radeon RX 6700 which is substantially faster in gaming.
I think they may have a hurdle of getting folks to buy into the concept though.
I imagine it would be analogous to how Arria FPGA’s were included with certain Xeon CPU’s. Which further backs up your point that this could happen in the near future!
Edit: Oh, thanks for the downvote, with no discussion of the question. I'll just sit here quietly with my commercial OpenCL software that happily exploits these vector units attached to the normal CPU cores.
I did decide not to engage because “you mean like <very common well known thing>” seemed a bit brusque and dismissive, commenting on this site is just for fun, so I don’t really see the point in continuing a conversation that seems like it is getting off on the wrong foot.
Given that most programming languages are designed for sequential processing (like CPUs), but Erlang/Elixir is designed for parallelism (like GPUs) … I really wonder if Nx / Axon (Elixir) will take off.
This is the best one imo.
I like attacking complexity head on, and have a good knowledge of both quantitative methods & qualitative details of (say) computer hardware so having an article that can tell me the nitty gritty details of a field is appreciated.
Take "What every programmer should know about memory" — should every programmer know? Perhaps not, but every good programmer should at least have an appreciation of how a computer actually works. This pays dividends everywhere — locality (the main idea that you should take away from that article) is fast, easy to follow, and usually a result of good code that fits a problem well.
If you don't want to follow the happy path, on Nvidia you get to beg them to maybe support your use case in future. On amdgpu, you get the option to build it yourself, where almost all the pieces are open source and pliable. The driver ships in Linux. The userspace is on GitHub. It's only GPU firmware which is an opaque blob at present, and that's arguably equivalent to not being able to easily modify the silicon.
The other problem is that there aren't any places to rent the high end AMD AI/ML GPUs, like the MI250's and soon to be released MI300's. They are only available on things like the Frontier super computer, which few developers have access to. "regular" developers are stuck without easy access to this equipment.
I'm working on the later problem. I'd like to create more of a flywheel effect. Get more developers interested in AMD by enabling them to inexpensively rent and do development on them, which will create more demand. @gmail if you'd like to be an early adopter.
Am I lazy to expect we’ll be getting a lot more “parallel-on-the-GPU-for-free” in the future?
If that's true, then I'm surprised they only see a 10x speed-up. I would expect more from only compiling that loop for the CPU. (Comparing to interpreted Python without numpy.) Given they already have a numba version, why not compile it for the CPU and compare?
Also, they say consumer CPUs have 2-16 cores. (Who has 2 cores these days?) They go on suggest to rent an AWS GPU for $3 per hour. You're more likely to get 128 cores for that price, still on a single VM.
Not saying it will be easy to write multi-threaded code for the CPU. But if you're lucky, the Python library you're using already does it.
Pretty sure my mom's laptop has 2 cores; I can't think of anyone whose daily driver has 16 cores. Real cores, not hyperthread stuff running at 0.3× the performance of a real core.
As for the 128-core server system, note that those cores are typically about as powerful as a 2008 notebook. My decade-old laptop CPU outperforms what you get at DigitalOcean today, and storage performance is a similar story. The sheer number makes up for it, of course, but "number of cores" is not a 1:1 comparable metric.
Agree, though, that the 10x speedup seems low. Perhaps, at 0.4s, a relatively large fraction of that time is spent on initializing the Python runtime (`time python3 -c 'print("1337")'` = 60ms), the module they import that needs to do device discovery, etc.? Hashcat, for example, takes like 15 seconds to get started even if it then runs very fast after that.
I don't think the price matters btw. If 4 cores is all I need (I can't think of a desktop application that I use which benefits from more than 4 cores more than from faster cores, which is typically the trade-off at >4 real cores) then that's what I'd get because that's the optimum for my wallet and performance profile.
Use CUDA, vice graphics APIs+compute. The latter (Vulkan compute etc) is high friction. CUDA is far easier to write in, and the resulting code is easier to maintain and update.
But for some reasons professional programmers are judged under a much higher moral standard.
If I start to work on a tool, then I cannot work anymore on what I actually wanted to do. And it just so happens ... that this is exactly what I did and I can just say, it usually takes way longer than the most pessimistic estimate one can come up with, so yes, one can decide to switch careers and try to get funding to (re)build what is not offered to acceptable conditions (but in my case the tool simply did not exist, though).
Just like an artist can switch career, study CS, build on his own a tool a professional company build with a team over years - and then someday work with his tool to acomplish his original work. In (simplified) theories, lots of things are possible ..
Not in the real world. Most programmers who are trying to get a job done won’t avoid CUDA or AWS or other tools just to avoid “contributing to a monopoly”. When responsible programmers have a job to do and tools are available to help with the job, they get used.
A programmer who avoids mainstream tools on principle is liable to get surpassed by their peers very quickly. I’ve only met a few people like this in industry and they didn’t last very long trying to do everything the hard way just to avoid tools from corporations or monopolies or open source that wasn’t pure enough for their standards.
It’s only really in internet comment sections that people push ideological purity like this.
I believe the key word there is "professional" -- one of the challenges of a venue like HN is the professional engineers and the less-professional ones interact from worldviews and use cases so distinct that they may as well be separate universes. In other spaces, we wouldn't let a top doctor have to explain very basic concepts about the commercial practice of medicine to an amateur "skeptic" and yet so many discussions on HN degenerate along just these lines.
On the other hand, it's that very same inclusiveness and generally high discourse in spite of that wide expanse which make HN such a special community, so I'm not sure what to conclude besides this unfortunate characteristic being a necessary "feature, not a bug" of the community. There's no way around it that wouldn't make the community a lesser place, I think.
Vote with your feet. Maybe you can't or can't afford it, then at least admit the problem to yourself and maybe don't try to persuade others in order to feel better for your own decision.
..and doesn't suck.
HIP is basically that, but they still make you jump through hoops to rename everything etc.
There are libraries written at a lower level that wouldn't be immediately portable, but surely that could be addressed over time as well.
Currently I've given up and use runpod, but still...
They just leverage parallelism in making a single prediction
The section goes on to teach Amazon-specific terminology and products.
A "bare minimum everyone must know" guide should not include vendor-specific guidance. I had this in school with Microsoft already, with never a mention of Linux because they already paid for Windows Server licenses for all of us...
Edit: and speaking of inclusivity, the screenshots-of-text have their alt text set to "Alt text". Very useful. It doesn't need to be verbatim copies, but it could at least summarize in a few words what you're meant to get from the terminal screenshot to help people that use screen readers.
Since this comment floated to the top, I want to also say that I didn't mean for this to dominate the conversation! The guide may not be perfect, but it helped me by showing how to run arbitrary code on my GPU. A few years ago I also looked into it, but came away thinking it's dark magic that I can't make use of. The practical examples in both high- and low-level languages are useful
Another edit: cool, this comment went from all the way at the top to all the way at the bottom, without losing a single vote. I agree it shouldn't be the very top thing, but this moderation also feels weird
Tbh the most basic question is: “are you innovating inside the AI box or outside the AI box?”
If inside - this guide doesn’t really share anything practical. Like if you’re going to be tinkering with a core algorithm and trying to optimize it, understanding BLAS and cuBLAS or whatever AMD / Apple / Google equivalent, then understanding what pandas, torch, numpy and a variety of other tools are doing for you, then being able to wield these effectively makes more sense.
If outside the box - understanding how to spot the signs of inefficient use of resource - whether that’s network, storage, accelerator, cpu, or memory, and then reasoning through how to reduce that bottleneck.
Like - I’m certain we will see this in the near future, but off the top of my head the innocent but incorrect things people do: 1. Sending single requests, instead of batching 2. Using a synchronous programming model when asynchronous is probably better 3. Sending data across a compute boundary unnecessarily 4. Sending too much data 5. Assuming all accelerators are the same. That T4 gpu is cheaper than an H100 for a reason. 6. Ignoring bandwidth limitations 7. Ignoring access patterns
Even when I was working at an Azure-only shop, I've never actually seen anyone use Windows Server. Lots of CentOS (before IBM ruined it) and other Unixes, but never a Windows Server.
It's useful to have experienced, but I do take issue with exclusively (or primarily) focusing on one ecosystem as a mostly-publicly-funded school.