The GPU is not always faster
cowfreedom.de
cowfreedom.de
Roofline plots [1] are framework to visualize system design from this perspective.
I don’t think it’s accurate that only trivial implementations use the direct o(n^3) algorithm. AFAIK high performance BLAS implementations just use highly optimized versions of it.
I just wish I understood the tricks done to make it so fast so I could implement my own for variations for which there are no pre-existing BLAS implementations. The best BLAS implementations are all closed source sadly.
NVidia open-sourced CUTLASS [0] some years ago and it achieves pretty competitive performance compared to e.g. the closed-source cuBLAS.
Keen observers will notice that Strassen is not used in CUTLASS.
CUDNN supports FFT to do matmul as well as convolution/correlation and can also be configured to automatically use the best algorithm.
In some cases the FFT method has the incidental side-benefit of data reuse, like in the case of FIR filters, where the data allows for partitioned convolution.
In this article the deciding factor seems to be the startup cost because the application has placed the data on the CPU memory side, and is considering shipping out to GPU memory just for this computation.
PCIe 3.0? What?
https://cowfreedom.de/#appendix/computer_specs/
> GeForce GTX 1050 Ti with Max-Q Design (PCIe 3.0 x16) (2016)
> Intel Core i5-8300H (2020)
This is a low-price 8 year old GPU and a 4 year old CPU. And he seems to be including loading the data to GPU. Newer cards have wide PCIe 5.0 or some faster interconnect, like Nvidia Grace-Hopper.
Also he is comparing his own CUDA implementation. He should use one of the many available in CUBLAS/CUTLASS. Making a good CUDA GEMM is a very difficult art and very hardware specific
There aren't any (edit: consumer) GPUs with PCIe5 yet, though they probably aren't far off. Plenty already have PCIe4 though.
Nvidia is withholding new releases but the current hardware has more legs with new matrix implementations. Like FlashAttention doing some significant improvement every 6 months.
Nvidia could make consumer chips with combined CPU-GPU. I guess they are too busy making money with the big cloud. Maybe somebody will pick up. Apple is already doing something like it even on laptops.
https://www.datacenterknowledge.com/data-center-hardware/nvi...
Some low-end or middle range GPUs really only use 8 lanes, because that's how the chip is designed. Fewer lanes -> less silicon area for the pcie logic needed -> cheaper.
The chips which use 16 lanes take advantage of it and can saturate the link.
"PCI Express 4.0 x8"
[root@gh200 nvbandwidth]# nvidia-smi | grep GH200
| 0 NVIDIA GH200 480GB On | 00000009:01:00.0 Off | 0 |
[root@gh200 nvbandwidth]# ./nvbandwidth | grep -E 'host_to_device_memcpy_sm|device_to_host_memcpy_sm' | grep ^SUM
SUM host_to_device_memcpy_sm 357.45
SUM device_to_host_memcpy_sm 352.05
[root@gh200 nvbandwidth]#
d2d is much higher:
[root@gh200 cuda-samples]# ./bin/sbsa/linux/release/bandwidthTest --dtod | grep -B1 32000
Transfer Size (Bytes) Bandwidth(GB/s)
32000000 2294.8
[root@gh200 cuda-samples]#For problems like Matrix Multiplication, it costs N to communicate the problem but N^2 operations to calculate.
For problems like dot product, it costs N to communicate but only N operations to calculate.
Compute must be substantially larger than communication costs if you hope to see any benefits. Asymptotic differences obviously help, but linear too might help.
You'd never transfer N data to perform a log(n) binary search for example. At that point communication dominates.
If you’re willing to work entirely with the gpu memory the gpu will of course be faster even in this scenario.
* If the GPU is a non-dedicated older style intel GPU, use CPU
* If the GPU is a non-dedicated anything else do anything super parallel on the GPU but anything that can be BRRRRRT via CPU on the CPU because the memory is shared.
* If the GPU is dedicated move everything to GPU memory and keep it there and only pull back small statistics if at all plausible.
- Either way, you're writing 'cuda style' fine-grained data parallel code that looks and works very different from regular multithreaded code. You are now in a different software universe.
- You now also have to think about throughput, latency hiding, etc. Nvidia has been commoditizing throughput-oriented hardware a lot better than others, and while AMD is catching up on some workloads, Nvidia is already advancing. This is where we think about bandwidth between network/disk=>compute unit. My best analogy here, when looking at things like GPU Direct Storage/Network, is CPU systems feel like a long twisty straw, while GPU paths are fat pipes. Big compute typically needs both compute + IO, and hardware specs tell you the bandwidth ceiling.
To a large extent, ideas are cross-polinating -- CPUs looking more like GPUs, and GPUs getting the flexibility of CPUs -- but either way, you're in a different universe of how code & hardware works than 1990s & early 2000s intel.
So, GPUs have the slight disadvantages that you have to think an about data movement and the drivers are a little less convenient to install, but it isn’t really a big deal.
Traditional AI/ML models (including smaller transformers) can definitely be optimized for mass scale/performance on cpu-optimized infrastructure.
The use case of transferring ALL data over every time is obviously misusing the GPU.
If anyone’s ever tried running a model that’s too large for your GPU you will have experienced how slow this is when you have to pull in the model in parts for a single inference run.
Is this surprising or obvious?
Is there a possibility that in the not too distant future that GPUs and CPUs will just converge? Or are the tasks done by GPUs too specialized?
My dev GPU is a 6800XT. Cheapish gaming card from a little while ago, 16GB ram on the card. 72 "compute units" which are independent blocks of hardware containing memory ports, floating point unit, register file etc. Roughly "a core" from x64 world. Each of those can have up to 64 tasks ready to go, roughly a "hyperthread". It's 300W or so.
There's some noise in the details, e.g. the size of the register file from the perspective of a hyperthread affects how many can be resident on the compute unit ready to run, the memory hierarchy has extra layers in it. The vector unit is 256byte wide as opposed to 64byte wide on x64.
But if you wanted to run a web browser entirely on the GPU and were sufficiently bloody minded you'd get it done, with the CPU routing keyboard I/O to it and nothing else. If you want a process to sit on the GPU talking to the network and crunching numbers, don't need the x64 or arm host to do anything at all.
CPUs aim to minimize latency (how many cycles have to pass before you can use a result), and do so by way of high clock frequencies, caches and fancy micro architectural tricks. This is what you want in most general computation cases where you don't have other work to do whilst you wait.
GPUs instead just context switch to a different thread whilst waiting on a result. They hide their latency by making parallelism as cheap as possible. You can have many more cores running at a lower clock frequency and be more efficient as a result. But this only works if you have enough parallelism to keep everything busy whilst waiting for things to finish on other threads. As it happens that's pretty common in large matrix computations done in machine learning, so they're pretty popular there.
Will they converge? I don't think so - they're fundamentally different design points. But it may well be that they get integrated at a much closer level than current designs, pushing the heterogeneous/dark silicon/accelerator direction to an extreme.
If you take a serial algorithm and put it on the GPU, it's easy to verify that a single GPU thread is much slower than a single thread on the CPU. For example, just do a bubble sort on the GPU with a single thread. I'm not even including the time to transfer data or read the result. You'll easily find the CPU is way faster.
The way you get GPU speed is by finding/designing algorithms that are massively parallel. There are lots of them. There are sorting solutions for example.
As an example, 100 cores * 32 execution units per core = 3200 / 20 = 160x faster than the CPU if you can figure out a parallel solution. But, not every problem can be solved with parallel solutions and if it can't then there's where the CPU wins.
It seems unlikely GPU threads will be as fast as CPU threads. They get their massive parallelism by being simpler.
That said, who knows what the future holds.
Here's mine
https://jsfiddle.net/jw7a6to9/ bubblesort
https://jsfiddle.net/y1w6s9tj/ taylor series
> in which case you can show those same GPU cores are 20x faster than CPU for other things
Which things? Remember, I wrote single thread, no SIMD, no samplers. It's the parallelism that provides the speed.
The CPU would always be slower if the data originated in GPU memory.
But about bandwidth, matrix multiplications happen mostly in cache and that has a lot more bandwidth than RAM. Blocks of the matrix are loaded to cache (explicitly in CUDA) and used multiple times there.
I'd exploit the better multi-level cache hierarchy in CPUs and make the code NUMA aware. But still I wouldn't bet against a recent GPU card.
The post is about dot product, not matrix multiply. Dot product has no data reuse
The author has updated the post with corrected AVX measurements, with the original ~340 GB/s revised down to 31.7 GB/s. (Thanks CowFreedom)
a GH200 will run miles around any CPU.
well, to play the devil advocate, for the outside event to affect the CPU, a signal will have to go through the PCI bus or equivalent.