Also note that it wasn’t well received 2 years ago. Example: https://news.ycombinator.com/item?id=23809335
Also note that it wasn’t well received 2 years ago. Example: https://news.ycombinator.com/item?id=23809335
This is bad with dGPUs over the PCIe bus, but not so much with GPUs that share a very fast memory bus with the CPU. In this case, the layout of the data may prove challenging to keep the same for when you use a CPU and a GPU.
For 64-bit keys, we sort about 1 GB/s per (5 year old) Skylake core, and perhaps 5-6 parallel.
This (2018) reports 3.5 GB/s: https://benkarsin.files.wordpress.com/2018/10/dissertation.p... And a 6-year old GPU radix sort reports 2.1 GB/s: https://github.com/Bulat-Ziganshin/Compression-Research/tree...
BTW I've worked on a product that used GPUs. That typically requires everything to move to the GPU, which is not always desirable or feasible.
https://dev.to/tishden/computing-with-gpu-why-when-how-and-s...
This shows a approximately 20x speedup (2021, graph 1 vs 3): https://www.irjet.net/archives/V8/i7/IRJET-V8I7714.pdf
The "25-fold speedup", as is often the case for such reports, comes from not optimizing the CPU side.
Have you seen the state of GPU software development? GPUs are very expensive in cloud, are poorly supported in containers and virtual machines If you want to use GPU compute, some stuff is Nvidia-only, some stuff is glitching and crashy, it's probably not avaliable in your language of choice, etc.
It is literally impossible for me to add GPU compute to any of our corporate workloads, but I can tap into AVX easilly in my language of choice.
I'm pretty sure GPUs closes that gap over time though.