VkFFT: Vulkan/CUDA/Hip/OpenCL/Level Zero/Metal Fast Fourier Transform Library
github.com
github.com
I remember VkFFT got a lot of initial traction thanks to Hacker News three years ago. Back then VkFFT was a simple collection of pre-made shaders for powers of two FFTs.
Nowadays it is based on the runtime code generation and optimization platform that supports all the mentioned backends, has a wide range of implemented algorithms (some of which are not even present in other codes) to cover all system sizes and can do things no other GPU FFT library can so far (like real to real transforms, arbitrary dimensional transforms, zero-padding, convolutions and more).
If you have some questions about the library, design choices, functionality or anything else - I will be happy to answer them!
* futuristic buildings and flying cars *
that’s pretty cool, the vk prefix for a non-vulkan-only library is kind of confusing tho
Here's a compact(!) implementation of "Hello, Triangle!" https://github.com/Planimeter/game-engine-3d/blob/97298715b2...
921 lines of doing nothing.
Vulkan is targeted at graphics and game engine developers who want to extract the maximum possible performance from their GPUs and know the limitations of a global-state API like OpenGL. Vulkan allows extremely fine-grained pipeline management and synchronisation primitives, and adds in hardware ray-tracing support natively without having to uncannily bolt it on.
If all you want is to draw a three-coloured triangle, you can do that easier and faster with ShaderToy instead of fudging with Vulkan. If you want to write a fast, powerful, modern graphics engine, then you use Vulkan or D3D12.
People forget that GPUs are now massive slices of silicon with memory and power subsystems in their own right, and are obviously extremely powerful hardware. At some point, OpenGL itself becomes a bottleneck, or is difficult enough to program with that Vulkan becomes easier, and that's when the true utility and power of its extreme verbosity is displayed.
For the record, the boilerplate isn't '921 lines of nothing'—it's effectively setting the GPU up from scratch, similar to bootstrapping a CPU from 16-bit real mode.
Use a middleware instead might be the answer.
Then we don't need really Vulkan, as the middleware already allows to use the best 3D API on each platform.
[1]: https://kompute.cc/
[1] - https://tinygrad.org/
Very impressive performances. I'd be happy to have a comparison with regular CPU performances... If you put together the GPU time + GPU upload & download, is it faster than CPU overall ?
Also Python: https://github.com/vincefn/pyvkfft
That always depends on sample size and the hardware you use.
And like for all those kind of problems on it also depends on the parallelizability of the computation you are doing.
The other factor is that AMD's data center hardware is not available at any cloud provider so nobody even has access to the supposedly supported hardware.
Also, maybe a bit obvious.. but that even if there is no huge benefit - sending compute to the GPU frees up your CPU/application to do other things .. like keeping your application responsive :)
That's the kind of stuff Nvidia has offered for the last decade while AMD did god knows what.
[0] e.g. https://jamesmccaffrey.wordpress.com/2021/09/02/example-of-a...
https://pytorch.org/tutorials/advanced/dispatcher.html
More complex but still not that hard.
Edit: as an example that I experienced first-hand, coremltools, which converts PyTorch to CoreML models, only gained FFT support very recently. It's also not really a PyTorch backend but a PyTorch code converter, though, so wouldn't benefit at all from PyTorch's FFT being backed by VkFFT. Still, good example that one shouldn't take FFTs for granted.