My program performed complicated processing of large 8-bit grayscale images. Naive use of portable SIMD made several important parts of the algorithm run ten times faster than my integrated-GPU implementation! (OpenGL ES 3.0 over ANGLE on an Intel HD Graphics 520)
I think the sheer power of modern SIMD has slipped past most peoples' radar. A multicore processor with AVX2 has number-crunching throughput comparable to a low-end GPU, and for some workloads, the extra flexibility of the CPU can unlock an additional order-of-magnitude performance uplift. (The comparison may be a little different if you're programming the GPU with modern compute shaders - I'm not sure.)
Portable SIMD programming is also just fun! Trying to keep my working set within the 32 KiB L1 cache, trying to keep my data packed within individual 8-bit lanes, and lacking access to some basic operations like integer multiplication and division - it almost feels like programming for a retro games console.
[0]: https://github.com/ispc/ispcIt is a shame that one needs to fiddle with RUSTC_FLAGS just for SIMD extensions to trigger though :/
I get this feeling writing CUDA code all the time, having to keep my stack frames and local data small, figuring out how to even out and time-slice the work. It’s a little like Arduino coding in a way, it’s just that instead of one of them running at 16Mhz, I have ten thousand of them running at 2Ghz (and as a bonus they have floating point superpowers).
There is an article on the web explaining the purpose of SVE/SVE2 ("Scalable Vector Extension"), which is supposed to be the successor of SIMD on ARM : https://levelup.gitconnected.com/armv9-what-is-the-big-deal-...
Extract : "[...] the addition of SIMD instructions has led to an explosion in the number of instructions, especially for x86. And of course not every x86 processor will support all these instructions. Only the newer ones will support AVX-512. The beauty of SVE is that the same code will work for both the super-computer and the cheap phone. That is not possible with the x86 SIMD instructions."
There is also a Java proposal to use SVE as the way of doing SIMD in the Java world : https://openjdk.java.net/jeps/417
The same principle will be extended on ARM to matrices with "Scalable Matrix Extension" : https://community.arm.com/arm-community-blogs/b/architecture...
We can speculate that everyone will migrate to ARM / RISC-V at some point, or x86 will have similar instructions.
> the same code will work for both the super-computer and the cheap phone. That is not possible with the x86 SIMD instructions
Actually, when the code is expressed using "portable intrinsics" (https://github.com/google/highway), the source code looks the same but compiles to SSE4/AVX2/AVX-512 and NEON,SVE,SVE2 and RISC-V V instructions.
Disclosure: I am the main author of this library.
An important nontechnical problem with these is that there is such a long way to travel for a language to reach widespread adoption, needing credible long term sponsorship and commitment from a broad shouldered user community or private organisation, etc.
What transformations do you think C++ doesn't support?
I've got a _lot_ of experience in C++ and I've found that it's almost always the case that whatever library being used needs to be rewritten for a small change. Rewriting libraries is rarely approved. And further, there's often very little unit testing being done towards that end too.