It's just not been an aim even in more recent-ish languages like Rust, Swift, Go etc. Partly because of the unworkably hard task of targeting fragmented, undebuggable, proprietary driver/OS programming interfaces from POV of portable languages.
It's just not been an aim even in more recent-ish languages like Rust, Swift, Go etc. Partly because of the unworkably hard task of targeting fragmented, undebuggable, proprietary driver/OS programming interfaces from POV of portable languages.
But many languages and common practices push people in a highly incremental piece by piece approach. I'd argue we're basically taught to think in terms of loops, from our earliest days of programming. We could just as easily be taught to think of things in terms of bulk arrays. The win in programming this way is extremely high.
I was very disappointed to find the SIMD story in Rust so underdeveloped.
Yes, people are not running weather forecasting on their home computers. But they're playing video, synthesizing audio, doing speech recognition, etc. all the time. And more of this ever day.
And if there's any chance of rescuing the uses of ML from a hellish privacy violating landscape worth of a Butlerian Jihad... it would by making sure most of this type of computation was done locally on-device.
Guess APL has the last laugh.
I think SIMD support is useful too but most of the examples you gave are things which are handled on-device now using dedicated hardware. Video playback has been hardware accelerated for ages, ML acceleration is multiple hardware generations in, etc.
ML is similar – many people are running these apps now but, for example, many millions of them are running Apple ML models on Apple hardware with acceleration features so again, while SIMD is unquestionably useful, I'm not surprised that the average working programmer doesn't feel a huge need to dive in rather than using something like a library which will pick from multiple backends.
EDIT: there was a period in the late 90s/early 00s when dedicated DSPs were a useful but costly add-on to digital audio PC systems. SIMD moving into the CPU put an end to that.
A) The vast majority of programmers have zero idea about any of this. They want a key-value store and have no idea about the underlying implmentation. And that's okay. You can write an awful lot of useful software without understanding Big-O notation.
B) Pointer chasing has been dogshit on modern architectures for quite a while, and most people have no idea. Look at all the shit Rust gets for being annoying to implement linked lists which are a garbage data structure on modern CPUs.
C) Array-of-struct to struct-of-array type transformations are just starting to get some traction in languages. Entity component systems are common in games but haven't really moved outside of that arena. These are the kind of things that vectorization can go to town on but haven't yet moved to mainstream programming.
D) Most of this stuff is contained inside libraries anyway. So, only a few people really need to care about this.
When what we should have is libraries that provide high level relational/datalog style manipulation of in-memory datasets, and you describe what you want to do and the system decides how. Personal beef.
Think of the average "leetcode" question. It almost always boils down to an iterative pass over some sequence of data, manipulating in place. C-style strings or arrays of numbers, etc. If you tried to answer the question with "I'd use std::sort and then std::blah and so on, on some vectors" they'd show you the door because want you to show clever you are at managing for loop indices and off-by one problems and swapping data between two arrays in a nested loop, etc.
So we're actually gating people on this kind of thing. And imho it's doing us a disservice. The code is not as readable. And it doesn't necessarily perform well. It's often a 1988 C programmer's idea of what good code is that is used as the entry bar.
It was never universally true. I remember reading about compiler design in the 1970s; I think it was somewhere in Wirth where he pointed out that, for local symbolic lookups, it was more efficient to use a simple loop over an array because the average number of symbols was so low that the constant management overhead from any more advanced data structure was more expensive than a linear scan.
Competing efforts including std::simd or Highway are either unusable because they don't expose pshufb or more difficult to use than the tools they are abstracting away.
Can you help me understand why Highway might be more difficult to use? That would be very surprising if we are writing a largish application (few thousand lines of vector code) and have to rewrite our code for 3+ different sets of intrinsics.
I don't think there are necessarily problems with the design of Highway given its constraints and goals. I would just personally find it easier for most of the SIMD code I've written (this tends to be mostly parsers and serializers) to write several implementations targeting AVX2, AVX-512, NEON, and maybe SSE3.
I suppose there is a tradeoff between zero type safety + less typing, vs more BitCast verbosity but catching bugs such as arithmetic vs logical shifts, or zero extension when we wanted signed.
Would be interesting if you see a difference in compile time. On x86 much of the cost is likely parsing immintrin.h (*mmintrin are about 500 KiB on MSVC). Re-#including your own source code doesn't seem like it would be worse than parsing actually distinct source files.
Many of the Highway swizzle ops are fixed pattern; several use permute4x64 internally. Can you share the indices you'd like? If there is a use case, we are happy to add an op for that.
It is impressive you are willing (and able) to write and maintain code for 4 ISAs. Still, wouldn't it be nice to have also SVE for Graviton3, SVE2 on new Arm9 mobile/server CPUs, and perhaps RISC-V V in future?
I'm using 2031 and 3120 (indices and the rest of this comment in memory order assuming an eventual store, so that's imm8 = (2 << 0) | (0 << 2) | (3 << 4) | (1 << 6) etc.). This is pretty niche stuff! It's also not the only way to accomplish this thing, since I am starting with these two vectors of u64:
v1 = a a b b
v2 = c c d d
then as an unimportant intermediate step v3 = unpacklo(v1 and v2 in some order) = a c b d
and I ultimately want v4 = permute4x64(v3, 3120) = c d b a
v5 = permute4x64(v3, 2031) = b a c d
> It is impressive you are willing (and able) to write and maintain code for 4 ISAs.This is mostly hypothetical, right now I'm writing software targeting specific servers and only using AVX2 and AVX-512. I may get to port stuff to NEON soon, but I've only skimmed the docs and said to myself "looks like the same stuff without worrying about lanes," not written any code.
> Still, wouldn't it be nice to have also SVE for Graviton3, SVE2 on new Arm9 mobile/server CPUs, and perhaps RISC-V V in future?
I don't know much about variable-width vector extensions like SVE and SVE2 but my impression is that they are difficult to use for the sorts of things I usually write. These things are often a little bit like chess engines and a little bit less like linear algebra systems, so the implementation is designed around a specific width. This isn't really a rebuttal though, it just means I am signing up for even more work by writing code on a per-extension per-width basis instead of a per-extension basis if I ever move on to these targets, which I certainly should if they have large performance benefits compared to NEON.
v4 = permute4x64(v3, 3120) = d c b a
v5 = permute4x64(v3, 2031) = b a d c> These things are often a little bit like chess engines and a little bit less like linear algebra systems, so the implementation is designed around a specific width.
hm. For something like a bitboard or AES state, typically wider vectors mean you can do multiple independent instances at once. Likely that's already happening for your AVX-512 version? If you can define your problem in terms of a minimum block size of 128 bits, it should be feasible.
> I ever move on to these targets, which I certainly should if they have large performance benefits compared to NEON.
SVE is the only way to access wider vectors (currently 256 or 512 bits) on Arm.