I'm not a fan of such vector libraries, AFAIK they all just inhibit auto vectorization.
At most, you can take advantage of 128-bit SIMD with a bunch of shuffles and extracts, whenever you are also working with scalar variables.
I did a small experiment comparing 6 possible implementations of the n-body [0] update loop: https://godbolt.org/z/sfehEfPGT
The implementations are:
* AOS: a simple scalar implementation with coordinates stored in an array of structs
* SOA: a simple scalar implementation with coordinates stored as a struct of arrays
* float3: uses a struct of three floats as a vector type
* float4: uses a struct of four floats as a vector type, ignores the last element
* vec4: like float4, but using a generic SIMD abstraction (so basically what glam does)
* floats3: attempts to do SOA with nice syntax. floats3 type has three arrays of floats and there are operations to extract and store a float3 type from a given index.
Since these abstractions are often used in games I'll start of looking at what the compiler produces when targeting Zen5 with -O3 -ffast-math:
* Zen5 O3 ffast-math:
AOS: gcc: 11119 ~SSE clang: 3688 AVX512, but quite messy
SOA: gcc: 1283 AVX512 clang: 1202 AVX512
float3: 11050 ~SSE clang: 10894 ~SSE
float4: gcc: 8646 ~SSE clang: 10815 ~SSE
vec4: gcc: 7913 ~SSE clang: 8196 ~SSE
floats3: gcc: 1284 AVX512 clang: 13351 ~SSE
The numbers next to the compilers are the cycle estimates from the llvm-mca model of Zen5 for processing 1024 elements.
AVX512 indicates whether the compiler was able to vectorize the loop with AVX512, and ~SSE means it could be partial vectorization with SSE.
Now let's also look at a different ISA, this time the RISC-V Vector extension:
* P670 2xVLEN O3 ffast-math:
AOS: gcc: 17445 clang: 3357 RVV
SOA: gcc: 3355 RVV clang: 3334 RVV
float3: gcc: 17445 clang: 17449
float4: gcc: 25668 RVV128 clang: 17470 RVV128
vec4: gcc: 45091 RVV128 clang: 23111 RVV128
floats3: gcc: 3333 RVV clang: 17446
This time the llvm-mca model for the SiFive-P670 was used, but I pretended it has 256-bit vectors instead of 128-bit ones, as the vector length is transparent to the codegen and this amplifies the effect I'd like to show.
RVV means it could be fully vectorized, while RVV128 is similar to ~SSE and means it could only partially take advantage of the lower 128-bit of the vector registers.
So if you are using such vector types to do computations in loops you are likely to end up preventing your compiler from optimizing it for modern hardware.
In general writing simple SOA scalar code seems to vectorize best, as long as you make sure the compiler isn't confused by aliasing.
But even the plain old AOS scalar code can be vectorized by modern clang, but not by gcc, and sadly also not the float3/float4 implementations, which should be very similar.
Modern ISAs like NEON/SVE/RVV have more complex vector load/stores that allow you to retrieve data more efficiently even from a traditionally bad data layout like AOS.
You can dress up the SOA code to make it a bit nicer, unfortunately my attempt with floats3 currently only works properly with gcc.
Below are the results when compiling without -ffast-math:
* Zen5 O3:
AOS: gcc: 11819 ~SSE clang: 10788 ~SSE
SOA: gcc: 4146 AVX512 clang: 13734 AVX512
float3: 11826 ~SSE clang: 11499 ~SSE
float4: gcc: 8662 ~SSE clang: 11810 ~SSE
vec4: gcc: 8575 ~SSE clang: 7451 ~SSE
floats3: gcc: 4148 AVX512 clang: 14367 ~SSE
* P670 2xVLEN O3:
AOS: gcc: 17464 RVV64 clang: 6122 RVV
SOA: gcc: 7140 RVV clang: 6118 RVV
float3: gcc: 17445 clang: 17464 RVV64
float4: gcc: 25665 RVV128 clang: 19184 RVV128
vec4: gcc: 17463 RVV128 clang: 56868 RVV128
floats3: gcc: 7140 RVV clang: 17444
Weirdly clang seems to be struggling with the SOA here, and overall vec4 looks like the best performance tradeoff for X86.
Still with proper SOA, and I bet you could coax clang into generating it as well, you can still get a 2x performance improvement.
Additionally, vec4 performs horribly with current compilers for VLA SIMD ISAs.
I'll try to experiment with some real world code, if I can find some that is bottle-necked by such types.
[0] https://benchmarksgame-team.pages.debian.net/benchmarksgame/...