Don't think any language has standardized SIMD that's particularly nice; Highway is probably a quite nice library on C++, though I haven't used it enough to get comfortable.
The thing I use for my projects is Singeli[1], a DSL specifically made for SIMD stuff (though it's capable of generally sanely doing abstractions over types/operations/loops; it's just a fancy code generator). Obligatory disclaimer that I'm one of the two people working on its design. It's far from a nice experience starting from nothing, but it's pretty nice for what I do.
CBQN's the main place it's used, can click around its source: https://github.com/dzaima/CBQN/tree/develop/src/singeli/src
Its goal isn't necessarily to unify architectures, but rather make it as easy as possible to make abstractions that do; as such its built-in includes for x86 don't have arbitrary shuffling, but do provide a sane interface over the cases that are supported in a single instruction (not including constant creation/loading), and those can run on NEON unchanged (assuming they're ran on 128-bit vectors, of course, as NEON doesn't support larger ones); and, with Singeli just generating C/C++ currently, you can just map in __builtin_shufflevector if desired. e.g. here's your AVX2 `pre`:
include 'skin/c' # defines infix a+b & a*b etc to run __add/__mul/... (yes, those aren't here by default, and you can define custom infix/prefix ops)
include 'arch/c' # defines __add & __mul to do C ops
include 'arch/iintrinsic/basic' # not necessary for a shuffle, but provides basic x86 arith ops
include 'arch/iintrinsic/select' # x86 shuffles; there's similar 'arch/neon_intrin/basic' & 'arch/neon_intrin/select' for NEON
fn pre(inp: [32]i8, out: *[32]i8) : void = {
store{out, 0, vec_shuffle{16, inp, merge{ # 16 specifies to repeat per 16-elt lane
range{8}*2+1, # lower half: 8 Y components; compile-time index calculations
range{4}*4, range{4}*4+2 # upper half: (4 * U), (4 * V).
}}}
}
As a more fancy thing, I've got this working (via bodging together the definitions in CBQN with some sugar to make this pretty; not including all that boilerplate here), compilable to SSE2/AVX2/NEON producing a 4x unrolled core loop, plus tail handling (via reading past the end and doing a load-blend-store if necessary because that's what CBQN's fine with; could easily define a fancy_loop such that it does a scalar tail though). (also can be compiled to RVV via currently-unpublished mappings; no need to unroll for RVV; can choose to do either a stripmined loop or one with a separate tail):
fn sigmoid{E}(r:*E, x:*E, n:ux) : void = {
def V = arch_preferred_vector{E}
@fancy_loop{V,4}(r in tup{'dst',r}, x, M in 'mask' over n) {
# this loop body is generated 3 times for x86 & ARM - with x being a 4-elt tuple (core unrolled loop); a 1-elt tuple and no masking; a 1-elt tuple and masking
if (any_hom{M, ...(x!=x)}) {
emit{void, 'abort'}
}
r{x / __sqrt{1 + x*x}}
}
# were it not for a bug in tuple loop var mutation in Singeli having undesired pervasion, this would be possible:
# @fancy_loop{V,4}(r, x, M in 'mask' over n) {
# if (...) ...
# r = x / __sqrt{1 + x*x}
# }
}
export{'sigmoid', sigmoid{f32}}
(e.g. generated C for AVX2:
https://godbolt.org/z/KTeGsazKP)
(I'm not actually particularly expecting interest in Singeli; I just like writing stuff :) )
[1]: https://github.com/mlochbaum/Singeli