Each processor family has its own supported SIMD instructions though. For a long time, SSE / AVX / AVX2 were the only game in town on x86. AVX-512 introduced in Skylake was primarily available only on server parts, and was a mixed bag, and didn't have support on AMD machines until just recently (with Genoa / Zen4).
A modern Mac laptop with an M-series chip will have ARM NEON vector instructions, just like on an iPhone or similar. There is a new-ish ARM vector instruction set called Scalable Vector Extensions (SVE) that Apple doesn't support, but Graviton 3 does. But generally improvements on your laptop often translate onto the server side, and you can definitely test correctness.
There are a few least-common denominator SIMD libraries out there, but realistically for many problems, you'll get the most mileage from writing some intrinsics yourself for your target platform.
Java's Panama project is working on their 6th preview (JEP-450) for the September 21 (LTS) release, but the released Java 20 previews JEP-426 features, including
- basic math
- transcendental functions (sin, cos...)
- load/store to foreign memory e.g., to interop with C
- compress/expand lanes, for parallel selection or filtering
- bit-wise ops (with count, reverse, compress, expand)
- all translated on Intel to SSE/AVX or on ARM to NEON/SVE
See e.g.,
- <https://jbaker.io/2022/06/09/vectors-in-java/>
- <https://openjdk.org/jeps/426>
If you go to the x86 cpu list it will tell you what features are enabled if you would optimize for a particular one https://docs.rs/target-features/latest/target_features/docs/...
On linux you can run lscpu to check what cpu features are detected.
If you go down the vectorization path, unless you are a crufty CPU DB vendor, I'd skip ahead to GPU for the same. Ecosystem is built to do the same but faster and easier. CPU makers are slowly changing to look more like GPUs, so just skip ahead...