Fast bitset decoding using Intel AVX-512
lemire.me
lemire.me
I believe it does 128 bits per instruction, but I'm still struggling with rust w/ asm.
Along my journeys, however, I found this repo https://github.com/WojciechMula/sse-popcount/ which has tons of competing simd implementations for both intel and arm.
If the rumours are true that AMD Zen 4 will have AVX-512, then it'll be yet another sad corporate failure of watching a small underdog competitor show you how you are supposed to be managing your own flagship products...
Can anyone with actual hardware experience comment?
[1] Have I got that wrong? Is it 25%
(edited to de-screw formatting)*
This brings a huge simplification to any loop which must process arrays of variable sizes and alignments.
With SSE or AVX you have to write very complex prologue and epilogue code for loops, to handle the sizes that are not a multiple of the register size and to cope with various alignments. A compiler may be not smart enough to add such code, so it may fail to vectorize a loop, while when the code is written by hand it is hard to avoid bugs.
With AVX-512 handling any size or alignment is trivial and it can also be done automatically by compilers when vectorizing loops.
The second feature is that in AVX-512 there are both gather load and scatter store instructions, which simplify the manual or automatic vectorizing of loops that must access some data structures that do not have the layout required by the SIMD instructions.
Also AVX-512 has a more complete and flexible set of instructions to do various permutations between the parts of the SIMD registers, which also allow the vectorization of some loops that would have been difficult to convert to SIMD instructions when using SSE and AVX.
The net effect of all these features is that many loops that were not worthwhile to convert to SIMD with SSE or AVX, because the resulting programs would have been very complex and would have contained a large number of instructions, can be converted to AVX-512 with little effort and the resulting program has much less instructions than in the SSE/AVX version.
For the loops that are easy to convert to SIMD, e.g. they access only aligned arrays whose size is a multiple of the SIMD register size and there is no need to access other kinds of data structures, e.g. arrays of structures with heterogeneous members, AVX-512 may have no advantage over AVX, except if the target CPU happens to have double throughput with 512-bit registers (true for a part of the server/workstation Intel CPUs).
(Treating it as completely opaque is, I believe, also why you can get away with things like inlining AVX2 ASM on an -mno-avx2 build, for example.)
[1] - https://news.ycombinator.com/item?id=31312175
We're each cells in a neutral meta network, triggering each other.
Still, modern processors have BMI2. For some practical applications, the PEXT instruction is pretty comparable, here’s an example: https://stackoverflow.com/a/72106877/126995