So, I can emulate this thing on my desktop PC at about 28 nSec/cell using some Pascal code[2]. I'm thinking that if I upgrade to a machine with AVX512 instructions, it might get radically faster. What I can't figure out is how this instruction actually works, and what gains I would actually get.
The Intel documentation on this instruction is as clear as mud. There's no example with all the bits shown and worked through, leading the reader to have no ledge on which to make some intellectual purchase towards understanding.
Questions:
If I were to fork over the cash for a machine with AVX512 instructions, how many of these instructions can actually execute/second?
Wouldn't moving a bunch of values to/from memory basically empty out all the caches and make this thing really slow anyway?
Does anyone have a worked out example with bits shown for all the sources and destinations before/after the instruction, so I can see what it does?
[1] https://esolangs.org/wiki/Bitgrid