How deep down the rabbit hole did you go with hardware optimization?
In an ideal world, would it be better to compile this on a processor more RISC-y?
In an ideal world, would it be better to compile this on a processor more RISC-y?
The focus is still on learning and pushing latency on regular hardware.