Is the goal here to increase the decode bandwidth of Intel CPUs?
Is the goal to reduce demands on load-store units by increasing the number of registers?
Are they hoping to make it easier to port or JIT armv8 asm to Intel CPUs?
Is the goal here to increase the decode bandwidth of Intel CPUs?
Is the goal to reduce demands on load-store units by increasing the number of registers?
Are they hoping to make it easier to port or JIT armv8 asm to Intel CPUs?
Nevertheless, there are a few instructions inspired by Armv8, mainly PUSH2 and POP2, which correspond to the load register pair and store register pair of Aarch64.
On ARM that instruction is a pain in the ass when reading disassemblies, because the on-fail condition bits are just specified as a number from 0 to 15; the disassembler doesn’t bother to label which bits are specified, let alone what conditions they correspond to. Unfortunately it seems like Intel is doing the same thing in their assembly syntax, at least if I’m reading the document correctly.