More seriously, yup: if you read the architecture manual this instruction can take a long time. It's also a pain for a OoO (not superscalar!) cpu to keep track of all the dependencies.
It made complete sense back in the early 80s when the instruction set was designed, but like universal condition code dependent instruction execution it was dropped from ARM64 for good reasons.
As mentioned before these were microcoded anyways so it may end up "compiling" down to N loads which should track just as well.
If instead you had executed 16 separate LD instructions, all of them would have retired and update state by the time the last one gets the bus error.
It’s all doable of course, but it means your OOO algorithm and physical implementation need to handle a lot more state in flight.
The microcontroller-scale cores have a few extra bits in the control/status register that govern restart of an interrupted ldm/stm or IT (if-then, a form of hammock predication).
Irrelevant. "Undo all of that" is just a matter of using an older version of the register alias table. The retirement process has do to that for any and all faults anyway.
Without that instruction, your hardware design could have assumed that reverting state can always be done in one cycle. That makes the hardware implementation easier.
Any time there is a fault, it is detected by the retirement stage as it inspects the ROB. There are already hundreds of state changes that must be rolled back: all of the succeeding entries in the ROB. So the work performed by the recovery mechanism isn't changed by the amount of clobbered state because it is completely clobbered! The ROB must rewind it all.
The register file doesn't need any additional ports, because the fault recovery mechanism doesn't touch the physical registers. It only updates the register alias table (RAT) to point to the appropriate registers. But that machine must be capable of restoring 100% of the register pointers very quickly no matter what. Whether the faulting instruction touched one, two, or 16 registers doesn't matter. The 100+ updates following it are enough state that the answer must be: "Everything".
At that scale, those mechanisms aren't micro-coded, they are state-machined. You just have to restore the state machine.
https://gab.wallawalla.edu/~curt.nelson/cptr380/textbook/adv... says it takes 2 + N cycles, plus 2 if the PC is in the list.
Storing them takes 1 + N cycles, according to that PDF.
(Also, this kind of instruction isn’t a novel idea. 68k had MOVEM, for example (http://68k.hax.com/MOVEM))
(The special case I accounted for is because there's a bubble in the instruction pipeline if the next instruction uses the last register)
If, as the article claims, pop is truly aliased to ldm, then I'd expect it to be very fast.
Very much so, either due to memory latency or just inherent complexity. For example, a single division instruction can easily take the same time as ten adds.