LDM: My Favorite ARM Instruction
keleshev.com
keleshev.com
As an example, consider page fault handling for a situation where part of access is valid and other part is not, especially for store operation. Or out-of-order execution. Or, for whatever sake, "load multiple" from PC (which is exposed to programmer in quite peculiar way) - you have to have it.
ARM is full of these quirks. I can see why they did that back in the day, but today or even 30 years ago these quirks are not a good decisions.
It has been replaced with LDP (load pair, i.e load two registers) which is not nearly as nice for software developers but made the life of the ARM designers much easier. For comparison:
https://elixir.bootlin.com/linux/v4.14.52/source/arch/arm64/...
https://elixir.bootlin.com/linux/v4.14.52/source/arch/arm/ke...
BUT, I don't think we will see real benefits of this unless we remove aarch32 support from ARMv8.
• Support AArch32 at EL0-EL3 (everywhere)
• Support AArch32 at EL0 (user mode)
• No AArch32 support
The most fielded cores fall into the first category. The newer cores - A76 and it’s contemporaries IIRC - fall into the second category. Only it’s recently disclosed data center cores fall into the third category.
As an example, consider page fault handling for a situation where part of access is valid and other part is not, especially for store operation.
You do a check for the full range before the access, because for best performance you would want to pull in as much as possible in one go anyway.
Or out-of-order execution.
Everything gets turned into uops, this instruction can emit more than one.
"load multiple" from PC (which is exposed to programmer in quite peculiar way)
What's hard about that?
Cache coherence across cores is easy. The number of corner cases is much smaller.
The number of corner cases of old architectures (MIPS, or SPARC) is staggering. I designed a MIPS core prototype and I was done in single week with all arithmetic and memory access commands and spent three weeks designing, implementing and debugging branch handling, due to branch delay slot "feature".
>You do a check for the full range before the access, because for best performance you would want to pull in as much as possible in one go anyway.
You cannot do that - pull as much as possible in one go, - for case page boundary is crossed. But the very fact that you have to check before issuing operation make things more complex than they can be.
>>"load multiple" from PC (which is exposed to programmer in quite peculiar way)
>What's hard about that?
State handling. Can this instruction be interrupted?
>>Or out-of-order execution.
>Everything gets turned into uops, this instruction can emit more than one.
There are more than "everything gets turned into uops" out-of-order execution engines - scoreboarding, for example, provides most of benefits for much less energy cost.
I think PowerPC's EIEIO wins on cute.
The OS has to handle this case anyway for misaligned loads and stores.
Are load and store (multiple or misaligned) atomic? If so, then hardware has to check things in advance and trap. If not, then there is a possibility for some part of state to be lost.
Misaligned loads and stores that cross page boundaries are handled by the OS and hardware together. The OS must ensure that the post-condition of the page-fault handler leaves both pages on either side of the boundary in an accessible state.
What you describe above makes hardware design harder (one has to split unaligned access into two (unnecessary) microoperations which preclude or make harder to use techniques simpler than out-of-order execution engine with content-addressable memory). It also makes software handling harder.
One can designate unaligned access as "undefined behavior" and leave it to compiler writers to make things right.
Lots of hardware architectures have tried to levy this requirement on the software. Every one has subsequently walked it back and provided unaligned access support in hardware. Even RISC-V, which is as orthodox a RISC as it gets, supports unaligned access in hardware.
So that's got to be the baseline for deciding how much additional work must be performed when supporting something new: Any memory instruction may touch two pages, but the vast majority of them won't.
Thus, I cannot agree with you on "every hardware architecture walked back".
https://en.wikipedia.org/wiki/Lexra - for an historical expose.
It contains "possibly via trapping" which I am not at all against of. What is interesting there that MIPS introduced BALIGN instruction which reminds me of first generation of Alpha ISA.
Alpha did not had byte or word (16-bit) memory access. These should be provided using 32-bit word access and some shuffling. They walked back on that for code size reasons and later Alpha ISAs had byte and 16-bit loads and stores. I do not have reference ready but I believe they were aligned accesses.
I could find some reasons though here https://stackoverflow.com/questions/35742570/why-is-the-loop...
People didn't get automatic updates from fast always-on connections. Lots of people had slow dial-up connections. For some people, even that wasn't an option, so an update would require physical media.
The main reason old instruction become slow is that the hardware implementation is replaced by a microcoded version.
There quite possibly been a point in time where certain MMX instructions run faster on some AMD CPUs and required less uOps to achieve.
https://agner.org/optimize/ This is a pretty good source for uOps and instruction cycles.
FWIW LOOP isn't the worse thing in the world once you have dedicated silicon for it anyway generating micro ops in the instruction decode pathway. It's just a pretty cute run length encoding scheme for the instruction stream.
Source: https://stackoverflow.com/a/35743699
See also https://stackoverflow.com/questions/35742570/why-is-the-loop...
"(My opinion: Intel is probably still making it slow on purpose, and hasn't bothered to rewrite their microcode for it for a long time. Modern CPUs are probably too fast for anything using loop in a naive way to work correctly.)"
I believe they could be mixed so you could have it.tft - in which case you’d execute the first and third instructions if the condition was true, and the second if it was false.
Needless to say this is challenging instruction to deal with any kind of out of order processor, or any kind of branch predictor (logically a single instruction could produce multiple branches).
I know that in modern arm processors the performance is much worse than a short jump, so I’m guessing that they just give up when they see it instructions and do it in order.
Alas this instruction is gone in arm64. It’s obviously better for cpu designers, but I will miss it forever.
It gets really fun when you combine condition codes with comparisons, or with ALU instructions that set the processor status bits: you can do some quite intricate logic in very little space.
In the analysis he counts instruction set features like number of registers, number of addressing modes, number of memory accessed per instruction. He compares over a dozen architectures.
ARM comes out as the least RISCy RISC, but definitely on that side of the line, and x86 as the least CISCy CISC. (This was before amd64.)
That is the reduced means each instruction does less, rather than specifying the number of them.
> A RISC is a computer with a small, highly optimized set of instructions
but later:
> The term "reduced" in that phrase was intended to describe the fact that the amount of work any single instruction accomplishes is reduced—at most a single data memory cycle—compared to the "complex instructions" of CISC CPUs that may require dozens of data memory cycles in order to execute a single instruction
1. https://en.wikipedia.org/wiki/Reduced_instruction_set_comput...
> Cocke and his team reduced the size of the instruction set, eliminating certain instructions that were seldom used. "We knew we wanted a computer with a simple architecture and a set of simple instructions that could be executed in a single machine cycle—making the resulting machine significantly more efficient than possible with other, more complex computer designs," recalled Cocke in 1987.
[1]: https://www.ibm.com/ibm/history/ibm100/us/en/icons/risc/
LDM can access banked user registers. As a special case, this disables base register updates.
LDM can copy SPSR to CPSR, causing an exception return with register banking. This is a way to go from kernel code to user code.
Sort of like serially launching N async XHRs. The “serially” part doesn’t really matter, since the time of each task is dominated by the latency of the remote getting back to you; so you may as well think of it as launching them in parallel.
The only practical difference I can see is that LDM results in smaller code and so increases cache-coherency.
Other CPUs were not as good as the ARM at using the available DRAM bandwidth.
That's not the only alternative. Here's AVX2 code that copies 32 bytes with 2 instructions:
vmovdqu ymm0, ymmword ptr[rsi]
vmovdqu ymmword ptr[rdi], ymm0
Pretty sure ARM NEON can do something similar.Also, the gate count where these instructions really shine don't tend to have vector units.
More seriously, yup: if you read the architecture manual this instruction can take a long time. It's also a pain for a OoO (not superscalar!) cpu to keep track of all the dependencies.
It made complete sense back in the early 80s when the instruction set was designed, but like universal condition code dependent instruction execution it was dropped from ARM64 for good reasons.
As mentioned before these were microcoded anyways so it may end up "compiling" down to N loads which should track just as well.
If instead you had executed 16 separate LD instructions, all of them would have retired and update state by the time the last one gets the bus error.
It’s all doable of course, but it means your OOO algorithm and physical implementation need to handle a lot more state in flight.
The microcontroller-scale cores have a few extra bits in the control/status register that govern restart of an interrupted ldm/stm or IT (if-then, a form of hammock predication).
Irrelevant. "Undo all of that" is just a matter of using an older version of the register alias table. The retirement process has do to that for any and all faults anyway.
Without that instruction, your hardware design could have assumed that reverting state can always be done in one cycle. That makes the hardware implementation easier.
Any time there is a fault, it is detected by the retirement stage as it inspects the ROB. There are already hundreds of state changes that must be rolled back: all of the succeeding entries in the ROB. So the work performed by the recovery mechanism isn't changed by the amount of clobbered state because it is completely clobbered! The ROB must rewind it all.
The register file doesn't need any additional ports, because the fault recovery mechanism doesn't touch the physical registers. It only updates the register alias table (RAT) to point to the appropriate registers. But that machine must be capable of restoring 100% of the register pointers very quickly no matter what. Whether the faulting instruction touched one, two, or 16 registers doesn't matter. The 100+ updates following it are enough state that the answer must be: "Everything".
At that scale, those mechanisms aren't micro-coded, they are state-machined. You just have to restore the state machine.
https://gab.wallawalla.edu/~curt.nelson/cptr380/textbook/adv... says it takes 2 + N cycles, plus 2 if the PC is in the list.
Storing them takes 1 + N cycles, according to that PDF.
(Also, this kind of instruction isn’t a novel idea. 68k had MOVEM, for example (http://68k.hax.com/MOVEM))
(The special case I accounted for is because there's a bubble in the instruction pipeline if the next instruction uses the last register)
If, as the article claims, pop is truly aliased to ldm, then I'd expect it to be very fast.
Very much so, either due to memory latency or just inherent complexity. For example, a single division instruction can easily take the same time as ten adds.
MOVEM was the key to fast buffer->screen copies.
The earliest appearance I know of is in the IBM System/360, first released in 1964. The instruction is probably even older than that.
The ARM did change the instruction, arguably improving it. The IBM instruction loaded/stored multiple registers, but the registers were required to be consecutive. The ARM instruction has a register mask.
See IBM System/360 Principles of Operation, page 26: http://bitsavers.trailing-edge.com/pdf/ibm/360/princOps/A22-...
As you note, it's a place were ARM shows itself to be a hybrid RISC/CISC design, required because the ARM1 didn't have caches.
As for why it's awful, it's the amount of loads and stores that can be in flight associated with a single instruction when an exception is taken, and how the instruction is restarted afterwards. ARM-M cores make this a little better by exposing "I've executed this many loads so far in the LDM" in architectural state so it can be restarted without going back and reissuing loads (so it can be used in MMIO ranges), but the instruction really only shines in simple, in order designs. They're almost as bad to implement on modern cores as the load indirect instructions you see on some older CISC chips.
Additionally, LDM/STM really shined in either cache-less or cache-poor designs where instruction fetch is competing for memory bandwidth with the transfer itself. That doesn't really apply to these modern cores with fairly harvard looking memory access patterns. Therefore, getting rid of these instructions isn't the biggest deal in the world.
So to answer your question, they absolutely could have done that, but chose to use the transition to AArch64 to remove pariahs like LDM/STM from around their neck because they're more trouble than they're worth from a hardware perspective on modern OoO cores. The LDP/STP instructions are the bone they throw you to improve instruction density versus memory bandwidth to/from the registers, but they don't really want each instruction being responsible for more than a single memory transfer for core internal bookkeeping reasons.
> If an interrupt occurs during an CPU AHB burst read access to an end of SDRAM row, it may result in wrong data read from the next row if all the conditions below are met:
> • The SDRAM data bus is 16-bit or 8-bit wide. 32-bit SDRAM mode is not affected.
> • RBURST bit is reset in the FMC_SDCR1 register (read FIFO disabled).
> • An interrupt occurs while CPU is performing an AHB incrementing bursts read access of unspecified length (using LDM = Load Multiple instruction).
> • The address of the burst operation includes the end of an SDRAM row.
(https://www.st.com/resource/en/errata_sheet/dm00068628-stm32..., section 2.3.5)
Admittedly, this was a silicon bug in ST's memory controller. But it was a bug which was only triggered by LDM/STM instructions!
> If an interrupt occurs during an CPU AHB burst read access to one SDRAM internal bank followed by a second read to another SDRAM internal bank, it may result in wrong data read if all the conditions below are met:
> • SDRAM read FIFO enabled. RBURST bit is set in the FMC_SDCR1 register
> • An interrupt occurs while CPU is performing an AHB incrementing bursts read access of unspecified length (using LDM = Load Multiple instruction) to one SDRAM internal bank and followed by another CPU read access to another SDRAM internal bank.
That wasn't the actual design goal, though: Acorn was trying to get away with not shipping a DMA controller, which was quite expensive kit at the time they were shipping the Archimedes. Having an instruction pair that let you use registers as DMA buffer let them get similar performance and save lots of per-unit cost.