Arm AArch64 Adds Memcpy() Instructions
community.arm.com
community.arm.com
Classic ARM had LDM/STM which could load/store from a list of registeres. While very handy, it was a nightmare from a hardware POV. For example, it made error handling and rollback much much more complex in out-of-order implementations.
ARMv8 removed those in aarch64 and introduced LDP/STP which only handled two registers at a time (the P is for Pair, M for multiple). This made things much easier but it seems the performance hit was not negligible.
Now with v8.8 and v9.3 we get this, which looks much nicer than intels ancient string functions that have been around since 8086. But I am curious how it affects other aspects of the CPU, specially those with very long and wide pipelines.
AFAICS x86 "rep" prefixed instructions are defined so that they can in fact be interrupted without problems. The remaining count is kept in (e)cx, so just doing an iret into "rep stosb" etc. will continue its operation.
I think VIA's hash/aes instruction set extension also made use of the "rep" prefix and kept all encryption/hash state in the x86 register set, so that they could in fact hash large memory regions on a single opcode without hampering interrupts.
Usually. Cortex-M3 and M4 cores allow LDM/STM to be interrupted by default, and offer a flag to disable that (SCB->ACTLR.DISMCYCINT).
https://developer.arm.com/documentation/ddi0439/b/System-Con...
The 8086/8088 have a bug (one of very few!) where segment override prefixes were lost after an interrupted string instruction:
https://www.pcjs.org/documents/manuals/intel/8086/
I believe it was fixed in later versions.
This includes a pipeline to re-align from source to destination; partial-fill of the pipe line at the start and partial dump at the end; and page-sensitive fault and restart logic throughout.
Multiple versions of memcpy is suspicious to start with: is the compiler expected to know the alignment statically at code generation time? It might be from arbitrary pointers. Alignment is best determined at runtime. Each pass through the same memcpy code may have different aligment and so on.
Years ago I debugged the standard linux copy on a RISC machine. It has a dozen bugs related to this. I remember thinking at the time, this should all be resolved at runtime by microcode inside the processor. It's been years now, and we get this. Sigh. It's a step anyway.
Note that the Pi also has a not-fully-standard PCIe bus implementation, so that doesn't really help things either.
Ninja edit: Apparently its $699 USD but it's been scalped by third party sellers. A shame, I'd love to pick one up for some of the work I'm doing!
EDIT: would you be satisfied with a Jetson Nano combined with the hack mentioned by consp in another comment? (I would be)
On Apple M1, trying to map the BARs as normal memory directly causes an SError.
As a reminder, the Arm Base System Architecture spec mandates that PCIe BARs must be mappable as normal non-cacheable memory.
I think I learned about that many years ago, but I couldn't find anything about that recently when skimming the 11/780 user manual.
Sophisticated customers hacking the instruction sets of their machines goes back pretty much to the beginning. The earliest I personally know of is Prof Jack Dennis hacking MIT's PDP-1 to support timesharing, sometime in 1961. Commercial machines like the Burroughs B1700 had a WCS that was designed so various compiled languages could be optimized - e.g. a FORTRAN instruction set, a COBOL instruction set, etc https://en.wikipedia.org/wiki/Burroughs_B1700 It was also in the IBM360s because they had to emulate the IBM1401 software (although I don't know if the capability was open to users to modify).
Today of course you have the various optional features of the RISC-V ecosystem --- easy to load up on an FPGA.
Perhaps we should remember that we are in the very very early days of Computers, and we should expect continued modification / experimentation.
https://en.wikipedia.org/wiki/Control_store#Writable_stores
Here’s an example usage:
You mean from RISC to CISC, right?
FISC ("fast instruction set") was a term used for POWER/PowerPC to describe a philosophy that started very much with the RISC world, but considered the actual number of instructions to /not/ be a priority. Instructions were freely added when one instruction would take the place of several others, allowing higher code density and performance while staying more-or-less in line with the "core" RISC principles.
None of the RISC principles are widely held by ARM today -- this thread is an example of non-trivial memory operations, Thumb adds many additional instruction encodings of variable length, load/store multiple already had pretty arbitrary latency (not to mention things like division)... but ARM still feels more RISC-like than CISC-like. In my mind, the fundamental reason for this is that ARM feels like it's intended to be the target of a compiler, not the target of a programmer writing assembly code. And, of the many ways we've described instruction sets, in my mind FISC is the best fit for this philosophy.
Many (or all?) of the RISC pioneers have claimed that RISC was never about keeping the number of instructions low, but about the complexity of those instructions (uniform encoding, register-to-register operations, etc, as you list).
"Not a `reduced instruction set' but a `set of reduced instructions'" was the phrase I recall.
Most of RISC falls out of the ability to assume the presence of dedicated I caches. Once you have a pseudo Harvard arch and your I fetches don't fight with your D fetches for bandwidth in inner loops, most of the benefit of microcode is gone, and a simpler ISA that looks like vertical microcode makes a lot more sense. Why have a single instruction that can crank away for hundreds of cycles computing a polynomial like VAX did if you can just write it yourself with the same perf?
Worth noting that, AFAIK, both of these were removed on aarch64 (and aarch64-only cores do exist, notably Amazon's Graviton2)
RISC was originally about getting reasonably small cores, which can do what they are advertised, and nothing more, and µOp fusing was certainly outside of that scope.
Now, silicon is definitely cheaper, and both decoders, and other front-end smarts are completely microscopic in comparison to other parts of a modern SoC.
RISC has been definitively dead since Dennard scaling ran out; complex instructions are nothing new for ARM.
Except this is still not agreed upon on HN. Every single thread you did see more than half of the reply about RISC and RISC-V and how ARM v8 / POWER are no longer RISC hence RISC-V is going to win.
If anything, I got the vibe that they were more concerned about cost of implementation and "scaling it down" than about a future-looking, high-performance ISA. And I'd prefer an ISA designed for 2040s high end PCs rather than one for 2000s microcontrollers..
I Am Not A RISC Expert (IANARE), but I think it boils down to how reprogrammable each core is. My understanding is that each core has degrees of flexibility that can be used to easily hardware-accelerate workflows. As the other commenter mentioned, SIMD also works hand-in-hand with this technology to make a solid case for replacing x86 someday.
Here's a thought: the hype around ARM is crazy. In a decade or so, when x86 is being run out of datacenters en-masse, we're going to need to pick a new server-side architecture. ARM is utterly terrible at these kinds of workloads, at least in my experience (and I've owned a Raspberry Pi since Rev1). There's simply not enough gas in the ARM tank to get people where they're wanting it to go, whereas RISC-V has enough fundamental differences and advantages that it's just overtly better than it's contemporaries. The downside is that RISC-V is still in a heavy experimentation phase, and even the "physical" RISC-V boards that you can buy are really just glorified FPGAs running someone's virtual machine.
It's the software tooling cost.
There's nothing exceptional in the spec because it's trying to insert itself into the industry as a standard baseline, so staying small and simple is pretty intentional.
Its whole deal is that you can design a 2040's ISA or whatever you want and run 2015 Linux on it.
Everyone is jumping on it because no longer do they have to deal with a GCC/LLVM backend, and a long tail of other platform support: they can focus on the hardware, and put their instructions on RISC-V (with some set of standard extensions).
The other thing, though less impactful on the industry adoption, is that the simplicity allows hardware implementations (aka "microarchitectures") to replicate intricate out-of-order designs that we're used to in high-performance x86 (and ARM) cores, with a small fraction of the resources (https://boom-core.org/).
The real question in the high-performance space is: who will be the first to get an OoO RISC-V core onto one of TSMC's current process nodes (N7, N5, etc.)?
That seems like why everyone in the low-end space would be jumping on it (like WD for their storage controllers). But that's not really an advantage over the existing ARM & X86 ISAs in the mid to high-end space since they already have that software tooling built up.
But that also seems rather narrowly scoped to those who are willing to design & fab custom SoCs, which seems to need both ultra-low margins and ultra-high volumes to justify. Anyone going off-the-shelf already has things like the Cortex-M with complete software tooling out of the box. And anyone going high-margin can always just take ARM's more advanced licenses to start with a better baseline & better existing software ecosystem (ex, graviton2, Apple Silicon, Nvidia's Denver, Carmel & Grace, etc..)
It's really more the small vendor ISAs that I expect to become rarer with time, not the existing ISAs to go away.
Frankly, RISC-V feels perhaps a decade too late, but so does LLVM, and alternate history is such a rabbit hole so I won't go into it (but I suspect e.g. Apple would've had a less obvious choice for the M1, if RISC-V had been around for twice as long).
> But that also seems rather narrowly scoped to those who are willing to design & fab custom SoCs
I'm expecting most of the (larger) adopters are already periodically (re)designing and fabbing their own hybrid compute + specialized functionality - like the WD example you mention (Nvidia replacing its Falcon management cores being another).
I don't know for sure, but I also suspect some of them also want to avoid having Arm Ltd. (or potentially soon, Nvidia) in the loop, even if they could arrange to get their custom extensions in there.
You don't have to design it yourself. The foundaries are working towards free hard cells of RISC-V cores in most of their PDKs. It's hard for ARM to compete with free.
RISC-V has some big names promoting it heavily.
What open ISA would be a real competitor to RISC-V?
The medium article [1]. The Ars Technica article it refers [2]. The paper which the Ars Technica refers to[3].
[1] https://medium.com/macoclock/interesting-remarks-on-risc-vs-...
[2] https://archive.arstechnica.com/cpu/4q99/risc-cisc/rvc-5.htm...
[3] http://www.eng.ucy.ac.cy/theocharides/Courses/ECE656/beyond-...
Consider a CISC instruction set with all kinds of exceptional cases requiring specific registers. Humans writing assembler won't care much. When code was written in higher level languages and compilers did more advanced optimizations, instruction sets had to adapt to this style: A more regular instruction set, more registers, and simpler ways to use each register with each instruction. This was also part of the RISC movement.
Consider the 8086, eg with http://mlsite.net/8086/
* Temporary results belong in AX,so note in rows 9 and A how some common instructions have shorter encodings if you use AL/AX as a register.
* Counters belong in CX, so shift and rotation only work with CX. There is a specific JCXZ, jump if CX is zero, intended for loops.
* Memory is pointed at with BX,SI,DI, the mod r/m byte simply has no encoding for the other registers.
* There are instructions as XLAT or AAM that are almost impossible for a compiler to use.
* Multiplication and division have AX:DX as implicit register pair for one operand.
* Conditional jumps had a short range of +/- 128 bytes, jumping further required an unconditional jump.
Starting from the 80386 32 bit mode, a lot of this was cleaned up and made better accessible for compilers: EAX EBX ECX EDX ESI EDI were more or less interchangeable. Multiplication, shifting and memory access became possible with all these registers. Conditional jumps could reach the whole address space.
I heard people at the time describing the x86 instruction set as more RISC-like starting with the 80386.
You then have others, like the PDP-10 and S370 which are also regular but doesn't have these register-specific requirements that the Intel CPU's are stuck with.
AFAIK it was a quick and dirty stopgap processor to drop in the 8080-shaped hole until they could finish the iapx432. Intel wanted 8080 code to be almost auto-translatable to 8086 code and give their customers a way out of the 8bit 64K world. So they designed instructions and a memory layout to make this possible, at the cost of orthogonality.
Then IBM hacked together the quick and dirty PC based on the quick and dirty processor, and somehow one of the worst possible designs became the industry standard.
Thinking of it, the 80386 might be Intel coming to terms with the fact that everyone was stuck with this ugly design, and making the best of it. See also the 80186, a CPU incompatible with the PC. Maybe a sign Intel didn't believed in the future of the PC ?
It's the clones we have to thank for the situation we're in today. If Compaq hadn't done a viable clone and survived the lawsuit we'd probably be using something else (Amiga?). But they did and the rest is history, computing on IBM-PC compatible hardware became affordable and despite better alternatives (sometimes near equal cost) the PC won out.
The 80186 was already well in its design phase when the PC was developed. And the PC wasn't even what Intel thought a personal computer should look like; they were pushing the multibus based systems hard at the time with their iSBC line.
It is the same problem as POPCNT on Amd64, and practically everything on RISC-V. Checking some status flag at program start is OK for choosing computation kernels that will run for microseconds or longer, but for things that take only a few cycles anyway, at best, checking first makes them take much longer.
I imagine monkeypatching at startup, the way link relocations used to get patched in the days before we had ISAs that didn't support PIC. But that is miserable.
Developers that cares about portability are obviously going to stay far away from such things.
For computer support, generally you would pass a -mcpu= flag (or maybe -mattr=, but that might be a compiler internal flag, I forget). Obviously then that's not portable and has implications on the ABI. I didn't read the article but I suspect they might be in ARMv9.0, hopefully, otherwise "better luck next major revision."
For monkey patching, the Linux kernel already does this aggressively since it generally has permission to read the relevant machine specific registers (MSRs). Doesn't help userspace, but userspace can do something similar with hwcaps and ifuncs.
MSVC emits them to implement extension _popcount, but does not use that in its stdlib. Gcc, without a directive, expands __builtin_popcount to much slower code.
You can check for a "capability", but testing and branching before a POPCNT instruction adds latency and burns a precious branch prediction slot.
Most of the useful instructions on RISC-V are in optional extensions. Sometimes this is OK because you can put a whole loop behind a feature test. But some of these instructions would tend to be isolated and appear all over. That is the case for memcpy and memset, too, which often operate over very small blocks.
On x86, it's been there since the 8086, and can do cacheline-sized pieces at a time on the newer CPUs. This behaviour is detectable in certain edge-cases:
Linus Torvalds has some good comments on that here: https://www.realworldtech.com/forum/?threadid=196054&curpost...
https://www.realworldtech.com/forum/?threadid=196054&curpost...
https://www.realworldtech.com/forum/?threadid=196054&curpost...
It seems to me that rep move is so bad that you want to avoid it, but trying to write a fast generic memcpy results in so much bloat to handle edge cases that rep move remains competitive in the generic case.
https://en.wikipedia.org/wiki/Cell_(microprocessor)#Power_Pr...
More importantly today, using DMA to do large memcpy for non-latency-sensitive tasks allows cores to sleep more often, and it's a godsend for I/O intensive stuff like modern Java apps on Android which are full of giant bitmaps.
In ARM assembly syntax, the exclamation point in an addressing mode indicates writeback. Its difficult to be certain without seeing the architecture reference manual, but it would be consistent for instruction to be writing back all three of the source pointer, destination pointer, and length registers.
A memcpy is interruptible without replaying the entire instruction (say, because it hit a page that needed to be faulted-in by the operating system) if it wrote back a consistent view of all three registers prior to transferring control to an interrupt handler.
In those ARM cores I've programmed, the core has a few extra DMA channels which can be used for such things. However, using them from userspace has always seemed a bit of a hassle.
movs[b|w|d]: move data in bytes/words/doublewords aka memcpy
stos[b|w|d]: put a value in bytes/words/dwords aka memset
cmps[b|w|d]: compare values aka memcmp
scas[b|w|d]: scan for a value aka memchr
ins[b|w|maybe d]: read from IO port
outs[b|w|maybe d]: write to IO port.
lods[b|w|d] : read from memory was probably not meant to be combined with rep as it would just throw everything but the last byte away. I once saw a rep lodsb to do repeated reads from EGA video ram. The video card saw which bytes were touched and did something to them based on plane mask. This way touching 1 bit changed the color of a whole 4 bit pixel, speeding up things with a factor 4.
Then one day, someone found that rep movs was not the fastest way to copy data on an x86 and they all went out of vogue. I think rep stos recently came back as fastest memset, as it had a very specific CPU optimization applied.
Update: See https://stackoverflow.com/questions/33480999/how-can-the-rep...
If everything resides in CPU L1 cache, it hardly matters at all. Other than L1 cache pressure, of course.
Other example is copying DMA transferred data followed by immediately consuming said data. Also in this case, the copy often effectively just brings the data to the CPU cache and the consuming code reads from cache. Of course it does increase overall memory write bandwidth use when the cache line(s) are eventually evicted, but total performance degradation can be pretty minimal for anything that fits in L1.
http://webcache.googleusercontent.com/search?q=cache%3Ahttps...
But then the SPU did not have direct RAM access (only 256 kB of local S-RAM addressible from the CPU instructions), so DMA was something that followed naturally from the general design. Also not having any cache meant there were none of the usual cache coherency problems (though you may run into coherency problems during concurrent DMA to shared memory from multiple SPUs).
[edit] note also that the SPUs did not usually do any multitasking / multi-threading, which also simplified handling of DMA. Otherwise task switches would have to capture and restore the whole DMA unit's state (and also potentially all 256 kB of local storage as these cannot be paged).
[1] https://en.wikipedia.org/wiki/Cell_(microprocessor)
[2] https://arcb.csc.ncsu.edu/~mueller/cluster/ps3/SDK3.0/docs/a...
That's the key right there. Many embedded SoC's I've worked with have DMA engines, but they are all behind the MMU and only work with physical addresses. It makes using them for something like "accelerated memcpy" kind of cumbersome and usually not even worth it unless it's moving HUGE chunks of memory (to overcome the page table walk that you have to do first).
I think had they provided (very slow but) normal path for accessing memory it would have made the situation much more acceptable to nominal developers.
The difficulty in adopting the PS3 basically killed the idea of Many-Core as the future for high performance gaming architecture.