edit: oh, and softfloat vs softfp vs hardfp vs vfp
edit: oh, and how they have two incompatible assembly language dialects that are mostly the same, but in non-trivial code, incompatible
Heh, having dealt with x86 for years, this is comparatively such a nothing burger. It's always a simple known fixed offset.
> softfloat vs softfp vs hardfp vs vfp
That's something not really unique to ARM per se. Any architecture with options for hardware FPU are going to practically need ABI specs for the soft and hard cases (you don't absolutely need anything more than hardfp since you can always emulate but it will be slow as shit) - and ARM is certainly not unique in having multiple hardware floating point implementations either.
> two incompatible assembly language dialects
Curious what you're referring to here - but I personally wouldn't consider assembly language dialects to be part of a CPU architecture.
I assume they're talking about ARM vs. Thumb?
- Every instruction being conditional (all instructions have a four-bit condition field, with one of the 16 possible modes being "always");
- The barrel shifter, which can be used on nearly every data processing instruction;
- The program counter being one of the general-purpose registers (and on the original ARM, the same register also containing the flags), so that any register move can alter the program flow;
- The load-multiple/store-multiple instructions, which can load or store up to 16 registers, plus incrementing or decrementing the base register; and since the program counter is one of these registers, it can restore several registers from the stack, update the stack pointer, change the program counter, and switch to Thumb mode (stored on the least significant bit of the program counter), all in a single instruction.
Being able to manipulate the program counter in the ARM processors directly is just being honest, simple and straightforward, rather then having it always done implicitly. Seems very intuitive to me now that I think about it.
There are four kinds of instructions that play hell with pipeline and OoO design:
- instructions that might cause traps, dependent on the values processed
- instructions that you don't know whether they will change the control flow
- instructions that you don't know where the control flow is going to go to
- instructions where you don't know how long they will take to execute
RISC-V, for example, bans the first category entirely other than load/store, and carefully separates the other three so any one instruction only had at most one of those problems.
ARM load multiple has all of those problems. At least you can examine the register mask at instruction decode time and know whether it will change the PC or not and tag the instruction in the pipeline as being a Jump or not. Imagine if there was a version that took the bitmap from a register instead of being hard-coded...
Load/store multiple don't increase performance much if at all on a CPU with an instruction cache and/or an instruction prefetch buffer. On an original 68000 or ARM without any cache, sure, a series of load or store instructions requires interleaving reading the opcodes with reading or writing the data, while load/store multiple eliminates the opcode reads. An instruction cache also eliminates them, leaving only the code size benefits. But load/store multiple is a perfect candidate for using a simple runtime function instead, at least if you have lightweight function call/return as RISC designs usually do.
The performance of movem.l in the MC68000 comes not from multiple load, but from multiple store, because the main memory access incurred a tremendous, extremely punitive penalty. This has not changed, even decades later, in systems with the fastest memory chips available: any writes to random access memory incur tremendous penalties.
PC being a general-purpose register was historically not uncommon. PDP-11 and VAX both did it and they were kinda popular at one time.
Load/store multiple was also fairly common with, for example, both 68000 and VAX having it. IBM 360 also, though using a register range rather than a bitmap -- a less general solution, but good enough, and much easier to make go fast.
Worse, iirc what looks like it should be the slot for the "never" condition actually does exactly the same thing as "always".
The A64 manual says 1111 on a Bcc etc disassembles as NV but does the same as AL.
The ARM7TDMI manual says 1111 is reserved and don't use it. I don't know the actual behaviour.
Aha. The Welsh&Knaggs "ARM Book" says before ARMv3 NV meant NV. In ARMv3 and ARMv4 NV is unpredictable. And in ARMv5 NV is used to encode "various additional instructions that can only be executed unconditionally".
ARM endianness is switchable at runtime.
They are? Care to point to a ratified RISC-V standard that lists said instruction sequences?
slli rd, rs1, {1,2,3}
add rd, rd, rs2 Fused into a load effective address
...this is so insane. Whoever thought that this is okay and good, has, in my opinion, severe psychological and psychiatric problems and would do well to seek professional help. If this gets "fused" into a lea, why just not implement a hex code for lea? I'm just completely at a loss as to how messed up that is.
You know what, I'd like to know what a person who thinks that this is okay looks and behaves like.
Thus they are replacing not two instructions but three. This addition was indicated because of the number of critical loops in legacy software that, against C recommendations, tries to "optimise" code by using "unsigned" for loop counters and array indexes instead of the natural types int, long, size_t, ptrdiff_t. This can indeed be an optimisation on amd64 and arm64, but it is a pessimisation on RISC-V, MIPS, Alpha, PowerPC.
One codebase that uses "unsigned" in this way is CoreMark and they explicitly prohibit fixing the variable type. But it's also common in SPEC and in much code optimised for x86 and ARM in general, where using "int" pessimises the code. If they used long, unsigned long, or the size_t or ptrdiff_t typedefs the code would run well everywhere.
While the .uw instructions were being added, it was very low cost to add the versions using all the bits of rs1 at the same time.
So, in the context of this discussion, having 32 bit operations sign extend the results to 64 bits is a weirdness. More ISAs do is than zero-extension, but the most common in the market zero-extend. Note that at the time RISC-V was designed arm64 was not yet announced, so only amd64 did zero-extension.
An unsourced table on Wikichip (which was added in 2019 and not updated since) barely counts as a proposal (in terms of it standing a good chance of becoming part of a ratified RISC-V standard).
No doubt many high performance implementations will choose to use fusion and no doubt they'll all go for different combinations, with different edges cases. Yes there likely will be significant overlap but it could become a bit of a nightmare for a compiler writers. A thorough standardized list of instruction pairs to fuse would definitely help here, but we don't have one.
That's not in the spec. It's just something high end implementations do, and low end ones don't.
So if fusion is a weirdness it's a nearly universal one.
The good thing about fusion is the program works fine if you don't do it, so low end minimal area CPUs such as microcontrollers can just not bother.
The poor ELF specification ends up quite tortured by this, IMO.
Arm64 is even worse! There is almost exactly the same instruction, but it also zeroes the low bits of the target address, so as you relocate code you also have to change the offset in the 2nd instruction even if the distance between the reference and the target stays the same.
RISC-V was originally going to implement a hypervisor mode which would only have worked with Xen-like hypervisors. Luckily we were able to head that off early and the actual hypervisor extension we got can run KVM efficiently.
The base is too base and their bitfield extension is weird.
I'm sure those that only need the base disagree. For the rest of us, there's G (IMAFD).
[0] See "Expanded Instruction-Length Encoding" in the user spec.
No one has done it, no one seems to be keen to be the first to do it, and even how the instruction length encodings work is not a ratified part of the spec -- it's just a proposal at the moment, even for the next step of 48 bit instructions.
There has been discussion of encodings better than the one proposed in the current spec, especially around instruction length encoding schemes that would make more opcode bits available in 80 bit instructions than in the scheme in the spec, so as to have a possibility of encoding 64 bit literals in an 80 bit instruction.
I've looked at wide RISC-V decoder design and the variable length is no problem at all out to at least decoding 32 bytes of code per cycle i.e. eight 32 bit opcodes or sixteen 16 bit opcodes, or somewhere between for a mix (average would usually be about 11-12 or so).
You just need 8 decoders that can decode any instruction, plus 8 decoders that only have to understand C instructions. The 16/32 decoders each need a 2:1 mux in front of them selecting either bytes 0..3 or 2..5 from a six byte window. They always output a real instruction. The C-only decoders will sometimes be told just to output a NOP instead [1]. Each decoder type needs a 1 bit input to tell it which option to take. Those inputs can be chained like carries in a simple adder, or they can be calculated in parallel like in a carry-lookahead adder. For an 8-16 wide decode you need a 2-deep network of LUT6 to do this (in FPGA terms .. also not very deep in SoC terms).
Note that this is a VERY wide machine. Possibly well beyond the point of usefulness given typical basic block lengths and what you can sensibly do in the OoO back end. x86 is currently doing 3-4 wide decode, and Apple M1 is doing 8 wide.
In short: no, it's not a problem.
[1] or not output an instruction at all. Outputting a NOP makes it easier to insert the decoded instructions into an output buffer. Then you need to filter out NOPs later -- which is needed anyway, as programs contain explicit NOPs, OoO machinery turns register MOVE instructions into NOPs by just updating the rename tables, etc.
+ 'char' being unsigned
+ handling of unaligned accesses (in early architecture versions a value is read from the aligned address and rotated, which is useless behaviour that falls out of the original implementation because of how it dealt with byte loads; subsequently it was at least made to fault, but it wasn't until I think v6 that unaligned accesses were made to Just Work)
+ the weak memory model
Sorry, its been a while, but i found the idea, to generate those instruction isles into the code quite weird.