The article cites Hermann pushing IBM's RISC approach, but as I recall (I worked at Acorn at the time) the inherently RISC-like simplicity of the 6502 that Acorn was using was itself part of the inspiration.
The article cites Hermann pushing IBM's RISC approach, but as I recall (I worked at Acorn at the time) the inherently RISC-like simplicity of the 6502 that Acorn was using was itself part of the inspiration.
The register had a much longer set of articles on the history of Acorn and ARM. I was hoping for new information in this Ars Technia piece, but it was scant.
https://www.theregister.com/2012/05/03/unsung_heroes_of_tech...
- Fixed-width encoding: No. It varies from one to three bytes.
- Large register set: No, it's an accumulator architecture. I don't count the zero page because every program that ran on the computer had to use the same zero page. IIRC, user programs on some late-era 6502 OSes would have only a handful of those 256 bytes to work with, rest being claimed by the system software.
- Pipelining: There's some optimization with prefetching of instructions. It is not the multi-step pipeline of RISC, nor does it need to be.
- Load/store architecture: Plenty of instructions work directly on memory.
I never had the conversation with Wilson/Furber, but I'm assuming that CISC was essentially off the table in the first place. Both had zero CPU design experience, so to have any hope for success it had to be simple. The 6502 provided plenty of validation that simple could be fast, and no-one at Acorn had been impressed with next-gen 16/32-bit CISC designs anyway. No doubt other early RISC designs were an influence, but I don't believe there was any RISC vs CISC religion involved.
It's amusing how today's worldwide domination of the processor market by ARM is really based on design decisions that were born of necessity for a chip that was really targeted for desktop use. It had to be simple, hence the die was small and cheap. It had to use low cost plastic packaging, so it had to be low power to avoid overheating. Low cost and low power is what led to ARM's mobile success.
All I'm saying is that I heard at the time that the well-liked 6502 was a source of inspiration for the design (and someone else has noted Hermann also remembering something similar), notwithstanding that there was no doubt much borrowing from the contemporary (UC Berkeley, Stanford) RISC designs as well.
As a 6502 programmer the utility of zero page variables is pretty obvious and there is a parallel there to a large register bank.. Maybe the 6502 wasn't "RISC like" in sense of incorporating what would become RISC design principles (it was too primitive for that), but it certainly validated a KISS design approach, just as contemporary next-gen CISC designs did nothing in Acorn's eyes to sell that approach. Back at that time we were all just seeing/discovering what was possible, and I'd have to guess that the ARM instruction set was more pragmatic than dogmatic - a combination of what seemed desireable, appeared reasonable to implement, and performed well in simulation.
> I do not remember IBM featuring so prominently in our deliberations. Our starting point was a 16 bit version of the 6502.
There is an interesting interview with Hermann where he says that Acorn asked Intel for a special version of the 80286. I've always been puzzled by that given how weak the 286 was!
First I've heard of that. Did he give any details of what it was being considered for, or what special features were being asked for ?
It’s here I think:
Yeah, that doesn't sound quite right. There must have been some approach to Intel, but that particular complaint applies to the 8086 not the 286. There was certainly a big emphasis on memory bandwidth at Acorn. There's a video by Sophie Wilson (who I knew pre-Sophie - worked with her to reuse BBC BASIC floating point libary for ISO Pascal) describing how Acorn found that, across various process architectures being evaluated, performance was directly related to instruction fetch speed regardless of how that broke down in terms of clock speed, bus width or cycles per instruction.
Acorn had of course been experimenting with most of the next-gen processors at the time, and put out a few in "2nd processor" form before a bit later repackaging them as the Acorn Business Computer series. There was actually a prototype 286-based version of the business computer, but by then the ARM was already being developed.
I don't recall hearing anything about the 286 being evaluated by Acorn the time (I was there from 82-85). It's interesting that the 286 and Nat Semi 16032 both came out in '82, and Acorn chose to go with the 16032 as a 2nd processor (I still have one - did some work on PANOS) rather than the 286. Perhaps this is the time frame Hermann is talking about? The "If Intel hadn't rejected us, there'd have been no ARM" story it cute to make Intel look foolish, but it's not obvious that a 286 2nd processor would have curtailed ARM development anymore than the 16032 had.
It must have been interesting to work with Sophie. I remember looking at some disassembly of the BBC Basic 6502 code and being fascinated by how she got such good performance out of it.
Does your 16032 still work? Must be fairly rare if it does!
I was at the Uni at the time ARM was launched. Remember going to a University Computer Society meeting where I think (my memory is a bit vague at this point) Steve Furber presented the first Arm. No one could quite believe what they'd done.
I'm working on a series of articles for my blog on the early days of RISC. Comparing with the other RISC architectures it's really interesting how they differ. I think Sophie / Steve did a fantastic job of turning RISC into a practical and cost effective reality. I've just had a twitter conversation with Ben Finn (of Sibelius music software fame) who was saying how pleasant writing ARM assembly was!
source: https://community.arm.com/support-forums/f/architectures-and...
If you have to save e.g. 10 registers, using 40 bytes of instructions on function entry and another 40 bytes of instructions on function exit is a lot of wasted code space.
On Intel/AMD CPUs, the deprecation of the loading and storing of multiple registers (POPA and PUSHA) had less impact, because for pushing or popping a single register the length of an instruction is only 1 byte or 2 bytes, vs. 4 bytes on ARM.
To alleviate somewhat the removal of the instructions that could load and store up to all registers, in AArch64 double load and double store instructions have been added to the ISA, which can load or store a pair of distinct registers.
That's why the ARM1 had a microcode engine for LDM/STM despite having 'RISC' as the center letter in it's initialism.
Modern processors do way way more to try and predict where branches will go and minimize breaking the pipeline flow.
http://www.righto.com/2016/01/conditional-instructions-in-ar...
I think both can be debated. It took 4 bits out of each 32-bit instruction to specify those condition flags. That’s 12.5% of the bits in each instruction that could have been used for other stuff such as supporting more registers and larger immediate constants, both of which could speed up code.
It was an interesting idea, though. Simple to implement and useful at times.