RISC-V Origins and Architecture, Part 1
thechipletter.substack.com
thechipletter.substack.com
One of the worst things about technology was that so many people focused on "If we own the road, we can charge EVERYONE a toll and get rich!" and so, so many, business plans were built around gaining exclusive control over key facets of technology. The idea that there could be an instruction set architecture (ISA) and related silicon implementations that anyone could use for free, is an existential threat to those who live on the margin generated by the licensing of their ISAs.
If anything can derail ARM's growing influence, it will be this. The winners will be people like Apple with the resources to have their own chips fabbed with their own extensions and put into their own products which can run rings around people trying to use "generic" processors to do the same thing. Could signal a coming resurgence in demand for "computer architects" (old skool type :-))
Another protection is that the membership application for RISC-V International includes certification that RISC-V does not infringe on the applicant's own patents. With companies such as Intel, AMD, IBM, Marvell, Microchip, MIPS, nVIDIA, NXP, Qualcomm, Renesas, Samsung, Siemens, Sony, STM, Texas Instruments being members that covers a lot of ground.
The base ISA (up to RV32GC, RV64GC) very deliberately breaks no new ground. It's just smoother, with known potholes filled in.
Many of the optional extensions do break some new ground. The V extension has some unique features, though the basics are inspired by the Cray-1. RVWMO is based on experience gained with ARM and DEC Alpha memory ordering models, but fixing the problems -- and was actually designed by a world-wide group of academic and industry experts. Both of these are examples of something that simply would not be done as well by a small group of people inside a single company, working to a deadline.
Looking at the May 2011 spec (https://www2.eecs.berkeley.edu/Pubs/TechRpts/2011/EECS-2011-...) there is `RDNPC Rd` which simply loads next PC into Rd. So that's 4 bytes different to modern `AUIPC Rd,0`, and functionally identical (other than return address predictor SNAFU) to modern `JAL Rd,.+4` or 2011 `JAL .+4` (only `x1` supported).
So 32 bit PC-relative addressing needed `RDNPC;LUI;ADD` and then `JALR`/load/store vs modern just `AUIPC` before the main instruction.
If there was something before `RDNPC` then it much have been very short-lived!
There is allegedly an earlier RISC-V design here...
https://inst.eecs.berkeley.edu/~cs250/fa10/handouts/lab2-ris...
... but I get "You don't have permission to access this resource."
There is an updated 2011 version which appears to conform to the May 2011 spec. The instruction listing omits load/store byte/half and there is no relative addressing support at all (except `JAL 4`) but that is probably just simplification for a hardware design class.
https://inst.eecs.berkeley.edu/~cs250/fa11/handouts/lab2-ris...
Chris, do you have access to that fa10 version?
I think RDNPC is what I was thinking of, but I can't at all remember what was perceived as "risky" about it. I may be overblowing something from memory. AUIPC is better anyways.
- 2010 opcodes left justified like MIPS, 2011 right justified and a func3 like modern RISC-V .. but rd on the far left, not between func3 and opcode.
- register names in formats changed from xa, xb, xc (dst) to rs1, rs2, rd
- 2010 lw uses xa for dst, sw uses xa for src like MIPS, ARM etc. 2011 has rd always in the same place, with sw using rs2 not rd, like modern RISC-V.
- 2010 jalr has no offset, 2011 has 12 bit offset like now.
- 2010 the literals for ANDI, ORI, XORI are zero-extended, 2011 unspecified (so I think all sign-ext)
- 2010 j/jal have 27 bit offset (28 with shift?), 2011 25 bit field.
- 2010 ADD/SUB/shifts have *W suffixes. 2011 is is like modern RV32.
- 2010 has SRA / SRAI, 2011 only has logical. Possibly just a simplification for the lab.
http://csg.csail.mit.edu/6.S078/6_S078_2012_www/handouts/smi...
In what ways?
I thought that it was fairly well understood that LR/SC stuff was suboptimal for (for example) lock-free data structures and that the x86 primitives were quite a bit more suitable for those.
Taking a superficial look at the RVWMO stuff doesn't seem to contradict this.
What am I missing? How was this fixed?
The memory model specifies what happens when you DON'T use atomic ops.
Do you have a link we could read about this?
https://people.eecs.berkeley.edu/~krste/papers/EECS-2016-1.p...
I remember hearing about it back in college in the 2000s and thinking "wow, open source hardware! This is revolutionary!". I'm sure there is a good reason it didn't work/scale, but does anyone here know?
The document contains a short discussion of the pros and cons of most architectures which were still somewhat popular at the time RISC V was designed (the Power/PPC is omitted) - MIPS, SPARC, Alpha, ARMv7/v8, OpenRISC and x86.
Major points that count against SPARC in Waterman's view are:
- SPARC register windows have significant power and area cost and complicate superscalar implementations
- Branches use condition codes which result in more complexity for OOO implementations with register renaming
- Load/store pair instructions also complicate implementations with register renaming (ARMv8 has introduced those, too)
- Moves between the floating-point and integer register files must use the memory system as an intermediary
- The only atomic memory operation is fetch-and-store, which is insufficient to implement many wait-free data structures
The openly available SPARC implementations are either embedded-style cores (e.g. Gaisler's LEON3) or relatively specialized cores for highly parallel workloads such as web serving (T1 and T2). Regular Sun 4m/4u cores were not available in source form, but the SPARC license would have allowed alternative implementations. The license used to be available for a one-time fee of $99, but it seems SPARC can be used free and without royalties now, the $99 are an optional registration fee at SPARC International (https://sparc.org/faq/#q3).
[0] https://yarchive.net/comp/register_windows.html
[1] https://en.wikipedia.org/wiki/Delay_slot
[2] https://en.wikipedia.org/wiki/Register_allocation#Graph-colo...
ARM too. The RISC-V design was fairly well advanced by the May 2011 spec, and IMAFD (and E) were in essentially their current form in the May 2014 spec.
Aarch64 was announced between those two in October 2011, with the first 64 bit Apple iPhone 5s shipping in September 2013 and with ARM's own cores such as Samsung Galaxy S6 in April 2015.
The first commercial RISC-V chip (SiFive FE310) shipped in December 2016.
IBM System/360 is still going strong after 60 years, with extensions not replacement. Only the 24 bit memory addresses were short-sighted. Motorola 68000 and ARM made the same mistake, 15 and 20 years later. Let's not even mention x86 memory addressing.
RISC-V has that covered with 64 bit from the outset, and 128 bit addressing included in the planning.
The basic RISC idea of separating load/store from arithmetic has stood the test of time for 60 years too, since 1964's CDC6600, the first supercomputer, and the fastest one for the rest of that decade. The Cray 1 in 1975 is also very much a RISC computer on the scalar side (plus vectors very similar to modern RISC-V V extension).
Both of those machines had two instruction lengths just like current RISC-V: 15 and 30 bits on the CDC6600 and 16 and 32 bits on the Cray 1. RISC-V has had plans for future longer instructions, if necessary, from the outset.
IBM 360 is actually very similar, with the same 2 bits in every instruction specifying whether the instruction is 16, 32, or 48 bits long.
This makes very wide instruction decode simple -- 8 wide, 16 wide, even 32 wide if that is ever practical. Not quite as simple as doing the same thing on a completely fixed length ISA such as Aarch, but simple enough to not be a problem. Unlike x86.
What else might want to change? More registers? That is easily handled with the planned-for 48 bit or 64 bit instruction formats. The long instructions can address all 64 or 256 or 1024 registers, while the 32 bit instructions can use only 32 of them, and the 16 bit instruction can use (mostly) only 8 of them. Or x86-style REX prefixes could be used. That works, but is worse for wide instruction decode unless the prefix itself indicates the full instruction length.
https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=144...
That's perhaps the biggest thing about RISC-V: no one company can kill it by going out of business (DEC), by getting bored with it (Oracle with SPARC), or by deciding to replace it with something else (Intel with Itanium).
If the company you were buying RISC-V disappears, you can choose from ten others -- maybe 100 or 1000 others one day. Or, if you have the skills, you can design your own ... maybe in an FPGA if technology has moved along and you were using an old design.
Or if another ISA takes over, you can run a RISC-V emulator on whatever the new ISA is.
All of this without any questions or worries about legality, licenses, patents.
FPGAs and special purpose CPUs don't change much, as they have also been around for decades.
You need (in a simple core [1]) a very un-RISC read/modify/write sequencer with bus lock. That's not cheap in silicon. Or trap and emulate using LR (which could be NOP) and SC (which could be a plain store) and a few other instructions. That's not fast.
If you have that tiny embedded device with more than one core then you implement RV32IA properly, simple. And if you have a single core and no interrupts or DMA then implement RV32I and save a good bit of silicon.
[1] big designs don't implement AMOs in the CPU core, but out in the peripherals, RAM controller, or cache controller. As far as the CPU core is concerned an AMO is a memory read that sends extra data (a constant and an operation code) out with the memory address.
A similar comment was in a dialogue in the Expanse series. I forget which character said it, but the gist was the same. It's a good lesson.
At least the content is quite good.