Glacial – Microcoded RISC-V core designed for low FPGA resource utilization
github.com
github.com
Like, sure, it's not meant to be a fast implementation, but even just a "mask byte with 0x7C and set PC to that value times 8" instruction (which in an FPGA implementation is just rearranging the wires) could save 5-6 cycles per instruction.
Is it really "microcoded" when all you're doing is writing a RISC-V emulator that runs on what looks to be a fairly standard 8 bit CPU?
[1] https://github.com/brouhaha/glacial/blob/master/ucode/ucode....
I understand that there are some FPGAs now that essentially have RISC-V ALU hard blocks, so use of them might be a speed and area improvements.
I have an "Icicle" board based on this chip. The FPGA comes preconfigured to boot Linux from the included pre=programmed SD card, so you can just use it to run Linux if you don't care about FPGAs. Or you can replace or enhance the default FPGA programming.
Don't know, but the amount of "microcode" or emulation code required is itself a reasonable measure of an ISA's complexity. Doing an x86 that way would surely take tons more code.
In those days the microcode ROM and ALU etc was substantially faster than RAM (core). At some point SRAM became as fast as or faster than ROM and machines copied the microcode into SRAM on startup. Some machines such as the Burroughs 1700 series loaded different microcode into SRAM depending on whether you wanted to run FORTRAN or COBOL programs.
Then companies started allowing users to write their own custom instructions in microcode. See for example the VAX "Writeable Control Store" which on the 11/780 (as an option) gave users 1024 words (12 KB) for custom microcode and a microcode assembler and debugger. Some people even wrote compilers targeting this for languages such as Pascal (see for example https://apps.dtic.mil/dtic/tr/fulltext/u2/a089424.pdf)
The next step was to turn the SRAM into a cache, and make a slightly more user-friendly microcode the actual instruction set used by all compilers, and thus RISC was born.
Of course you are correct that specialised instructions to make instruction decoding easier are helpful in an ISA emulator. It would not surprise me to see RISC-V itself get an extension along those lines in the near future, to help M-mode software emulate unaligned loads and stores and other unimplemented instructions, but maybe also to help emulate other instruction sets.
In a sense, yes, it is indeed!
But when I think "microcode" then I think something like the 8086's horizontal microcode [1], where each line of microcode is wired directly to the various functional units and the microcode jumps and branches are (in some sense) determined based off (some of) the bits in the instruction register.
Characteristics of microcode include: wide instructions (20-40 bits) that perform multiple operations in parallel (e.g. ALU operation and register-register copy) and hardware dispatch to the appropriate microcode to handle each microcoded instruction via zero-overhead multiway branches.
I wouldn't call a 6502-like assembly language microcode, even if it implements an interpreter for a user-level instruction set, because it lacks the relevant characteristics of true CPU microcode.
[1] https://www.reenigne.org/blog/8086-microcode-disassembled/
From time to time, I have been tempted to design a RISC-V implementation out of discrete 74xx components. Sure, there are plenty of projects out there to build your own processor from scratch like that, but most of them aren't LLVM targets!
The 32-bit datapaths and need for so many registers makes it a bit daunting to approach directly. That approach would probably end up similar in scale to a MIPS implementation I once saw done like that. (Can't find the link, but it was about half a dozen A4-sized PCBs).
Retreating to an 8-bit microcoded approach and lifting all the registers and complexity into RAM and software is a very attractive idea. Might even fit on a single Eurocard. It's not like a discrete TTL RISC-V implementation would ever be a speed demon, either way.
If one cuts even more corners, that number could come down much further. For example, an adder isn't actually necessary and can be replaced with lookup tables and bit-twiddling, again at the cost of cycles and more microcode.
See https://hackaday.io/project/161251-1-square-inch-ttl-cpu -- while not RISC-V it demonstrates some of these principles in action, taken to the extreme.
That design consists of: one 4 bit counter, four 8-bit flip-flops, one quad OR gate, one dual 2-to-4 demux, and one 128 KB Flash ROM.
Including the flip-flops (and obviously excluding the Flash memory) that comes out to about 200 or so gates by my count, and the microcode/emulation program implements a fairly typical CISC 16 bit processor. It's not even all that inefficient, with under 100 cycles per instruction on average.
In my quick searches, I found that high speed FPGAs are around $10k+, and have much more fabric than I really need or want.
I haven't synthesized it though, so I can't say for sure.
I feel like that slow evolution can make it appear "suddenly popular" despite being around for a few years. =)
It's pretty new.
So why do we need RISC-V? Is it another case of NIHS?
1. RISC-V is completely unencumbered from an IP perspective. There is no possibility of a rightsholder reasserting rights on IP they had previously released (like what happened with MIPS in 2019).
2. RISC-V is legacy-free. It's an extremely "clean" design, free of weird quirks like the MIPS branch delay slot or SPARC register windows.
3. There are subsets of the RISC-V architecture defined for different sizes of systems, e.g. 32/64 bit versions, an embedded subset with fewer registers, etc. They all share an instruction set and a general architecture, and most compilers can target any subset. Some of the smaller subsets are well within the realm of what a single student can be taught to implement within a semester.
4. Numerous real implementations of RISC-V exist -- both as hardware and HDL -- are being maintained, and the hardware is available on the open market.
What happened with MIPS in 2019?
Shouldn't that kind of thing be impossible, the way it's impossible to revoke the GPL licence on software?
https://www.hackster.io/news/wave-computing-closes-its-mips-...
Astonishing.
memory latency bottleneck and the resulting topology challenges
how can an isa solve those issues?Also saying it's like "a mediocre 90s design" is pure bias. It's a nice modern design.
I think MIPS had a similar issue but eventually fixed it. Maybe RISCV can do similar.
Yeah, that's RISC.
Instructions that can trap -- but almost never do unless you have a program bug -- cause a large complication in pipelines, and especially in OoO implementations. Even a single-issue pipeline can run faster and be smaller without conditionally-trapping instructions, and as soon as you have even 2-wide execution it's just much better in every way to use explicit checks that use the same conditional branching facilities as the rest of the code.
The code size penalty is very minor in practice.
64 bit RISC-V code density is far better than 64 bit ARM code density, which is similar to AMD64 (i.e. quite a bit bigger than i386)
Other than a few extra instructions to zero-extend 32 bit values to 64 bits at times, 64 bit RISC-V and 32 bit RISC-V code are identical in size.
High performance general computing these days means 64 bit, and RISC-V has by far the highest code density of any 64 bit ISA.
How open is OpenSPARC? Are there patent concerns?
RISC-V isn't aiming to revolutionise CPU architecture with a radical new design, it's aiming to offer a Free and Open, patent-unencumbered, fairly conventional RISC ISA. They're quite open about their emphasis on openness. [0]
For a project that aims to turn CPU design on its head, there's the Mill processor, although it's broadly thought to be vaporware.
Or would be.
The fact is RISC-V is a distinct improvement on early 90s designs such as MIPS III and has also learned lessons from Alpha, PowerPC, Itanium, and AMD64.
In many ways RISC-V and Aarch64 (which were being designed in parallel unknown to each other) learned the same lessons from those earlier ISAs, though they made several trade-offs differently.