RISC-V Stumbling Blocks
x86.lol
x86.lol
The best source of simple assembly examples I know of would be the tests from my project [3]. Though those could be a lot more helpful for a human reading them.
As for simulators, there are quite a few. The website lists a bunch [4]. I am not sure how many would support any kind of machine mode memory mapping though.
[0]: https://github.com/TheThirdOne/rars
[1]: https://github.com/riscv/riscv-teach
[3]: https://github.com/TheThirdOne/rars/tree/696ef9ed82285f0a73d...
Anyway, I got this example from Project Trellis to work on this board. It uses this "PICO RISC-V" (the entire CPU is in a single Verilog file):
https://github.com/SymbiFlow/prjtrellis/tree/master/examples...
I think it would be straightforward to build a Linux capable RISC-V system for this board using this free tool chain. The board has SDRAM, FTDI UART and SPI-flash. It does not have an SD slot, but maybe could be wired in with the expansion headers.
This ULX3S board also looks promissing, but is not available yet:
https://www.crowdsupply.com/radiona/ulx3s
What's very cool is that low end ECP5 FPGAs are around $5. I think RISC-V on an ECP5 with LCD touch interface might be the cheapest available option for a Linux microcontroller with a touch LCD interface. By available, I mean where you can easily buy the chips on mainstream distributors (otherwise it would for sure be some AllWinner chip).
https://github.com/litex-hub/linux-on-litex-vexriscv
also, step by step using a different core (rocket): https://insights.sei.cmu.edu/sei_blog/2019/10/how-to-build-a...
The rocket core mentions 102% of space used of an 85K LUT FPGA and having to deal with it. I wonder which core is using so much space? I have more experience with vendor soft cores, and they would use much less..
> If you want to do any OS work, you need a system that implements the Privileged ISA. Supporting the Privileged ISA is synonymous to being able to run UNIX-like systems, because it brings user-/supervisor-mode distinction and paging. Working on real hardware is generally preferred to working in emulators, because the code eventually has to run on metal anyway and emulators can be too forgiving for certain classes of problems. There is a plethora of RISC-V microcontrollers for a couple of dollars that all don’t support the Privileged ISA. At the time of writing, the only board you can buy that does support it is the $1000 HiFive Unleased, which was beyond my “I just want to play around with this” budget.
Suggests to me that it's at least part of why Stephen Marz [0] isn't trying to target a specific board with his project (which has been posted here a few times and I'm sorta following along). That pretty much kills most of the interest in RISC-V that I had in me.
I guess I'll wait a few more years and check back again later.
The HiFive Unleashed is interesting in that it has a lot more RAM than cheap ARM dev boards, for the same money though I could get a HoneyComb LX2K ARM system.
HiFive Unleashed is $1000 for the board with onboard RAM… but it does not have PCIe onboard, so if you want that, you could pay $2000 extra (!!) for the "HiFive Unleashed Expansion Board" that has a giant FPGA that has a PCIe root complex programmed on it.
HoneyComb LX2K is $750 for the board with 64GB eMMC + bring your own SODIMMs for RAM.
SolidRun also offers the MACCHIATObin for $339. 4-core A72 + 4GB full-size DIMM included. (The single-channel-ness of the RAM is not good for performance, but you can install a big 16GB DIMM no problem.) Both the LX2K and the mcbin have PCIe of course :)
Perhaps the biggest was that the comparison instruction should have had the condition being looked for in the instruction, and put the result in register 0 as a 0 or ~0 (= -1). Register renaming can then carry it through pipelines. Subsequent instructions may AND register 0 with other values, or subtract it from them, avoiding frequently mispredicted branches.
Are you saying that instead of (or in addition to) "branch if equal", "branch if less than", etc there should be a "set if equal", "set if less than", etc? In case you are, there is a "set if less than" instruction. The floating point extension does have a "set if equal" instruction despite there not being an integer equivalent.
Alternatively if you are complaining about 1 representing True rather than ~0, I don't think that matters much from a technical POV.
The extra step to get a comparison result into a register, and another to extend it to the whole register, is an inefficiency I would prefer to leave behind.
RISC-V was designed to be a good target for a C compiler, given that's what people almost always need. Compilers frequently see code line `int x = (a > b)` and SLT is perfect for that. Compilers don't do clever masking tricks to avoid branches, in general (there are a few hard-coded tricks like divide by signed constant).
That compilers fail to do clever masking tricks to avoid branches frequently results in 2x slower programs.
And your proposal - using the zero register as a flag register - breaks way more micro-optimizations (like the simple ones described in "The RISC-V Reader" book) than it allows
They actually do a lot of such tricks. Check out the GCC sources, for example tree-ssa-phiopt.c[1] and ifcvt.c[2]
[1] https://github.com/gcc-mirror/gcc/blob/master/gcc/tree-ssa-p...
[2] https://github.com/gcc-mirror/gcc/blob/master/gcc/ifcvt.c
a = (b&m)|(c&~m)
into cmov, but gcc won't.RISC-V doesn't have condition codes.
> The extra step to get a comparison result into a register, and another to extend it to the whole register, is an inefficiency I would prefer to leave behind.
Can you provide RISC-V assembly that demonstrates this problem? Common branching patterns in RISC-V are all a single instruction.
The core branching comparison instructions:
`a >= b` maps to `bge a, b, label`
`a = b` maps to `beq a, b, label`
`a != b` maps to `bne a, b, label`
`a < b` maps to `blt a, b, label`
Comparisons with 0 just use the x0 register and for <= and > just flip the order of the operands.
I boggle.
If I was going to change anything I'd add compare eq/ne with an immediate
Cmp rx ry rz # compare rx ry, write result to rz
Jmple rz, ra, imm # jump +- ra+imm if rz < 0
Which would avoid the use of a implicit condition register, as well as the short offsets that a combined compare+jump suffers from?
r31 = (rx<=ry)? ~0:0
(or <, etc.), result in r31 implicitly. Followed by pc = r31? pc + n : pc
i.e. a single conditional branch instruction, for all conditions; or rx &= r31
effectively a conditional move instruction, or rx -= r31
effectively a conditional increment.I.e., actually reduced, but more powerful. (Implicit destination followed by mov is free, because of register renaming.)
Decoding the RISC-V compare-and-branch instruction is a whole project.
There is no need for an implicit destination register, just do it explicitly and you get the same result but without any magic and with more control. Register renaming can do the same job with an implicit or explicit register destination, there is no incidence on register dependencies.
> r31 = (rx<=ry)? ~0:0
Reading your first post I know that you want to use 0 and ~0 instead of the traditional 0 and 1 as the result of condition statements. I don't think that's a particularly great idea because: - there are so many places where you need 1 instead of ~0; - translating 1 into ~0 only costs one "SUB" instruction. If this feature is required for a specific domain, it's sufficient to add an extension with this feature or - more drastically - to add a macro-operation fusion step to detect this pattern.
> Decoding the RISC-V compare-and-branch instruction is a whole project.
RISC-V encoding is made to keep decoding as simple as possible. Have you ever checked the RISC-V encoding? It looks a bit confusing at the beginning to a human but if you think about the implementation, it immediately becomes clear.
In RISC-V ISA, we can criticize: the absence of a standard SIMD extension (with a big abstract vector extension instead) or the fact that some extensions require 3 read ports on the register bank instead of 2, which can make the implementation more complex. But it's trade-offs
If the cost of negating before storing got to seem excessive, a single negate-and-store-word instruction could translate, but it would be used a very great deal less than negating the 1 to make a mask.
However simple it looks to decode the thicket of compare-and-branch instructions, decoding and executing exactly one conditional branch instruction that just checks for zero will be simpler and faster.
You still have the SET family to decode, but you need them anyway.
The "optional extensions" have all the useful instructions in them, but typically just a couple of the instructions in each extension would provide most of its value.
So, I don't get popcount because it's in the bitmanip extension along with 10M transistors' worth of clever stuff I don't need. Similarly, each of the other extension sets.
It would be much more useful to define slices through all the defined ones, so you get either the absolute minimum core instructions, or add just the coremost of each extension set, or a more comfortable subset of each, or half of each, spiraling outward as transistor budget grows.
The current system of fairly granular extensions, a useful base architecture, and profiles defining what extensions are guaranteed for particular application classes (eg "Unix server") is a nice compromise between extensibility and software/distro complexity.
If it is supposed to be simple, then great! Actually make it simple, instead of copying old complexities.
It isn't architecturally innovative, isn't meant to be, and isn't a good place to start if you're thinking about doing architectural innovation. A lot of people are talking up RISC-V like it'll solve every problem but there are many problems in both academia and the real world where it represents a solution.
That's either a lie, or the author is severely traumatized from coding on x86 and is suffering from the Stockholm syndrome. Poor author; I feel sorry for the bloke.
Code on UltraSPARC or MC68000 family of processors and then you will know what simple means. The instruction set of RISC-V is insane, absolutely insane.
Here is a good example[1]:
For general signed addition, three additional instructions after the addition are required, leveraging the observation that the sum should be less than one of the operands if and only if the other operand is negative.
add t0, t1, t2
slti t3, t2, 0
slt t4, t0, t1
bne t3, t4, overflow
...good luck programming this insane design with anything other than a high level language compiler.But your example shows no internal states exposure, it's just arithmetic to detect overflow. What are you trying to show ?
And the difficulty isn't really about exposing the internals, this isn't like a branch delay slot for instance. It's about making the relationship between instructions conceptually simple where there's always up to 2 inputs and up to one output. That doesn't help much in cases where the processor is doing something close to what the instructions are saying but it makes the simplest possible out of order processor, where the internals diverge wildly from the ISA, much easier to create.
EDIT: Also, the using C and compiling rather than doing assembly by hand is really a fundamental part of the whole RISC philosophy and RISC-V is just a more extreme form of that.
Contrast this with a leading CISC design, the VAX, which had plenty of addressing modes which practically all opcodes could use. A single opcode could expand into a whole stream of micro-operations to compute the addresses being loaded from, actually perform the named opcode, and then compute the address being stored to, and possibly more. Doing that in the presence of virtual memory, where an opcode might have to be stopped and restarted so the OS could take a page fault, was rather nontrivial.
But that's the only big issue, saying that the whole ISA is insane due to this specific point of design is totally overstated
> We did not include special instruction-set support for overflow checks on integer arithmetic operations in the base instruction set, as many overflow checks can be cheaply implemented using RISC-V branches.
So this boils down to "there's no platform overflow (V) flag", which is also a problem on other RISC systems, and is in any case rarely used in C code because it doesn't check for overflow.
add t0, t1, t2
in the occasional case where you want overflow as well add the other 3 instructions - if you're hand coding in assembler you probably know which you want, if you're using a compiler, it can map it's semantics where appropriate
That's been the RISC idea from day one. CISC ISAs were nice to the assembly programmer, RISC ISAs were nice to the C programmer and the hardware implementer. Do you know how to do integer multiplication on the MIPS? I'll tell you:
li $a0, 5
li $a1, 6
mult $a0, $a1
... do some other stuff here ...
mfhi $a2
mflo $v0
Why do you want to do other stuff? What's that mfhi and mflo stuff? Why is mult not three-operand? Because RISC! The MIPS model is so RISCy it exposes the fact its multiplication isn't single-cycle. It does this by making it somewhat asynchronous: You issue the multiplication opcode, and your code continues to run while the arithmetic is performed. If you try to get the result immediately, you'll stall the pipeline, so you were encouraged to schedule opcodes that don't rely on the result of the multiplication to run in that hazard.In a CISC chip, that would have been handled by microcode, and you would have gotten nice three-operand multiplication opcodes which could even do memory-to-memory operations, like on the VAX. In a RISC chip, the compiler was assumed to have all the intelligence the microcode used to have, so you got simpler-and-supposedly-faster hardware with really funky assembly language behavior.
https://devblogs.microsoft.com/oldnewthing/20180404-00/?p=98...
TL;DR: Compared to the MIPS, RISC-V is a model of rationality and clarity.
Note that Apollo Core isn't OSHW, so you won't be synthesizing that one.
Compared to the earlier CPU's like Z80's, or god help us DSP's and GPU's, this thing an absolute delight. In fact even compared to x86, it's a model of clarity. Have you ever looked at the Intel architecture manuals for x86? The description of how a "jmp" operates runs for pages and pages of pseudo code. I'll guess you will counter "but most people don't use the more complex features like gates", but the reality is I've used jumps via gates far more times than I've tested for overflow - and I've had to digest those pages and pages of pesudo code and the related data structures many times because they are so damned complex and unintuitive you keep forgetting the details.
I have written lots of MOS6502, MC68000 and UltraSPARC assembler, as I come from the demo and cracking scene. That's where I grew up.
I'm not a RISC-V fan-boi by any stretch of the imagination but I have to admit that they certainly cleaned up RISC which 40+ years later was in order for a thorough house cleaning. X86 is gloriously messy but RISC had become quite a mess as well.
That's kind of a nonsensical statement — "RISC" is not a particular ISA, RISC is a vague methodology/philosophy.
RISC-V is not officially related to any previous ISA, but people like to call it "MIPS in a trenchcoat". RISC-V got off the ground in academia, and MIPS was often used for teaching before. They are quite similar in some ways, but it's really independent.
https://people.eecs.berkeley.edu/~krste/papers/EECS-2016-1.p...
Is Power clean? Myself, I've never liked its special purpose Condition, Link and Count registers. I don't see the need for a Branch Processing Unit being an architectural division. The last time I checked, pushing a register value across that boundary was slow. It wasn't handled by a renamer.
Not suggesting it as a substitute for RV, just wondering why they saw nothing in it to learn from. It's not defunct like PA-RISC, Alpha, and (forgive me) SPARC, and quite some research has gone into it.
https://en.wikipedia.org/wiki/Sunway_(processor)
Is there an architectural feature in Power you think should have been considered by RISC-V for inclusion or omission?
My 2030 prediction is that Intel will have a simpler x86 where they discard a shit ton of legacy.
It is conventional to complain about condition codes, but I don't see what would be debilitating about just appending a set of condition codes to each register, so that renaming a register carries them along. Going from storing 64 bits to 70 bits each would't break the bank.
I believe that an element of your idea is done in modern x86 microarchitectures. I was having an argument on LLVMdev about a peephole optimization and partial register update stalls. Eventually one of the Intel compiler people popped in on the thread and settled the argument, saying:
Regarding the partial EFLAGS write, modern OOO processors independently rename the carry flag, et al, so this is no longer a problem.
The "Reduced" refers to reduced workload per instruction, not number or complexity of instructions for programmer. For example adding explicit load/store instructions.
Many modern RISC ISA's have more instructions than CISC processors.
This trade-off equation changed enormously the moment that L1 I-caches moved on-chip - which was when the RISC revolution flowered