BOOM Open Source RISC-V Core Runs on Amazon EC2 F1 Instances
cnx-software.com
cnx-software.com
https://twitter.com/firesimproject/status/103126763730350899...
The other advantage of AWS is scale-out once you want to get benchmark data and not wait forever (or to simulate warehouse scale computers). In the past, maintaining our own FPGA cluster was a nightmare.
But you can still get them for $20 or so on eBay.
They have a Spartan-6 LX150 (DVI version, revB) or Spartan-6 LX100 (DVI version, revC), which are really huge FPGAs. It's by far the best deal in town if you're interested in that kind of stuff.
Google for "github panologic-g2" to find up-to-date proceedings about the reverse engineering effort.
Right now, the FPGA is up and running and outputting an image over DVI.
Do you have any idea about how many FFs are needed for a BOOM CPU?
That's a lot of logic!
There are also a lot of new FPGA dev boards with good community support for folks looking at getting their feet wet, like the icebreaker (crowdsupply) and alchitry (Kickstarter).
The high-level idea is that you queue up your instructions in a table of sorts and execute them as they're ready (i.e., based on dependencies). You then re-order the instructions at the end of the pipeline.
The main challenges with this technique are: (1) handling exceptions and (2) speculative execution. For (1), you can use a re-order buffer. For (2), you would typically flush the instruction table noted above.
5 years ago it had already been in development for 10 years and there was supposed to be silicon shipping in 2-3 years. 2 years ago they were talking about an FPGA prototype that has yet to show up.
- A post-compiler compilation step that adapts a binary for a specific chip, replacing a complex hardware control module by a one-time software run, so the only thing actually on the chip is the routing;
- A very nice set of primitives that push the non-determinism into data, instead of flow control.
- Cheap interruptions, cheap memory access, and every other detail rethought to be cheap.
RISC-V has an advantage here: other than load/store and instruction fetch/decode, the instruction set is designed to not cause exceptions. For instance, integer division by zero is defined to return a specific value, instead of being an error.
I'm not sure I agree with that particular design decision, considering newer languages such as Rust that want to detect overflows. Then again, if/when such languages become more important, I guess RISC-V could add an extension featuring some kind of trapping overflow (or whatever mechanism is deemed best).
(Ref: search for the "Integer Computational Instructions" section in the RISC-V "user mode" reference. It should be section 2.4)
If you're going to be doing a lot of signed overflow checking (e.g. JavaScript) it may be better to store your numbers as N+0x8000000 (for 32 bit). Then each addition needs to subtract the offset afterwards, and each subtraction needs to add it (which works out to the same thing). The overflow check is then a single branch after the add and the adjust, for a total of three instructions instead of a total of four.
This seems bad, but if you're loading the values from memory and putting them back afterwards then it's pretty minor.
A later extension (such as J) could add instructions such as BADC/BADV/BSUC/BSUV a,b,label -- "Branch if ADd generates Carry" etc. These would fit into the existing instruction encoding and execution pipelines no problem.
Or, if this is a rare thing (as it should be in JavaScript) then making them "Trap if ADd generates oVerflow" would take a tiny amount of opcode space and allow the trap handler (which can be delegated to User space) to perhaps transparently fix up the result.
One option I just thought of is simply to have multi-output variants of the instruction which would take an extra flags register as an output operand.
Anyway that's really interesting to find out. It would complicate the optimized implementation of JS engines, which need it to promote int32 arithmetic to doubles where needed, which is what I was thinking of when I asked the question.
1. You typically won't get very far without also implementing register renaming, such that the reorder buffer is actually much larger than the register file.
2. You need to distinguish between instructions that are "complete" (result is available for internal bypass to another instruction that is being speculatively executed) versus "retired" (all instructions before it have retired, the instruction can no longer raise an exception, you know for certain that it was not executed because of incorrect speculation). Canonical, architecturally visible state is only written on retirement.
3. Without good branch prediction, it is kind of pointless. A reasonable rule-of-thumb-Kentucky-windage estimate is that 20-25% of the instructions that you encounter in execution order (not static code analysis) are branches. So if you have 30 or more instructions active in the pipe, your branch prediction needs to be pretty good or you waste too much work. (Note that some branches are just "flakey" (hard to predict) but the compiler can usually do a decent job of identifying those ahead of time. Thus, modern processors include a conditional-move instruction which allows the compiler to manage speculation and replace the flakey branch with an instruction that does not pollute the branch target buffer.
4. Watching what happens in the simulator at the end of a pointer chasing loop is good humor. A strongly predicted branch gets mispredicted on: while(p!=NULL) {}, and all sorts of page-fault, memory-exception, arithmetic-exception logic lights up all over the place until something asserts "OOPS_NEVER_MIND", at which point all that fun stuff collapses in a smoldering heap. Watching that never gets old.
The main advantage is higher IPC -> better performance since there is less idle time spent waiting for slow stuff like reading from disk.
Like reading from disk: no.
The max amount of out-of-order instructions is typically around 300 these days?
Reading from disk takes millions of cycles.
It’s about avoid waits when reading from L1 and L2 cache and DRAM.
They can do non-MMU SMT which is kinda cool too.
There is lots of useful things you can do without C. Also, it has not stalled:
The C extension is mentioned in that video as a priority but it sounds like some of the other priorities are being worked on (maybe there's funding for spectre mitigation work). Anyone who wants to use BOOM outside of research is going to need the C option because that's the standard for OSes. They said it's nice not to have to maintain a fork of Rocket Chip, the same will be true of not needing to recompile everything without C.
It is more active than I thought.