An Open-Source FPGA-Optimized Out-of-Order RISC-V Soft Processor (2019) [pdf]
rsg.ci.i.u-tokyo.ac.jp
rsg.ci.i.u-tokyo.ac.jp
http://www.rsg.ci.i.u-tokyo.ac.jp/members/shioya/pdfs/Mashim...
Typically, I think of an FPGA as something used to accelerate specialized operations. But sometimes, in the middle of one of these specialized operations, you might want to do something more general, like run a network stack, without returning to the CPU. A soft processor like this allows you to run an ordinary network stack (with ordinary code) inside the FPGA.
Is that right?
I thought one of the things people used FPGAs for was accelerating network stacks, so I don't quite know why you'd want to use a soft processor for that. But it does make sense that you'd want to be able to run ordinary code in an FPGA (as part of a larger FPGA operation that is not ordinary code).
EDIT: Also, I don't understand this statement: "for example, one main compute kernel, which is too complex to deploy on dedicated hardware, is run by specialized soft processors". What do the authors mean "too complex to deploy on dedicated hardware"?
You can think of FPGA as an ASIC and the softcore to control this ASIC. The hot data path and heavy processing is done in the FPGA and processing options for the ASIC can be done using the softcore firmware.
If you go looking for a microcontroller for your project, you have to choose among what is available. Maybe this microcontroller has two hardware UART interfaces and 1 SPI interface. If I don't need any UART but instead need a CANBUS interface that microcontroller won't work for me. Sure I can bitbang the protocol on GPIO ports but that uses up a lot of the limited processing power on the microcontroller... Usually that means a more expensive microcontroller.
There is a threshold that you can cross where a small FPGA is cheaper than a microcontroller that has enough pins and processing power for your application. This does come with an additional upfront design cost of also writing (but more often integrating) the soft cores but sometimes that makes sense.
Sometimes peripherals just don't exist at the price point you need. Try and find a microcontroller that has a MMIO controller for under $5 (I probably couldn't do it at under $10 but I haven't gone looking recently), it's rare they're needed but sometimes a design requires one.
There has also been a lot of recent interest in doing formal verification of hardware logic. A lot of the microprocessors and even that full CPU in whatever device you're reading this on has a lot of undocumented black boxes and undefined behaviours both of which prevent that verification from being meaningful.
A high end FPGA already has hardened CPU blocks, USB and PCIE interfaces, and lots of other things built in. Then it has a large area of generic reconfigurable logic that you can customize to do whatever you want.
This reconfigurable logic will not be as fast as a custom chip but it is still far faster than software and can be used to implement your own CPU (assuming you don't want to use the hardened CPU or got an FPGA without them)
It costs time and effort to translate a software function to HDL/FPGA, so it's not always worth doing. For example you could do TCP/IP in hardware, but unless you have particular performance requirements (say HFT) you're probably better off with a soft core and a tested software TCP/IP stack.
Also each feature translated to hardware takes up space in the fabric. When you crunch numbers on an FPGA, it's ideal if you can lay out the entire sequence of operations as one big pipeline, so you can keep throughput as high as possible. Sufficiently long or complex sequences may not translate efficiently to FPGA.
It will obviously be much lower than the IPC of an actual high performance CPU (modern x86-64), but how big is the difference? And how does it compare to typical mobile processors?
[1] https://people.eecs.berkeley.edu/~krste/papers/SonicBOOM-CAR...
I'm curious if this will work on Lattice ECP5- I'm not really sure if Synplify supports system verilog to the same degree as the Xilinx tools. ECP5 is interesting because it's a $10 FPGA..
On the developer tooling side, there is https://github.com/google/verible for linting, code formatting and code indexing.
On the actual compilation side with there are https://github.com/alainmarcel/Surelog and https://github.com/alainmarcel/UHDM which are then being coupled with open source tools like Yosys to allow targeting Xilinx 7 Series and Lattice ECP5 FPGA ICs with fully open source flows using fully open source FPGA tools like symbiflow.github.io
I can't comment on how well it performs myself without testing it, but a quick skim of the paper reveals that it apparently performs well in comparison to its rival open-source out-of-order soft processor.
Comparing soft processors to other soft processors is fairly easy if they can both run on the same hardware, but comparing them to real silicon is inherently kind of meaningless, as they don't really compete at the moment, and the performance of the design in absolute terms will depend on the FPGA it is implemented on. Nonetheless, you could compare the raw numbers presented in the paper for curiosity's sake and see that indeed, it isn't very fast compared to modern silicon processors.
It's 32 bit so it's not desktop-class necessarily but this should blow a microcontroller out of the sky, for scale.
It's worth saying here, that a big ooo CPU is pipelines differently to a small/old risc processor - even amongst discussions about compiler optimizations people still use terminology like pipeline stall, when a modern CPU has a pipeline that handles fetching an instruction window, finding dependencies, doing register renaming and execution, that pipeline is not like an old IF->IF->EX->MEM->WB - it won't stall in the same way a pentium 5 did. The execution pipes themselves have a more familiar structure.
Many modern designs aren't fully bypassed and involve clusters of pipelines to manage this. IBM's recent Power chips and Apple's ARM cores are particularly known to do this.
Is this something that could be automatically optimized via simulation?
Is it something that could be made dynamic and shared across a bunch of schedulers so that cores could move between big/little during execution.
I'm sure it could be automatically optimized in theory, even without the solution being AI complete, but I don't think we have any idea how to do it right now.
No, not unless you're reflashing an FPGA. You'd have better luck sharing subcores for threadlets I think.
It's pretty easy to slap down pipelines. What is far harder is keeping them all fed and running without excessive stalling and bubbles whilst avoiding killing your max frequency and blowing through your area and power goals.
On the other hand, I think there is still space for a small in-order dual-issue FPGA RISC-V with 2 DMIPS / MHz performance.
Actually there is one:
https://opencores.org/projects/biriscv
1.9 DMIPS / MHz..
Still, that's a good bit of work they should be proud for putting out there and I hope other people build on it.
EDIT: Oh, wait, they don't mention register renaming. Hmm, well, I guess no speculating over multiple iterations of a loop then.
EDIT2: No, the PDF the link mentions a rename unit. http://www.rsg.ci.i.u-tokyo.ac.jp/members/shioya/pdfs/Mashim...
[1] https://stackoverflow.com/questions/53007782/what-benefits-d...
> There is an answer on Stack Overflow regarding Chisel benefits, it is just embarrassing [1].
I don't understand what is embarrassing about the answer ? As a software guy above answer make sense. Some problems you want to use C (or similar) for and some problems you want to use scripting language for, and then again sometimes the right tool is erlang, rust or go-lang ...
But like I said, that's my software guy perspective, so I am wondering what I missed?
So in the end, the answer doesn't provide any specific answer regarding SystemVerilog and Chisel. All I found is one mention of negotiating parameters, which Verilog doesn't do. I would have loved to hear a lot more about examples of what Chisel makes more convenient than SystemVerilog.
SiFive values programmability above everything and for that Chisel is pretty clearly an advantage.
The VexRiscv is written in SpinalHDL, which is a close relative of Chisel.
The advantage of SpinalHDL/Chisel is that it supports plug and play configurability that’s impossible with languages like SystemVerilog or VHDL.
You can read about it here: https://tomverbeure.github.io/rtl/2018/12/06/The-VexRiscV-CP...
That said: there must be at least 50 open source RISC-V cores out there, and only a small fraction is written in Chisel. I don’t see how the use of Chisel has held back RISC-V in any meaningful way.
TL;DW: Chisel is beautiful/fun to write in, with a definite productivity bonus, but has a pretty large learning curve and had a much greater verification cost, partly because it's an HLS (most have that problem) and also lack of any tooling. Both of those costs are gradually being reduced (though in my opinion, not enough to not make verification a PITA).
Chisel is NOT HLS at all. Chisel is one of many languages that generates HDL, that is, you describe code to built a circuit whereas in eg. Verilog you just describe the circuit (Verilog has a limited ability to do dynamically with generate statements).
An HLS is one that raise the abstraction. Almost all of them today allows you to write "lightly" annotated C[++] that gets translated into a circuit. Almost universally, the timing relationship isn't explicit at all.
Opinions:
All existing HDLs and HLSes are terrible and there's fertile ground for creating something to really advance the art. Personally I'm looking for something that is more productive that HDLs, but with more control than an HLS. Some promising examples: Handle C, Google's XLS (assuming promised development), and Silice.
I like chisel as a concept but the learning curve is too high: Scala is kind of a mess and when you add a custom DSL + lots of functional programming on top a different hardware design methodology, it becomes overwhelming to your typical ce/ee, who probably doesn’t have that exposure. I simply ran out of time to learn it properly.
It also a second class citizen when it comes to rtl tools. The verification engineers have to work with the generated verilog and it looks like a nightmare. There was some improvements recently but the engineers knowledgeable enough to work on this stuff seem pretty bandwidth constrained.
The biggest headwind to chisel is the breadth of knowledge required to work and improve it IMO.
I’m really hoping that pymtl gets a firrtl backend soon. Python has a pretty decent record for building DSLs.
Wouldn't we expect much higher numbers (more parallelism) considering the number of frontend/backend pipelines?
I ran the sysbench CPU test on each, and the M65 trounced the RasPi4, being over 3x as fast in single-core (and about 1.5x as fast in multicore, which makes sense with the T2400 being 2-core to the RasPi's 4-core).
So the RasPi4 (a cheap-class SoC) remains slower than a performance-class PC from 14 years prior. Moore's law certainly helped in peformance-per-watt and performance-per-dollar, but if pure perfomance is what you want... I don't think there's anything available to consumers outside of Apple's offerings.
The Getting Started Guide indicates it comes with a micro-SD with a bootable Linux image, but mostly goes on to describe console access. That said, it does recommend a GPU, but it's unclear whether it can boot to a graphical desktop out of the box.
This is all non-trivial and would make the design ~ twice as big and likely impact the cycle times in a rather sad way. But possible of course.
Anecdata: Full RocketChip (RV64GC) built for ECP5 85F comes in at 54k LUTs (out of 84k) and clocked at 14.8 MHz. However, the cycle time is related to the FPU which assumes retiming which yosys can't do. Without the FPU it's a more reasonable 50-60 MHz.
1: https://www.seeedstudio.com/SeeedStudio-GD32-RISC-V-Dev-Boar...
It's admittedly a niche use-case but it's an option for playing with RISC-V hardware...
The real question is: Is there anything hidden in the silicon? That's something you can only solve by owning your own fab - the US military approach.
This can be hidden in an FPGA - for example attached to the input pins or SERDES - without needing to know anything about the application.
(Article: https://spectrum.ieee.org/semiconductors/design/the-hunt-for...)
This does not mean that there is no way vendor can backdoor the chip you are getting, but it does narrow the possibilities significantly.