A FPGA friendly 32 bit RISC-V CPU implementation
github.com
github.com
https://tomverbeure.github.io/rtl/2018/12/06/The-VexRiscV-CP...
There is a little bit of industry usage, with the biggest user being SiFive - the founders come from the UC Berkeley group that developed Chisel.
Also, VexRiscv has some industry presence.
do ask sifive how much they regret that decision though <shrug>
When I've started doing FPGA consulting a few years ago I've started using Chisel, but eventually had to go back to SystemVerilog due to client reluctance.
I was dramatically more productive with Chisel than with SystemVerilog.
i didn't say that as a supposition - i know that they regret it. the chisel compiler has been an enormous (enormous) technical debt/burden for them because of how slow/resource intensive it is.
compared to what?
It's not like all the other EDA tools are really fast or not resource intensive. For smaller design firms I would think things like FireSim [1] would be a significant advantage.
I can imagine it is a disadvantage in other ways, i.e. it's only possible to do single phase positive edge synchronous design, which could be an impediment to high performance digital design.
But I wouldn't imagine that scala performance is particularly significant.
[1] https://fires.im
> But I wouldn't imagine that scala performance is particularly significant.
Imagine all you'd like - reality is much less imaginative though.
In the FPGA-based space (accelerators, RF/SDR, trading), hard disagree. There's plenty of boutique FPGA work going on in these.
In general: No, alternative HDLs don't see a lot of use, and I'd argue that we qualify as 'academia' since the ASICs are NIH funded and we tend to work with a lot of academic partners and on low-quantity R&D projects.
Having said that, every time we've deployed SpinalHDL for a commercial client they've been blown away by the results. The standard library, developer ergonomics, test capabilities, and little things like having clock domains as a part of the type system make development so much faster and less error prone that the NRE for doing it in verilog just doesn't make sense.
You get access to the entire Java and Scala ecosystem at elaboration and test time. We deploy ScalaCheck in our test harnesses to automatically generate test cases that can reduce inputs to identify edge cases. It's incredibly powerful.
If the design is low volume then minimizing NRE, which is mostly set by engineering hours, makes sense. At low volume, the semiconductor unit cost is mostly irrelevant so you can potentially use things like SpinalHDL to keep engineering hours down, and therefore potentially save NRE, and eat the higher unit cost which occur due to toolchain inefficiencies.
At high volume NRE is mostly irrelevant and unit cost is everything. So even if a tool or language is hard and annoying to use, if it gives a lower unit cost, you use it. Here you see things like an engineers hand tuning the layout of a single MUX to eek out a bit more of something good in the PPA space.
I only have experience with high volume HW and there something like Chisel or SpinalHDL wouldn't be considered as it just adds complexity to the flow, and makes it hard to do the optimizations that high volume enable us to consider, for a potential benefit we're not interested in.
Also SV has an absolutely enormous feature set, and often alternative HDLs miss out important parts like support for verification, coverage, formal verification, etc.
Getting away from SV is like getting away from JavaScript. The network effects are insane.
There was an attempt to make a kind of IR for RTL that would break the tie with SV (kind of like WASM has for JS)... I can't remember the name (LL..something?) but it seemed to have died.
Maybe this is similar I'm not sure: https://github.com/llvm/circt
Anyway the only really interesting new HDL I've seen is https://filamenthdl.com/
It's probably FIRRTL and CIRT is the compiler for that [1], [2].
[1] The specification for the FIRRTL language:
https://github.com/chipsalliance/firrtl-spec
[2] Original FIRRTL compiler that's now been replaced by CIRT:
It does seem to be part of CIRCT in some way though. Maybe it inspired FIRRTL or something. Slightly unclear relationship between the projects!
VexRiscv is aware of this unofficial standard, and asks for four 16x64 multiplies and adds the result together on the next cycle. This produces a much better fmax on FPGAs, but if you were targeting an ASIC, you would be better off asking for a 64-bit multiplier, or not trying for a single-cycle multiply.
Most modern CPUs tend to target a 3 cycle pipelined multiplication, which means 22-bit wide multipliers. Doing this on an FPGA each 22-bit multiplication would require two 18-bit multiplier blocks, for a total of six multipliers, wasting more resources.
-----
In general, "FPGA friendly" means optimizing your design to take advantage of the things which are cheap on FPGAs, like the 18-bit wide multipliers and the block ram. Such designs tend to run faster on FPGAs and use less resources, but it's wasteful to synthesize them to ASICs.
As opposed to, say, interfacing with an FPGA which could be totally different way to be "FPGA-friendly".
Performance on FPGA was better than most open-source RISC-V cores out there as of 2020. Rocket might have been better on silicon, but that's it. I haven't looked much into it since then through.
I mean I understand that its nice for the development stage of a CPU, but for all practical purposes, a FPGA is a thing where you can do hyper specialized things in massively parallel fashion, and essentially don't do something to run general purpose code.
I am not saying that people should stop doing this things, everybody is free to do what they want, still i don't understand why most of FPGA talks are about soft CPU's when the really interesting stuff is something completely different.
FPGA-specific soft cores like VexRiscv and NaxRiscv are immensely useful for anything involving state machine logic or glue code that you do not want to implement in-fabric.
Peripherals like on-chip MMCMs/PLLs, on-board I2C and SPI peripherals, etc. with complicated initialization routines or communication flows or sequencing are very easily handled in a soft CPU.
Soft CPUs can also be used like high-powered programmable in-circuit logic analyzers: without rebuilding a potentially massive FPGA bitstream, you can probe/observe/inspect, inject/alter any signals or buses you pipe to the CPU. VexRiscv is far more pleasant to use than any vendor ILA IP.
Soft CPUs also normally utilize FPGA LUTRAM/BRAM resources, enabling whatever program to run with hard real-time latency consistency.
HW is actually really hard. If you can use a soft core to simplify the overall design and suck up a bunch of peripheral logic it's probably a good idea. Then the engineers can spend their time focusing on getting the hard parts of the design correct.
For example I worked on a project that used fpga to mux audio/video. It simply redirected digital pins. However the internal cpu was used to control/decide what to mux, when and how.
It could’ve been all done in fpga but that would’ve been more work (difficult/tricky/inflexible). Instead we had a small core that run a simple program and communicated to external world.
You wouldn't have only a soft CPU on an FPGA, that's a waste of time and money.
One example: our vendor had an FSM to quickly save and restore trained SERDES parameters. We replaced that with a tiny CPU and it allowed us to make training decisions that could be changed without resynthesis.
Similarly, Altera themselves use a Nios CPU for their DDR4 DRAM controller IO training.
There are so many other possibilities. In one case, we fixed a corner case bug in a HW I2C controller by bit-banging the protocol.
Soft CPUs cost a few thousands gates, one or 2 BRAMs which is totally fine if you have some left. It’s no different than having tons of tiny controller CPUs in large ASICs (which literally everybody does these days.)
What makes the Vexriscv (and Nios and Microblaze) FPGA friendly is that they don’t require zero latency access to the register file. You can use BRAM instead. FF based register files are murder on the FPGA routing fabric.
SeRV implements RV32I and uses 125 LUTs in Artix-7, 198 LUTs in iCE40, 239 LUTs in Cyclone 10LP. Plus 164 FF in each case. Or, apparently, 2.1kGE in CMOS.
Being bit-serial, instructions take either 32 or 64 cycles, vs 3 or 4 for many of the other small RISC-V cores, but it will run at whatever Fmax the rest of your design does, and it's often plenty fast enough.
There's also now QeRV, with the same basic design but with a 4 bit datapath: 3x faster for 15% larger size.
All the acronyms