Nios V – Intel's RISC-V Processor
intel.com
intel.com
It will be interesting to see how it compares against the VexRiscv, an open source RISC-V CPU that's quite popular in the FPGA world with excellent performance (at ~1 Dhrystone MIPS/MHz it's roughly double that of the Nios V) for a very little resources.
An incomprehensible decision is to make the Nios V only support uncompressed 32-bit instructions. This is a CPU that will live inside an FPGA and that will often be paired with a single FPGA block RAM to store instructions and data. By not including 16-bit compressed instructions, they'll add some 30% bloat to the code.
A major negative of Intel's offering is that the Nios V is only available for the Pro version of Intel's Quartus toolchain.
The biggest benefit of the Nios V is that its debugger logic will seamlessly integrate with Intel FPGA's virtual JTAG infrastructure. This means that, as long as you stay within Intel's development ecosystem, embedded SW debugging should be a breeze. With other CPUs, there are additional hoops to make that work. (I wrote about that here: https://tomverbeure.github.io/2021/07/18/VexRiscv-OpenOCD-an...)
The point of these very small CPU cores is basically as a flexible state machine that can be updated (inlcuding fix bugs) in the control path. In a big FPGA design you can have tens, hundreds of them inside larger cores. So worst in terms of performance, but far from worse in terms of functionality, power consumption, flexibility in comparison with complex FSMs.
Today you could use a number of RISC-V implementations (the SERV for example is really tiny: https://github.com/olofk/serv) or other soft CPU cores.
But the Nios was really great when it came, esp Nios II, which removed the issue of having to use different tool chains for the compact, large version was very useful. You can easily extend the ISA, connect co-processors, had a clean interface to the rest of the design. And as long as you used Altera FPGAs free as in beer.
In large systems such a 3G, LTE base station there can be thousands of Nios II cores. They have (are) used in automotive and provides hard real time control.
I mean people can import the VexRISCV verilog and get a faster and better chip for "free" so where does Intel think the value add is here? It is quite curious.
Now it would be interesting if they made hard-logic versions to compete with say the Cortex-M series from ARM or something but this seems like it isn't in the cards.
MicroBlaze, Xilinx's own big soft-CPU core was a combination of slow and large, with a buggy compiler to boot.
They are used for state machines and control algorithms that are complex enough that you don't want to implement them in HDL. High-speed interface (eg, DDR4, high-speed ADC) initialization and training is a common example.
Agree, but Xilinx now basically put 2+ ARM cores in everything which solves the problem extremely well.
Okay, let's do a better one.
On a stratix 10:
II/e uses 414 units, II/f uses 1006 units, and V/m uses 1580.
II/e gets .107 DMIPS/MHz, II/f gets .753, and V/m gets .464.
Coremark for II/e is 19, for II/f is 229, and for V/m is... only 16?
II/e runs at 320MHz, II/f at 300MHz, and V/m at 362MHz.
Overall, not looking great for the cost, but if you really want RISC-V I'm sure that has value.
And what's with CoreMark? I do notice they gave the II/f a big cache for that test.
I think the memory system in that example has a 3 cycle memory access latency, and the CPU waits for one instruction to arrive before requesting the next one. I'm almost surprised that they manage 0.464!
The Coremark score doesn't make any sense indeed. Where did you find these numbers?
I'm sure the 4+2KB of cache helps.
> Where did you find these numbers?
https://www.intel.com/content/www/us/en/programmable/documen...
https://www.intel.com/content/dam/www/programmable/us/en/pdf...
The Nios V core is RV32IA. No multiplier.
If you assume that you're trading LUTs for LUTRAM at 4 bits per LUT then each instruction that uses C saves 16 bits or 4 LUTs. If half the instructions in the program can use C then you can save about 2 LUTs per instruction by implementing C.
If it takes 200 LUTs to implement C (that might be a little low?) then your program code needs to have 100 instructions before you even start to win.
I'd imagine a lot of places where you're replacing something that could almost be a state machine would use smaller programs than that.
Some popular small RISC-V cores don't implement C, and I think all the others give you the choice.
I'm not familiar with Intel's FPGAs but Google suggests "M9K" is the unit of block RAM across at least some of their range. That's 8192 bits or 1 KB. The RV32I [1] register set is 128 bytes. A lot of things wouldn't need any additional state, leaving 896 bytes for code -- 224 instructions without C, maybe 300 with C. If you're under 224 instructions then there's no point implementing C.
[1] you can save 64 bytes of registers by using RV32E. Tests using Embench show, IIRC, up to a 30% code expansion by using RV32E due to register spills and reloads. So even with 1 KB total space for registers+program+RAM that might be a net loss. Or your code might not expand at all. RV32E does save interrupt latency / thread switch time. You're probably not doing either of those.
[0]: https://lore.kernel.org/kvm/82568eff-1eff-5e63-4264-ef0c25f7...
I can see RISC-V gradually taking over, since it now has a foot in the door. But it’ll take a long time and depends on third party development.
One thing that we will have to see about is whether the extension model they have now actually holds in the future e.g. will we see a high performance fork of RISCV in the future, or simply (say) a chinese-only company who don't particularly care using parts of the instruction encoding space for their own purposes.
If it does, would it be a pain to program for (see Cell)?
(1) Unfortunately compressed instructions were included in the Unix profile which causes a lot of pain and makes it mostly impossible to do any partial predecode at I$ fill time. Ironically Arm64 wised up and removed the Thumb variants. (Preempting the code density crowd: there are other ways to deal with this). EDIT: Pain includes dealing with instructions that crosses cacheline or even page boundaries (hello double fault), but worse, not being able to tell what might be an instruction in the I$. It was a sad day when RISC-V with essentially no input from the higher-end concerns decided to force it on us.
You can also use a QEMU or Valgrind-like thing to JIT the C away.
Or you can do what modern x86 does and annotate your cache lines the first time they are actually decoded instead of at cache fill time. That's just a little slower for cold code but the same for hot code. Having the C extension keeps 40% more of your code hot for any given cache size.
There might be some implementation point at which Aarch64 wins by having bigger but fixed length code, but it's not at the low end and I don't think it's at the very high end either.
More compact code takes pressure off icache size and refill rates, and transfer rates from icache to fetch buffer. And ROM size in embedded, sich as on FPGA.
As more designs get produced, there will be more and more pressure on the chip designers to stick to the published specifications, instead of going their own way. Why? Because everyone will want to just use the existing toolchains (GCC, LLVM, etc.) and not have to support proprietary extensions. It helps that the specs themselves are easily available.
The implementations that exist now are not at the leading edge of performance, but they are more than good enough for many applications. This will improve over time.
One of the best features of the ISA is the design for extensions... meaning that extensions were planned from the very beginning. Unlike other architectures, where often the ISA is designed to be "complete", and adding extensions inevitably becomes difficult, expensive (cost or run-time performance, or both), or complex (which slows adoption).
In what market segment?
Will RISC-V win a substantial market shares in embedded / micro- controller usage? Very likely. It is sort of started happening already.
Will RISC-V wins over a substantial market in current x86-64 and ARMv8 segment. Highly Unlikely. I have seen zero argument or proposition that even make sense for this to happen. Other than people want it or certain country want it for whatever ideology reason.
Or you mean designing an ISA from scratch? Which is precisely what ARM have done to ARMv8. And further refined in ARMv9.
There are A32 features that A64 carries over which no one else in the 35 years of RISC ISA design since 1985 has seen fit to copy -- the optional shift/rotate of one ALU operand being the obvious one. Would they really have done that if it was a clean sheet design? If it's so great, why hasn't anyone else done it?
Also no other clean sheet ISA designed for high performance since 1990 (it's a short list: DEC Alpha, Intel Itanium, RISC-V) has included condition codes. Even POWER acknowledges that a single condition code register is a bad idea and has eight of them instead (the others just use integer registers to hold long-lived conditions).
Dreamers gotta dream I suppose.
Moreover, they have no motivation to introduce too many non standard/proprietary extensions as it would require them to add support for these new instructions in mainstream compilers, while they could - with standard extensions - take advantage of the work already done for standard extensions.
Today, a major difficulty a newcomer to the industry faces is that they have 2 options:
1 - Taking an existing ISA and paying a license fee, but benefiting from the work already done on the compiler side.
or
2 - Building their own new ISA for free, but having to do all the work on the compiler side by themselves from scratch.
With RISC-V this problem does not exist anymore which opens the door to newcomers.
There is no way today to build competitive opensource hardware, and RISC-V is not about providing opensource hardware. With RISC-V (or any other ISA) we remain mainly dependent on the manufacturer who implement whatever they want within your processor. I think you are expecting too much from RISC-V, it's not about winning a war against manufacturers, it's about providing a royalty-free high-performance standard ISA which makes it easier for newcomers to enter the market.
RISC-V is built around the idea that manufacturers can make their own non-standard extensions and benefit from existing work.
The first silicon with the final ISA design (or at least RV32IMAC) was the FE310 which shipped on the HiFive1 in December 2016, less than five years ago.
It takes up to two years to go from RTL working in verilator or on an FPGA to manufactured chips on boards in shops.
RISC-V is all over the place in deep embedded. Samsung announced that their 2020 high end phones have RISC-V cores controlling the camera and also the 5G radio. Qualcomm is using RISC-V in their 5G. The very popular ESP32 series of WIFI/BT chips is switching to RISC-V, with the last three (?) models being either partially or fully RISC-V instead of Xtensa.
The last several months have seen announcements of startups who have chip designers from Intel, AMD, and Apple's M1 team and who are now working on *lake/Zen/M1-class RISC-V designs. Those will of course take maybe three or four years to appear on the market, but there is zero reason to think they won't be technically successful.
They said "5 stage", which sounds to me like no out-of-order fancy stuff of the sort we've been used to.
> Calling it a "little FPGA program" seems very dismissive.
Well, I wrote a little 5 stage CPU FPGA program once (In Haskell compiled to Verilog :), but that's another story). It wasn't very hard.
I haven't made a production IC, but I'm told that's much harder. I would be awesome if Intel made a RISC-V chip, even a slow one just good for arduino-type toys, but that's not what happened here.
That's inefficient for an FPGA softcore; wires are too expensive, CAMs are straight up awful, and memory latencies aren't too far off relative to the core clock frequencies to justify OOO stuff in the normal case.
> Well, I wrote a little 5 stage CPU FPGA program once (In Haskell compiled to Verilog :), but that's another story). It wasn't very hard.
Writing a 5 stage is easy.
Writing a 5 stage with no bugs is much harder.
Writing a 5 stage that talks industry standard busses and provides Debug/JTAG Support at high frequency and small gate counts is, well, an actual job.
Not saying you can't do better, but it's not a trivial effort.
There's a lot of processors out there. The vast, vast majority being shipped are in order and a handful of stages.
> Well, I wrote a little 5 stage CPU FPGA program once (In Haskell compiled to Verilog :), but that's another story). It wasn't very hard.
> I haven't made a production IC, but I'm told that's much harder. I would be awesome if Intel made a RISC-V chip, even a slow one just good for arduino-type toys, but that's not what happened here.
The FPGA vendor provided soft cores have about the same amount of engineering rigor as a hard core. They have enough customers for the designs to have enough reach to have the same financial implications for bugs.
The FPGA vendor provided soft cores have about the same amount of engineering rigor as a hard core.
I'm not sure what you mean by engineering rigor, but there is indeed plenty of engineering to get from a soft core to a hard core, let alone a fabricated, working chip: synthesis, layout, place and route, timing closure, pads, PLL/DLL, clock tree insertion, BIST, thermal, packaging, etc...Intel has a terrible track record with maker-ish stuff. They routinely launch products, sell them for a year, then discontinue them. (Intel Euclid vision SBC, Intel Edison x86 microcontroller, the realsense stuff, etc)
You can't build an ecosystem around a product if you kill it after a few months. Arudino is only Arduino because they actually stuck around.
Providing a good FPGA CPU to run all the custom stuff is a great way to sell more chips.
This isn't aimed at makers so much as aimed at huge corporations with loads of money to burn on FPGA hardware, circuit design, and custom software.
You can still buy tons of 386EX chips and boards on eBay and even a couple first-hand (JK Flashlight) and there's a few modern homebrew designs out there.
Probably completely overkill for a soft-core.
On top of that, designing and verifying a core that people actually want to use is obviously much more difficult to do than going through HLS.
This is meant to compete with Cortex-M1 and MicroBlaze, not with Cortex-A78 or Core i9 or something.
You get the core you need that meets the processing budget with as little cost, area/resources needed and with as little power consumption. That was true then and it is just as true today.
Also Intel's not intimidated by the gate count niche of classic five stage RISCs; they gave that market up decades ago.
Xilinx has Microblaze which is similar to Nios II, both can boot 32-bit linux.
It is not bit patterns stored in a memory and interpreted, executed by something as a sequence of instructions. Routing congestion for example doesn't exist as a concept in SW.
Given a clock and access to memory it will execute programs.
Is it that it is a moderately pipelined, single issue, RISC like (explicit load/store, fixed width instructions, fairly large register file) CPU design? Why then, all RISC-V, ARM etc cores targeted for low cost embedded, IoT systems are 1990 style.
https://www.cadence.com/en_US/home/tools/system-design-and-v...
https://eda.sw.siemens.com/en-US/ic/precision/
They contain hundreds up to many thousands of FPGAs. Using support dev tools, you can take a large ASIC design and partition it over these FPGAs.
We once used four interconnected machines like these to emulate a large superscalar, OOO CPU design. It took a few days, but could go from release of reset to loaded OS. Something impossible with a SW simulation. Fun times.
But FPGAs are also used as final target technology. When volume is too low, time to market too tight, where flexibility is important, FPGAs are used. Building an ASIC takes 12-15-24 months (depending on design complexity, target process), and a respin (due to an error) of parts of the design adds months before the product can be released.
Base stations for 3G/LTE/5G is a good example. They used to be built with a combination of ASICs and DSPs. But this took too long to develop, and was to rigid to handle upgrades (soft upgrade from 3G to LTE was/is a great selling point). Using FPGAs reduced development time, and provided generational upgrade in a much better way. Don't replace boards, just SW and FPGA bitstreams.
The Nios II family has existed for years, with very little ongoing development. FPGA soft cores don't go above the ~1.5 Dhrystone MIPS per MHz because the increased complexity to go above the 5-stage pipeline doesn't play well with FPGA-style logic. The Nios II/f already occupied that space.
They have to maintain all the compiler stuff, all the libraries, all the related OS kernel and library changes.
The core also needs updates to keep it current with newer hardware interfaces.
With just a team of 50 people to do all these things, that's a minimum of 10-20 million per year. That's not including the initial development team being much larger. Over a decade or two, hundreds of millions or more is not unlikely.
The libraries are all in C, CPU agnostic, and primarily there to support various Intel Platform Designer modules (you can easily verify that by looking at the BSP). None of that goes away when switching to RISC-V.
I have no idea what you mean with “newer hardware interfaces”, especially in the context of reducing cost: when using RISC-V, you’ll have these newer hardware interfaces (whatever you mean by it) just the same.
It’s not about how much money it costs to maintain the Nios infrastructure, its about how much they’ll save by adding(!) an extra CPU to support.
Somewhat aside: most FPGA IP development of Intel happens in Malesia, so don’t assume US salaries… But if you estimate it at $10M-$50M, shall we agree that writing “billions” was more than a bit over the top?
Intel will save a tiny bit on not having to maintain a custom GNU toolchain. That’s about it.