Why I will be using RISC-V in my next chip
adapteva.com
adapteva.com
http://www.adapteva.com/announcements/an-open-source-8gbps-l...
Love that they developed and open-sourced a 8Gbps, 1us, I/O interface. That might come in handy. :)
https://www.eecs.berkeley.edu/Pubs/TechRpts/2014/EECS-2014-1...
The good news is that RISC has been around since the early 80's so any killer hardware patent will likely have expired by now. Note that the RISC-V ISA is separate from any hardware implementations. If you get fancy with your micro-architecture you can definitely trample on one of the 1000's of micro-architecture patents filed across the industry.
Benchmarks are controversial but by most measures the RISC-V performance should be “good enough” for most use cases
In my experience, statements like this really mean "it's actually really bad, but we don't want to say that"... after all, who would call benchmarks "controversial" if they were winning? There do tend to be few comparative benchmarks, but here's one that shows how ARM and x86 are competitive, but MIPS is behind in energy efficiency:
http://www.extremetech.com/extreme/188396-the-final-isa-show...
Even modern ARM cores, an architecture that started out being very RISC, use microinstruction-based translation in their front end and probably have more in common with Intel's microarchitectures than MIPS and the like.
I think a "CISC-V" could be more interesting - something x86-like (AFAIK the remaining patents are only on the Pentium and above instruction set extensions, and those may expire soon, so 486 and below are now public-domain ISAs) with a small dense instruction set, but extended in a different direction. "x86, but free". ("ARM, but free" would likely not work for legal reasons.)
Perhaps the issue is that the internals of eg. modern Intel CPUs are locked up at Intel, so the missing step from CISC assembly to efficient RISC is missing (while we've got high-level/C/C++ to CISC covered)?
If you look at the current results for Rocket and BOOM, you'll find that these microarchitectures seem competitive with Cortex A5/A8 and Cortex A15 respectively at comparable frequencies.
In addition to this, they seem to have some significant amount of space savings on die. Power/performance ratios already seem to be better (likely due in part to smaller die area).
On a power-performance front, a bigger OoO RISC-V could likely be competitive in peak performance to an intel, though you'd have to find a market for that. There has also been a lot of discussion about a real vector machine extension, and that would likely wipe the floor with x86's packed SIMD if done right.
Also SPEC isn't a synthetic benchmark, it contains real applications like bzip2 and gcc.
For people designing their own cores for whatever reason (there can be many, research and commercial), RISC-V is very attractive because of all the reasons already given, but it's obviously not the only option.
RISC-V is carefully designed compromise; it scales down to extremely cheap cores and up to superscalar. Like Alpha before it, extreme attention has been paid to avoid features/choices that would be bad for OoOE implementations. Some examples:
- rs1, rs2, rd fields are always in the same location and all register sources and destinations are explicit (makes decoding faster and you can start fetching/renaming without having decoded)
- there are no branch delay slots
- instructions produce at most a single result
- no condition codes etc (dependencies are explicit)
- the sign-bit for all immediate fields is in a fixed location (cheaper sign-extension)
and so forth.
I'm also a big fan of the conditional branch instruction which unlike the Alpha can compare two registers.
ARM wasn't really a pure RISC from the beginning (e.g. multicycle instructions like LDM/STM, pre/post-increment addressing modes, built-in shifts), and I'd say the market success, especially more recently, is attributable to the fact that ARM cores are becoming more x86-like.
On the other hand, MIPS is the quintessential RISC, and hit as enjoyed some success, but was never really known much for amazing performance or efficiency. I suspect RISC-V will be similar, and all the cheap Chinese tablets/phones/etc. that are currently using MIPS may switch to RISC-V instead, although many are unlicensed clones so cost may not be a factor to them.
Alpha completely dominated the performance segment for years, but they fell behind for reasons that has nothing to do with CISC vs. RISC.
Alphas relied primarily on clock frequency to achieve their performance, and were very power-hungry as a result. This is absolutely the RISC philosophy of performance via simpler designs and increasing clock frequency, which stopped being viable long ago. A 200MHz 21064 can do 0.675DMIPS/MHz and consumes 30W, or 0.0225DMIPS/MHz/W; a Pentium (P5) 100MHz has 1.88DMIPS/MHz and consumes 10W, for 0.188DMIPS/MHz/W.
I feel I have to repeat again that none of this had to do with RISC vs. CISC. It had to do with implementation and market muscle.
"Alphas relied primarily on clock frequency to achieve their performance, and were very power-hungry as a result"
So was the P4 and it failed, but Intel had enough capital to take a power efficiency clue from Transmeta and turn the ship 180 degrees.
What killed the Alpha was three things: delays in getting the new models out, the cost of staying in this game, and of course, the arrival of Itanic.
EDIT: An aside: the 64-bit extension of ARM is a completely new and different ISA called AArch64 which incidentally is a lot more RISC-like than the original ARM. IOW, it looks like the ARM designers disagree with you.
> I'm surprised to see a lot of RISC proponents still around, because I think it's quite clear that things didn't quite work out the way they thought it would --- the vision of cheap, simple, high-performance CPUs just didn't happen. Thus I'm not of the opinion that another "MIPS, but free" architecture is such a good idea.
You seem to be implying that there's RISC failed. If you ask anyone who mattered in computer architecture ranging from academics like Patterson who literally wrote the most used books in the field to Intel fellows, they'll all tell you that the essential point the RISC folks were making was proven right. This point being that RISC is requires less logic (=area,power) to implement, it's easier to write compilers for, processors have faster cycle times, and so on.
In the end though, it turned out the market valued software compatibility over performance. That, good marketing and a phenomenally efficient supply chain and fab resulted in Intel winning. But don't confuse this with CISC winning or RISC losing. If Intel had to do a clean slate design today, they'd do RISC themselves. In fact, Intel have a few proprietary microcontrollers hidden inside their chipsets and SoCs that were designed post-2005ish and these have RISC architectures.
> In my experience, statements like this really mean "it's actually really bad, but we don't want to say that"... after all, who would call benchmarks "controversial" if they were winning? There do tend to be few comparative benchmarks, but here's one that shows how ARM and x86 are competitive, but MIPS is behind in energy efficiency:
The benchmark performance depends only on how much effort is put into optimizing the implementation. Intel have on average maybe 200 architects and 200 more designers tweaking their processors for performance for the last 30 years. It is completely absurd to expect a team of 10 or so grad students to compete with them.
It did fail to deliver on all its promises.
RISC is requires less logic (=area,power) to implement
True, but ultimately it is overall energy usage that is important. A tiny low-power CPU with less performance will take longer to complete a task than a larger faster higher-power one, meaning it consumes more energy. That's what the article I linked to shows.
processors have faster cycle times
That's not necessarily a good thing, as Intel's failed NetBurst microarchitecture shows. Making a CPU with such short delays that it can run at upwards of 10GHz is futile, as power dissipation becomes a huge problem long before that.
In fact, Intel have a few proprietary microcontrollers hidden inside their chipsets and SoCs that were designed post-2005ish and these have RISC architectures.
If you're referring to the ARC4 in the Management Engine, I have a feeling that was chosen for reasons other than being RISC or otherwise.
It is completely absurd to expect a team of 10 or so grad students to compete with them.
I remember it said that RISC would be so simple and performant that such small teams could easily design CPUs which vastly outperform the big CISCs at much lower prices, so I don't think it's that absurd of an expectation.
[1] https://github.com/ucb-bar/rocket-chip [2] https://github.com/ucb-bar/fpga-zynq
http://i.imgur.com/0FmTYcH.png
(Fullpage screenshot as Google cache lacked styling)
1. Still the core is not yet solid.
2. Not much powerful debugging tools exist.
3. Chisel itself. Every engineer who is willing to use RISC-V should understand output of chisel source code in order to do ECO or some other low level jobs.
4. BSD itself. It is strong and also weakness. There's no way to merge revision into main repository. It will not be quite trouble if RISC-V remains as it is. But will be a big problem if RISC-V evolves by its own development team.
I hoped this RISC-V becomes mainstream processor since I attended the lecture from Mr.Yunsup Lee 4 years ago. I really want to say "I was wrong at that moment" after all.
https://www.eecs.berkeley.edu/Pubs/TechRpts/2014/EECS-2014-4...
-only 2% of the standard FPGA fabric does useful work..
&
-the really expensive part to develop is the IO (too many standards)
Having an open FPGA is a a good start, but it will only be truly useful if the IO and the optimization is brought up to date with modern fpgas. Today's fpgas are really structured asics.
Very much for an open fpga, but let's not underestimate the investments made by fpga companies to get to the current state of the art. Building platforms is very expensive [ref mythical man month]
VLIW has been the future since the 80s, and we're still waiting for the magic wonder compilers that can actually spit out efficient VLIW code. Even GPUs have abandoned VLIW (AMD TeraScale) in favour of RISC (Nvidia, AMD GCN).
Also wouldn't you say that most of the stuff that used to be implemented on DSPs is now moving into ASICs? I wouldn't be so sure that VLIWs are going to be around forever.
[1] https://github.com/NervanaSystems/maxas/wiki/Control-Codes
Hwacha Vector-Fetch Architecture Manual: https://www.eecs.berkeley.edu/Pubs/TechRpts/2015/EECS-2015-2...
Hwacha Microarchitecture Manual: https://www.eecs.berkeley.edu/Pubs/TechRpts/2015/EECS-2015-2...
Preliminary Evaluation Results: https://www.eecs.berkeley.edu/Pubs/TechRpts/2015/EECS-2015-2...
M.S. Thesis on Mixed Precision in Hwacha: https://www.eecs.berkeley.edu/Pubs/TechRpts/2015/EECS-2015-2...
And the concept of baking into your ISA what the designer believes is the "perfect functional unit mix" is an anti-pattern. What's the perfect mix depends on the benchmark, and it changes from basic block to basic block. A history of failed VLIW projects can attest to this. A dynamic superscalar is far superior, even in power-efficiency.
Has anyone ever done a study / experiment of a VLIW with multiple hardware threads and how that would impact the need for an even mix?
I'm not sure I see how MT would solve the mix problem, if each thread gets an issue cycle (and each thread itself has a bad mix).
The Itanium is a great example of this. Benchmarks that pushed the CPU to its limits looked amazing, but real-world performance with general-purpose code was awful. There just isn't enough parallelism in general-purpose code to justify extremely huge instruction bundles that will mostly be wasted with NOPs.