I don't expect to see competitive RISC-V servers any time soon
utcc.utoronto.ca
utcc.utoronto.ca
Given that moving to RISC-V will make people's life harder, these servers
and their CPUs need to be unambiguously better than the x86 (and ARM)
server systems available at the same time.
While this would generally be true, it misses an extremely important current factor.China really wants to find a way to circumvent US soft power limiting their future growth.
The current tech embargo's by the US and friends are only going to make China try harder, and RISC-V seems perfectly positioned to become that circumvention measure.
Note the plethora of Chinese RISC-V boards and chips coming out... which is exactly what they need to do (iterate as fast as possible).
Because Tenstorrent is licensing 8-wide RISC-V cores and... No one is taking them up on it to turn around and sell.
Because OPENPower is ahead of RISC-V in most ways, with much better software support + IBM pushing it, yet has basically gained zero traction.
RISC-V is nice ISA for high end hardware, in theory, but like other theoretical hardware its irrelevant if its not economical.
In fact, I think ARM was kinda a fluke, and gained traction because of its huge embedded presence and the stars aligning among some big companies.
And you know that how? Because you can't buy a server with it in it?
> Because OPENPower is ahead of RISC-V in most ways, with much better software support + IBM pushing it, yet has basically gained zero traction.
OpenPower was simply not open. It was 'marketing' open. It was tightly controlled by a few companies and they wanted to sell you expensive CPU.
Only way after RISC-V was successful did OpenPower finally change its license terms, by then it was already to late as RISC-V already had all the momentum.
Being from IBM is not actually a benefit for many companies.
We have seen far, far, far larger investments in RISC-V then we have ever seen in OpenPower. And I question if the software support is really that much better by now.
Ventana, Alibaba, Esperanto, Tenstorrent, SiFive are all investing lots of money, there is a lot of traction. I think we should be careful predicting 10 years ahead with any confidence.
Nor a desktop/laptop. And running a linux desktop on it (much less a Windows desktop) is not easy from what I can tell.
I want to be enthusiastic about Tenstorrent in particular, but their supposedly open market PCIe accelerator cards never materialized.
Sure there aren't many boards so far (good list: http://krimsky.net/articles/riscvsbc.html ), but come on, give it a couple more years! The progress is insanely fast for the hardware world.
We are just now in the first half of 2023 getting mass-production chips and boards using cores announced in October 2018 (SiFive U74 -- VisionFive 2, Pine64 Star64 and PineTab-V, Milk-V Mars) and July 2019 (THead C910 -- Sipeed Lichee Pi 4A, Milk-V Pioneer).
This is NORMAL.
It took the same amount of time from Arm A53 announced (October 2012) to Raspberry Pi 3 (February 2016) and Arm A72 announced (February 2015) to Pi 4 (July 2019) and Graviton 1 installed at AWS (Movember 2018).
It takes a while, but once set in motion it is basically clockwork inevitable.
You realize the world doesn't evolve around you and what you need right?
The market for chips is large selling them to you might not be top priority.
It might not seem like it because you never see those.
Given that it's already a success and not going anywhere I don't think it's a stretch to think that it will eventually break into the mobile/server market. Google already announced that they're going to support RISC-V on Android.
The only possible link I can see is that you can arbitrarily modify RISC-V on your core, so it is easier to tivoize your software... but this seems like a reach.
I agree that would be nice. Though for most of these chips it wouldn't help too much without documentation for the chip and all its peripherals/registers, and probably a signing key to get it to actually run your code.
The fact that the code is closed source doesn't really change the trust model much either since the code is usually supplied by the chip vendor. If they wanted to do nasty things and keep them secret they could do it in hardware.
Not every service needs high performance. A web site serving static content, for example, or a low-traffic forum. i co-admin multiple public forums, e.g. sqlite.org/forum, which run on single-CPU nodes and get along just fine with that.
I suspect RISC-V won't have taken over datacenters by then, but I could see it being an emerging, stably growing phenomenon.
Without sounding too critical, I feel like this is sort of a discussion about strawmen. Keller set up the strawman so this is responding to that, but I think the title or something is misleading and so it doesn't do much to address the fact it's still a strawman. "No competitive RISC-V servers" is not the same as "everything will be competitive RISC-V servers."
Now Arm makes up 25% of AWS.
The RISC-V Linux kernel, gcc & llvm, Debian and Ubuntu are all in fine shape for generic C/C++ datacenter stuff. All the interpreted languages are there, the things with their own JITs are in the worst shape because they are the only things that need a lot of special RISC-V work (given gcc and llvm are long done) but coming along nicely.
There are a ton of "Raspberry PI" size and cost RISC-V boards now, with quad core 1.5 GHz dual-issue JH7110 SoC (like Pi 3) or quad core 1.85 GHz OoO TH1520 SoC (like Pi 4).
There is a server chip (SG2042) available RIGHT NOW that has similar cores to the first Graviton generation four years ago, but 64 cores per chip instead of 16. You can get it on a single-socket board today, and dual- and quad-socket boards are coming within months. That's up to 256 cores, hundreds of GB of RAM, lots of PCIe, lots of L3 cache.
The $100 TH1520 boards make a great development device for things to run on the SG2042 as they have the same C910 CPU cores.
RISC-V chips in the same class as current (or very recent) x86 and Apple M1 will be available for development purposes NEXT YEAR and in mass production in 2025 or 2026.
Is RISC-V going to take over entire datacentres? No, of course not. Is it going to have a significant and growing place in them within the next five years? Yes.
Not only going by literal sense about completely take over Server in 10 years. As in every server in use. I am willing to bet in 10 years time half of the shipping CPU used in server wouldn't even be RISC-V. A feat that even ARM is still far off from claiming it.
It is sad that most if not all CEO ( even if they are supposed to be engineers ) likes to paint some overly optimistic pictures just because they have vested interest.
Computational speed, RAM capacity, networking, amount of power used, all play a part.
For instance, say a RISCV 1024-core multi-chiplet CPU with ability to have 4TB on the motherboard; even if it used 400W for the CPU, it might make sense because of how many VMs you could pack per-system. For datacenter it sometimes may make sense to talk about how "wide" rather than how "deep" you can go, especially for cloud or VM tasks. Each core being slightly slower might not matter, in that you have a lot of concurrent tasks that may well be disk or network IO-bound part of the time...
If Intel decided to ship a core that could execute either x86 or RISC-V instructions, we could see massive adoption in no time at all.
The Top 500 supercomputer OS mix is maybe the easiest example to use:
https://en.wikipedia.org/wiki/History_of_Unix#/media/File:Op...
I was specifying large systems for banking applications in the late 90s, and we never installed Linux anywhere. It was all Solaris/HP-UX/AIX (and mainly Solaris!). I actually had Linux running at home at the time, so I was totally aware of it's capability, but there's no way i'd have seen it displacing the Unixes in the sort of timeframe that happened with the Top500.
So, i'd probably conclude that the 'out of nowhere' rise of RISC-V is a possibility. I'd say 5 years sounds quick, but 10 years is a long time. But hey, i'll be able to look back on this comment and reflect what an idiot I was about this in a few years time :)
It is currently almost* impossible to add rvv support to things like SIMDe, because neither gcc nor clang can eliminate redundant rvv vector load stores. (https://godbolt.org/z/ocs5rnzrs)
*you can actually do it now with clang, but only if you use statement expression macros instead of functions and set hardcode -mrvv-vector-bits=n: https://godbolt.org/z/r3haM9Yar
The first batch of hardware that's compliant with these (which include Tenstorrent Ascalon) is expected to show up in 2024.
0. https://wiki.riscv.org/display/HOME/Specification+Status#Spe...
The base instruction set (RV64GC + priv 1.10) was only set in stone in July 2019.
And of course the Graviton project will have started in more like 2015 when RISC-V was a rapidly-iterating draft spec with no chips, just starting to transition from being a university research project to being managed by the RISC-V Foundation.
A lot has changed in the last 4 1/2 years. There are now RISC-V server chips that are better than Graviton 1 (A72-class, 64 cores vs 16), with much better ones in the pipeline.
anyone who's ever run a company knows that people's time is not figurative money but literal money
However, the author is taking a very techno-centric line of argumentation. I think the success of RISC-V will mainly decided on a combination of economical and cultural factors. RISC-V wasn't really created to address a fundamental technical limitation of the current ISA lanscape (power,arm,x86,mips etc...) and it's debatable how much of an impact an ISA has on top-line metrics (Jym kellers said so i think).
IMO RISC-V came to be because of the frustration around the lack of innovation and progress around ISA/uArch/CPU design. The author set out to create a common platform to foster innovation and collaboration.
It's mainly a play on better license and modularity which would enable a wider range of innovative design, better collaboration between the research world and the industry.
The closest analogy i can think of is the way LLVM gained market share on gcc on the last 10 years. LLVM/Clang didnt need to be "better" C++/C compiler than GCC to be viable. It was a combination of friendlier code base, a focus on user experience (mainly error message, compilation speed etc...) and a licensing that friendlier to companies which foster massive investment from some of the biggest players. In parallel, the modular nature of LLVM made it the de-facto research platform, so most of the innovation happened there.
I think something similar will be the deciding factor for RISC-V success in the server landscape : - Do we have enough big, deep pocketed player for whom ARM/X86 is a sufficient pain in the rear to warrant large investments - Will the research community come up with enough innovative design to make the transition interesting.
I can't really comment on the research side. But i think if RISC-V succeed on the server side, it will start will some private design from the hyper-scaler to run large in-house software (google search etc...)
The dynamic from the hyper-scaler perspective is definitely changing and interesting : - The hyper-scaler have a lot of in-house software and services where they have much more control on which software runs. Sure they all use OSS software, but the contact surface is much lower there. Making the transitions pain easier to handle. - The economy of producing/designing a chip is very different from a chip manufacturer vs an hyper-scaler.
PS : powerISA is also open, but doesn't seem to gain any traction.... IBM... always to early at the party
GCC and clang already have enough support for risc-v that we have no concerns for support being there. If GCC and clang have what's necessary, everything else is fine.
The moment we see close to a juncture where we can put an order in for risc-v with very many cores, our investors are going to dump a massive round of funding.
Whether risc-v is popular or not doesn't matter. Our multi-tenant services won't have fundamental security isolation problems, everyone serving on x86_64 and ARM will. Looks like risc-v will dominate where things actually matter, unless some new iteration of OpenPOWER comes out.
Odd then that the one of the first RISC-V hardware implementations called itself "Berkeley Out-of-Order Machine".
In fact the vast vast majority of RISC-V chips shipped to date are in-order, including the JH7110 SoC in this year's VisionFive 2, Star64 and PineTab-V, and Milk-V Mars. The only shipping OoO RISC-V chips are the TH1520 and SG2042 with C910 cores, as used in the Sipeed LicheePi4A and Milk-V Pioneer boards.
BOOM was (and is) an experimental university project, not a product. But as Jim Keller said in a recent interview "with RISC-V a couple of university students created an OoO core that was better than the first Intel i7s which hundreds of engineers worked on".
You could just disable out-of-order execution but there’s a major performance hit. Some simpler ARM designs don’t have speculative execution.
Also, why not have dedicated CPUs per customer? Designing a RISCV CPU for multi tenancy shouldn't be impossible regardless of whether it uses out of order execution or not.
It's a huge problem for cloud providers, but most customers don't fully understand the issue and cloud providers aren't forthcoming that "mitigations" are partial at best, so it's really just a huge problem for cloud customers.
This is not to mention the fact that you can use transient execution itself (without any side channels) to amplify a single cache line being present/not present into >100ms of latency difference. Unless your plan is to burn 100ms of compute time to hide such an issue (nobody is going to buy your core in that case), you can't solve this problem like this.
ldr x2, [x2]
cbnz x2, skip
/* bunch of slow operations */
ldr x1, [x1]
add x1, x1, CACHE_STRIDE
ldr x1, [x1]
add x1, x1, CACHE_STRIDE
ldr x1, [x1]
add x1, x1, CACHE_STRIDE
ldr x1, [x1]
add x1, x1, CACHE_STRIDE
skip:
Here, if the branch condition is predicted not taken and ldr x2 misses in the cache, the CPU will speculatively execute long enough to launch the four other loads. If x2 is in the cache, the branch condition will resolve before we execute the loads. This gives us a 4x signal amplification using absolutely no external timing, just exploiting the fact that misses lead to longer speculative windows.
After repeating this procedure enough times and amplifying your signal, you can then direct measure how long it takes to load all these amplified lines (no mispredicted branches required!). Simply start the clock, load each line one by one in a for loop, and then stop the clock.
As I mentioned earlier, unless your plan is to treat every hit as a miss to DRAM, you can't hide this information.
The current sentiment for spectre mitigations is that once information has leaked into side channels you can't do anything to stop attackers from extracting it. There are simply too many ways to expose uarch state (and caches are not the only side channels!). Instead, your best and only bet is to prevent important information from leaking in the first place.
You have to stop the leak into side channels in the first place, it's simply not practical to try to prevent secrets from escaping out of side channels. This is, unfortunately, the much harder problem with much worse performance implications (and indeed the reason why Spectre v1 is still almost entirely unmitigated).