Faulty instructions in C910 RISC-V CPUs
ghostwriteattack.com
ghostwriteattack.com
Sure, anybody can make a RISC-V CPU, but who really has the capabilities to verify them?
There's a reason the ARM model has succeeded-- that is, providing totally off-the-shelf IP with pre-verified cores (because of their own large verif team). The logical end of RISC-V is that we have custom cores literally everywhere, but verifying them is quite costly.
The (equally) hard part with CPU design is funnily enough not in creating the design, but the verification. (That's kinda one small reason why I think CoPilot-esque tools haven't permeated the hardware design space very much).
Sure the first revisions of a new design will be buggy, but over time with iteration and continuous improvement they'll only get better.
I don't think too many folks will be designing new RISC-V cores from scratch, in the same way that very few people build their own OS's. It'll be contributing features and bugfixes to existing designs (and designing custom extensions).
While it might be true, and remain true for high-end designs, we're really seeing a proliferation of mid-to-low end RISC-V SoCs out of china based on open source IP.
I also made my own assembler as part of learning the ISA and put it at https://github.com/Arnavion/riscv . You can see a screenshot and video of the CPU in action there.
The game's campaign walks you through creating a simple 8-bit CPU and then a larger CPU similar to old ARM. But then it also has a sandbox to do whatever, and you can use whatever design you want to solve the puzzles intended for the original two CPUs.
I think the real question is not how you verify a CPU - we know how to do that. It's how you know how well a CPU has been verified. This is all based on reputation currently.
https://en.wikipedia.org/wiki/Pentium_FDIV_bug
It would be disastrous if a widely-used chip had a bug in a lock prefix that allowed crafted untrusted binaries to soft-brick the whole machine.
https://en.wikipedia.org/wiki/Pentium_F00F_bug
I'm not even mentioning the weirdnesses from the eight-bit era. We all know those chips were heroic efforts to get anything done given transistor budgets, and error checking was simply not possible. The point is, even mature proprietary companies have had severe hardware bugs for longer than many here have been around, even if you discount stuff like spectre and meltdown.
More details at https://www.theregister.com/2024/08/07/riscv_business_thead_...
When I wanted to benchmark their implementation last year I patched a kernel to enable it, and needed to consult the open source part of the core [0] to figure out that they placed the enable CSR bit in a different location than the final ratified spec. [1]
[0] https://github.com/T-head-Semi/openc906 (doesn't include XTheadVector extension)
> No, software updates or patches cannot fix this vulnerability because it is a hardware bug. The only mitigation is to disable the vector extension in the CPU, which unfortunately impacts the CPU’s performance.
https://old.reddit.com/r/RISCV/comments/1emhhzs/alibabas_the...
Disclaimer: I don't build CPUs, and as such I don't know what I'm talking about.
Most modern ISAs like RISC-V provide no way of directly accessing physical memory regardless of privilege (you have to either disable paging or setup page tables to point to the physical memory you want), so it seems unlikely that one could accidentally implement one.
In case of an intentional backdoor it seems surprising that it would not be authenticated with a secret key, but maybe they are very incompetent.
Highly unlikely. It's an actual vector instruction that you would otherwise use. The problem is it bypasses page table protections when invoked with memory operands. There is no debugging utility in this mechanism.
Some of these platforms are incredibly janky atm, so I'm not at all surprised that something like this could slip through.
The real surprise is scaleway rushing them into production.
There seems to be some debugging utility in such a mechanism, e.g. you could use it to run CPU tests, benchmarks or debug code in userspace of a stock OS and then directly communicate with serial port MMIO without needing to pollute the CPU state with a system call or change the kernel to directly map the MMIO into userspace.
Likely an oversight rather than malicious intent.
> The C910-based TH1520 SoC is used by French cloud Scaleway.
https://www.theregister.com/2024/08/07/riscv_business_thead_...
This has been expected, and it will happen again.