Godson-3: A Scalable Multicore RISC Processor with x86 Emulation (2009)
computer.org
computer.org
What they came up with really reminds me of the inside of a modern x86 core that sort of looks RISC-esque after the decoder, they're just doing the decode in software. It's a really interesting halfway point between a traditional x86 core and transmeta's that way.
But, regarding Godson, Zhaoxin, Hygon, Kaixian and 10 more other...
I can only give another poke to Chinese government bodies behind the "Chinese x86" program for continuing wasting state money for two decades straight, trying to make Chinese x86 clones.
There is a following sentiment that is widely spread in industry bodies that "Chinese chips can't run Windows, thus they are not as good as American ones." That's a complete nonsense to anybody moderately technically literate, but this premise lets those guys keep soliciting government money.
I wouldn't be surprised if someone sticks a x86 decoder onto a RISCV core sort of like the PowerPC 615 did.
We don't need Threadripper performance there, but we do need cheap, predictable, and not subject to supply whims. If you're a Chinese firm on a 10-year time horizon, "supply whims" includes making sure you've got domestic production.
Once we started adding instruction caches, that big win started to not be as big. Ultimately it's the caches that killed off that machine, and the ucode interpreters you saw on the Smalltalk and Lisp machines.
One idea I would like to make it's way from ucode that hasn't yet though is being able to define interrupt boundaries for the real time tasks you allude to.
I've figured this was a good way forward for a while, glad to see someone with the knowledge and access to a fab make it a reality :) I'm hoping that someone proposes a similar extension for RISC-V
Edit: Well I guess there's a J extension for JIT acceleration in development, guess I should have finished reading the comments before replying lol
I've been watching the RISC-V J extension but it seems to be a blank space, a placeholder hoping that someone has a good idea for accelerating JITs. I don't even know what that would mean.
> Azul Systems has built a custom system (CPU, chip, board, and OS) specifically to run garbage collected virtual machines. The custom CPU includes a read barrier instruction. The read barrier enables a highly concurrent (no stop-the-world phases), parallel and compacting GC algorithm.
From the paper, including some bits on the MMU you mentioned:
> Our read-barrier allows us to intercept and correct individual stale references, and avoids blocking the mutator to fix up entire pages. We also support a special GC protection mode to allow fast, non-kernel-mode trap handlers that can access protected pages.
> Having the read-barrier implemented in hardware greatly reduces costs. In our case the typical cost is roughly that of a single cycle ALU instruction.
> The hardware TLB supports an additional privilege level, the GC-mode, between the usual user- and kernel-modes.... Several of the fast user-mode traps start the trap handler in GC-mode instead of user-mode.
> The TLB is managed by the OS in the usual ways, with normal kernel-level TLB trap handlers being invoked when normal loads and stores fail an address translation. Setting the GC privilege mode bit is done by the JVM via calls into the OS. TLB violations on GC-protected pages generate fast user-level traps instead of OS level exceptions.
> The hardware supports a fast cooperative preemption mechanism via interrupts that are taken only on user-selected instructions, allowing us to rapidly stop individual threads only at safepoints. Variants of some common instructions (e.g., backwards branches, function entries) are flagged as safepoints and will check for a pending per-CPU safepoint interrupt. If a safepoint interrupt is pending the CPU will take an exception and the OS will call into a user-mode safepoint-trap handler.
Then read starting with section 3.3 Hardware Read Barrier for the details.
See also C4: The Continuously Concurrent Compacting Collector by Gil Tene, Balaji Iyengar, and Michael Wolf for how they moved this to x86 hardware: https://www.azul.com/files/c4_paper_acm1.pdf
you're probably thinking about Transmeta