It will obviously be much lower than the IPC of an actual high performance CPU (modern x86-64), but how big is the difference? And how does it compare to typical mobile processors?
It will obviously be much lower than the IPC of an actual high performance CPU (modern x86-64), but how big is the difference? And how does it compare to typical mobile processors?
Still, that's a good bit of work they should be proud for putting out there and I hope other people build on it.
EDIT: Oh, wait, they don't mention register renaming. Hmm, well, I guess no speculating over multiple iterations of a loop then.
EDIT2: No, the PDF the link mentions a rename unit. http://www.rsg.ci.i.u-tokyo.ac.jp/members/shioya/pdfs/Mashim...
[1] https://people.eecs.berkeley.edu/~krste/papers/SonicBOOM-CAR...
I'm curious if this will work on Lattice ECP5- I'm not really sure if Synplify supports system verilog to the same degree as the Xilinx tools. ECP5 is interesting because it's a $10 FPGA..
On the developer tooling side, there is https://github.com/google/verible for linting, code formatting and code indexing.
On the actual compilation side with there are https://github.com/alainmarcel/Surelog and https://github.com/alainmarcel/UHDM which are then being coupled with open source tools like Yosys to allow targeting Xilinx 7 Series and Lattice ECP5 FPGA ICs with fully open source flows using fully open source FPGA tools like symbiflow.github.io
I can't comment on how well it performs myself without testing it, but a quick skim of the paper reveals that it apparently performs well in comparison to its rival open-source out-of-order soft processor.
Comparing soft processors to other soft processors is fairly easy if they can both run on the same hardware, but comparing them to real silicon is inherently kind of meaningless, as they don't really compete at the moment, and the performance of the design in absolute terms will depend on the FPGA it is implemented on. Nonetheless, you could compare the raw numbers presented in the paper for curiosity's sake and see that indeed, it isn't very fast compared to modern silicon processors.
It's 32 bit so it's not desktop-class necessarily but this should blow a microcontroller out of the sky, for scale.
It's worth saying here, that a big ooo CPU is pipelines differently to a small/old risc processor - even amongst discussions about compiler optimizations people still use terminology like pipeline stall, when a modern CPU has a pipeline that handles fetching an instruction window, finding dependencies, doing register renaming and execution, that pipeline is not like an old IF->IF->EX->MEM->WB - it won't stall in the same way a pentium 5 did. The execution pipes themselves have a more familiar structure.
Many modern designs aren't fully bypassed and involve clusters of pipelines to manage this. IBM's recent Power chips and Apple's ARM cores are particularly known to do this.
Is this something that could be automatically optimized via simulation?
Is it something that could be made dynamic and shared across a bunch of schedulers so that cores could move between big/little during execution.
I'm sure it could be automatically optimized in theory, even without the solution being AI complete, but I don't think we have any idea how to do it right now.
No, not unless you're reflashing an FPGA. You'd have better luck sharing subcores for threadlets I think.
It's pretty easy to slap down pipelines. What is far harder is keeping them all fed and running without excessive stalling and bubbles whilst avoiding killing your max frequency and blowing through your area and power goals.
On the other hand, I think there is still space for a small in-order dual-issue FPGA RISC-V with 2 DMIPS / MHz performance.
Actually there is one:
https://opencores.org/projects/biriscv
1.9 DMIPS / MHz..