High-end means building and optimizing miroarchitecture to specific node and prosess technology. Typically it costs $100-200 million to do and it's not high-end after 3 years after anymore.
High-end means building and optimizing miroarchitecture to specific node and prosess technology. Typically it costs $100-200 million to do and it's not high-end after 3 years after anymore.
It has some comparisons so you can judge for yourself. Of course, since this is a project driven by only 2-3 grad students it isn't completely fleshed out. However, you would assume that if the ISA just wasn't suitable for high performance implementations that a project like BOOM would have uncovered that by now.
(The experience of Intel's Itanium, IBM's Cell processor and many others shows that it's not enough to have a few good ideas but you have to have ZERO bad ideas that slow you down to get a high performance design.)
Just doing something about the one potential bottleneck that you feel like doing something about doesn't necessarily get you a gain in performance at all.
The idea of scheduling parallelism in the compiler (VLIW) doesn't work for mainstream workloads because the time it takes for data to come back from the DRAM is highly variable.
A super-scalar processor can possibly run some instructions at the hardware level while other block wait until data gets back from DRAM.
A VLIW processor packs N instructions together (say N=3) and if one of them is blocked by DRAM, they all block. (If one of them is blocked by Optane they all block for a very long time...)
It looks obvious in retrospect but it's amazing how most of the RISC workstation vendors missed it and put themselves out of business by getting on the Itanic train.
(VLIW is successful for DSP and GPU, but that's because workloads like that can have completely predictable fetches)
I don't know what the problem w/ Cell was exactly, but it was the same in that it couldn't pull data from DRAM fast enough to keep the silicon busy.
Also Itanium was a heavily superscalar design; I think you meant out of order.
Cell just plain didn't really let you directly address main RAM, so you never saw unknown DRAM accesses stalling the cores. Access to the local memory was always single cycle.
In neither case was unknown RAM latencies an issue with the design.
And time to market is quite important in this area.
- ARM might license you cores for years and suddenly stop
- you might want to switch to a different vendor but all your code, tooling and knowledge is locked-in
- there is no guarantee that ARM will negotiate with you. Especially true for small companies, community projects, embargoed countries.
- ARM will not allow to relicense their "IP" to 3rd parties, either paid of for free