I like the way the mill approaches these architectural problems.
I like the way the mill approaches these architectural problems.
And so you fall into the pit of tar and despair that is relying on Sufficiently Smart Compilers. Same one that couldn't save Intel's shiny new IA-64 architecture (the "Itanic").
Static analysis is just not a tractable way to replace hardware predictors. Even profile-guided information is a questionable answer. Could it be we're missing some key idea that'll take us far beyond all of this? Maybe, but it's probably not the same WLIV designs that require impossible feats of software comparable to asking for the weather 2 years in advance so you can save yourself the hardware cost of an umbrella.
Now, perhaps it would be unfair to criticize the Mill for not taping out an actual chip, after all having new ideas is valuable and fabs are very expensive. But we know this is a tricky subject. Performance optimization ideas only have value after you benchmark them, and they don't have an RTL implementation. Nothing that can run on an FPGA, not even "abstract" Verilog running in a testbench on a software simulator.
If what you want is a DSP that runs super simple hot loops very fast, then build a DSP and make it VLIW all you like. But that's not the workloads people run on a general purpose computer.
How will we ever know that a Sufficiently Smart Compiler is impossible, if we never have processors where compiler intelligence is useful?
Do you have an HP zx6000 lying around, perchance?
Itanium failed early; it had many more problems than compilers. And if we don't have a replacement, we'll never get those compilers from anyone but the most obsessed.
The reason is extremely simple: a speculative OOO processor optimizes dynamically. If you switch that with static compile time optimizations, you are bound to only be as fast as before in some quite limited parameter ranges (like: number of entries in a hash table, size of an image, etc.)
EPIC stood for Explicitly Parallel Instruction Computing, and took the "let's push it all to compiler" to the extreme. Itanium was never supposed to have any form of branch prediction or OOB, because it was supposed to be handled by the compiler.
This led to someone quipping that Itanium was a very fast DSP, but ridiculously expensive.
IMO we need to bring back segmentation hardware. Not the x86 version of segmentation (there's not a feature that x86 wasn't able to make twice as complicated as it needed to be while only giving you half the use cases), but the cleaner object capability on top of paging hardware versions of segmentation. That solves the really rough Spectre cases like even NetSpectre where you can slurp out kernel state remotely from untrusted network packets. Just stick them in a "this memory is untrusted" segment. From there the CPU's dynamic dataflow optimizations can include speculation where memory is marked as trusted.
https://millcomputing.com/topic/on-the-lack-of-progress-repo...