I'm very unconvinced by these arguments. Providing enough information to the processor to claw back all the performance that predictors give you, that requires some way of knowing all those things in advance. Statically.
And so you fall into the pit of tar and despair that is relying on Sufficiently Smart Compilers. Same one that couldn't save Intel's shiny new IA-64 architecture (the "Itanic").
Static analysis is just not a tractable way to replace hardware predictors. Even profile-guided information is a questionable answer. Could it be we're missing some key idea that'll take us far beyond all of this? Maybe, but it's probably not the same WLIV designs that require impossible feats of software comparable to asking for the weather 2 years in advance so you can save yourself the hardware cost of an umbrella.
Now, perhaps it would be unfair to criticize the Mill for not taping out an actual chip, after all having new ideas is valuable and fabs are very expensive. But we know this is a tricky subject. Performance optimization ideas only have value after you benchmark them, and they don't have an RTL implementation. Nothing that can run on an FPGA, not even "abstract" Verilog running in a testbench on a software simulator.
If what you want is a DSP that runs super simple hot loops very fast, then build a DSP and make it VLIW all you like. But that's not the workloads people run on a general purpose computer.