VLIW is not a panacaea, engineering being all about tradeoffs after all. But it was intended to not have the complex instruction dispatching logic, with things like speculative execution and branch prediction, in the processor. Instead, using a process called if-conversion the compiler combines the two possible results of a conditional branch into a single instruction stream where predicates control which instruction syllables are executed.
* http://web.eecs.umich.edu/~mahlke/papers/1996/schlansker_hpl...
* https://www.isi.edu/~youngcho/cse560m/vliw.pdf
* https://www.cse.umich.edu/awards/pdfs/p45-mahlke.pdf
* http://web.eecs.umich.edu/~mahlke/papers/1996/schlansker_hpl...
Observe, in considering this alternative history, that the Itanium had 64 predicate registers. People have, in the past few days in various discussions of this subject, criticized Intel for holding on to a processor design for decades and prioritizing backwards compatibility over cleaner architecture. They have forgotten that Intel actually produced a cleaner architecture, back in the 1990s.
Consider the following simple C code: "if (arr[idx]) { ... }". Without speculation, the core must stall until the condition has been read from memory, which can be hundreds of cycles if it's not in the cache. With speculation, these wasted cycles are instead used to do some of the work from most probable side of the branch, so when the condition finally arrives from memory, there's less work left to do.
The pipeline depth only affects what happens when the speculation predicted the wrong way: since the correct way is not on the pipeline, it has to fill the pipeline from scratch.
Also modern OoO CISCs and RISCs have very similar pipeline depths for the same performance/power budget.
https://millcomputing.com/docs/prediction/
It has to. The problem is the speed of light here, not a simple slipup by a CPU designer.
Side-effects could obviously been mitigated better, but hindsight 20/20.