> You're basically encoding microarchitectural details (which operations each execution port can run, how many execution ports, etc.) into the ISA, which makes changing that microachitecture difficult. (See also branch delay slots, which have a similar issue).
The HP/Intel people who designed the Itanium did have an answer for this one: Stop bits, which allow the software to indicate which opcodes can run in parallel even across words such that processors with more parallelism could run instructions from multiple words in parallel. Itanium was designed around EPIC, or Explicitly Parallel Instruction Computing, which was designed as a next-generation VLIW that took into account lessons learned from previous VLIW designs:
https://en.wikipedia.org/wiki/Explicitly_parallel_instructio...
> Each group of multiple software instructions is called a bundle. Each of the bundles has a stop bit indicating if this set of operations is depended upon by the subsequent bundle. With this capability, future implementations can be built to issue multiple bundles in parallel.
Also:
> Several instructions have data-dependent execution time, and are very difficult to statically schedule. Dynamic scheduling can handle it much better.
They tried to handle this with prefetching and speculative loads, but, you're right, they didn't handle it well enough.
> Static scheduling is limited by the inability to schedule around barriers, such as function calls. Dynamic scheduling can overlap in these scenarios.
Itanium had speculative execution and delayed exceptions to try to get around this. Again, though, not good enough.
Itanium was an interesting design, but it seems that VLIW plus modern superscalar techniques isn't as good as superscalar techniques alone.