As a quick recap, superscalar processors have multiple execution units, each of which can execute one instruction each cycle. So if you have three execution units, your CPU can execute up to three instructions every cycle. The conventional way to make use of the power of more than one execution unit is to have an out-of-order design, where a complicated mechanism (Tomasulo algorithm) decodes multiple instructions in parallel, tracks their dependencies and dispatches them onto execution units as they can be executed. Dependencies are resolved by having a large physical register file, which is dynamically mapped onto the programmer-visible logical register file (register renaming). This works well, but is notoriously complex to implement and requires a couple of extra pipeline stages before decode and execution, increasing the latency of mispredicted branches.
The idea of VLIW architectures was to improve on this idea by moving the decision which instruction to execute on which port to the compiler. The compiler, having prescient knowledge about what your code is going to do next, can compute the optimal assignment of instructions to execution units. Each instruction word is a pack of multiple instructions, one for each port, that are executed simultaneously (these words become very wide, hence VLIW for Very Long Instruction Word). In essence, all the bits of the out-of-order mechanism between decoding and execution ports can be done away with and the decoder is much simpler, too.
However, things fail in practice:
* the whole idea hinges on the compiler being able to figure out the correct instruction schedule ahead of time. While feasible for Intel's/HP's in house compiler team, the authors of other toolchains largely did not bother, instead opting for more conventional code generation that did not performed all too well.
* This issue was exacerbated by the Itanium's dreadful model for fast memory loads. You see, loads can take a long time to finish, especially if cache misses or page faults occur. To fix that, the Itanium has the option to do a speculative load, which may or may not succeed at a later point. So you can do a load from a dubious pointer, then check if the pointer is fine (e.g. is it in bounds? Is it a null pointer?), and only once it has been validated you make use of the result. This allows you to hide the latency of the load, significantly speeding up typical business logic. However, the load can still fail (e.g. due to to pagefault), in which case your code has to roll back to where the load should be performed and then do a conventional load as a back-up. Understandably, few, if any compilers ever made use of this feature and load latency was dealt with rather poorly.
* Relatedly, the latency of some instructions like loads and division is variable and cannot easily be predicted. So there usually isn't even the one perfect schedule the compiler could find. Turns out the schedule is much better when you leave it to the Tomasulo mechanism, which has accurate knowledge of the latency of already executing long-latency instructions.
* By design, VLIW instruction sets encode a lot about how the execution units work in the instruction format. For example, Itanium is designed for a machine with three execution units and each instruction pack has up to three instructions, one for each of them. But what if you want to put more execution units into the CPU in a future iteration of the design? Well, it's not straightforward. One approach is to ship executables in a bytecode, which is only scheduled and encoded on the machine it is installed on, allowed the instruction encoding and thus number of ports to vary. Intel had chosen a different approach and instead implemented later Itanium CPUs as out-of-order designs, combining the worst of both worlds.
* Due to not having register renaming, VLIW architectures conventionally have a large register file (128 registers in the case of the Itanium). This slows down context switches, further reducing performance. Out-of-order CPUs can cheat by having a comparably small programmer-visible state, with most of the state hidden in the bowels of the processor and consequently not in need of saving or restoring.
* Branch prediction rapidly grew more and more accurate shortly after the Itanium's release, reducing the importance of fast recovery from mispredictions. These days, branch prediction is up to 99% accurate and out-of-order CPUs can evaluate multiple branches per cycle using speculative execution. A feature, that is not possible with a straightforward VLIW design due to the lack of register renaming. So Intel locked itself out of one of the most crucial strategies for better performance with this approach.
* Another enginering issue was that x86 simulation on the Itanium performed quite poorly, giving existing customers no incentive to switch. And those that did decide to switch found that if they invest into porting their software, they might as well make it fully portable and be independent of the architecture. This is the same problem that led to the death of DEC: by forcing their customers to rewrite all the VAX software for the Alpha, the created a bunch of customers that were no longer locked into their ecosystem and could now buy whatever UNIX box was cheapest on the free market.