15 years ago I thought Itanium was the coolest thing ever. As a compilers student, a software scheduled superscalar processor was kind of like a wet dream. The only problem is that that dream never materialized, due a number of reasons.
First, compilers just could never seem to find enough (static) ILP in programs to fill up all the instructions in a VLIW bundle. Integer and pointer-chasing programs are just too full of branches and loops can't be unrolled enough before register pressure kills you (which, btw, is why Itanium had a stupidly huge register file).
Second, it exposes microarchitectural details that can (and maybe should) change quite rapidly. The width of the VLIW is baked into the ISA. Processors these days have 8 or even 10 execution ports; no way one could even have space for that many instructions in a bundle.
Third, all those wasted slots in VLIW words and huge 6-bit register indices take up a lot of space in instruction encodings. That means I-cache problems, fetch bandwidth problems, etc. Fetch bandwidth is one of the big bottlenecks these days, which is why processors now have big u-op caches and loop stream detectors.
Fourth, there are just too many dynamic data dependencies through memory and too many cache misses to statically schedule code. Code in VLIW is scheduled for the best case, which means a cache miss completely stalls out the carefully constructed static schedule. So the processor fundamentally needs to go out of order to find some work (from the future) to do right now, otherwise all those execution units are idle. If you are going out of order with a huge number of execution ports, there is almost no point in bothering with static instruction scheduling at all. (E.g. our advice from Intel in deploying instruction scheduling for TurboFan was to not bother for big cores--it only makes sense on Core and Atom that don't have (as) fancy OOO engines).
There is one exception though, and that is floating point code. There, kernels are so much different from integer/pointer programs that one can do lots of tricks from SIMD to vectors to lots of loop transforms. The code is dense with operations and far easier to parallelize. The Itanium was a real superstar for floating point performance. But even there I think a lot of the scheduling was done by hand with hand-written assembly.