2) There is way more ILP available at run-time than compile time, and what ILP is available in both is much more tractable at run-time. An out-of-order CPU is constantly filling a buffer with instructions (or microcode), and hardware is determining the dependencies dynamically to issue them to the ALU. This is a much more tractable problem than trying to guess the control flow at compile time. A large enough prefetch buffer can overcome a really dumb compiler.
3) Requiring software to be aware of details of your hardware implementation is a really tempting idea, but it has been historically much worse than the opposite. Consider a modern x86 that is nothing like an 80386, but runs the same software, often at higher IPCs than an original 386. Now compare to MIPS which has its delay-slot for branches which often just gets filled with a NOP. Furthermore on modern MIPS cores, which have longer pipelines and branch-prediction, that slot is more-or-less useless!
4) Assuming IA64 doesn't die out, by the time compiler writers figure out how to make code that runs fast on an Itanium of today, Intel will be performing hardware gymnastics to make that code run fast on the hardware of tomorrow.