We had two Itanium servers for a project I worked on, both donated by HP. They weren't anything special, certainly didn't feel like an upgrade over our Sun machines.
We had two Itanium servers for a project I worked on, both donated by HP. They weren't anything special, certainly didn't feel like an upgrade over our Sun machines.
* You're basically encoding microarchitectural details (which operations each execution port can run, how many execution ports, etc.) into the ISA, which makes changing that microachitecture difficult. (See also branch delay slots, which have a similar issue).
* Several instructions have data-dependent execution time, and are very difficult to statically schedule. Dynamic scheduling can handle it much better. The common instruction classes are division, branches, and memory accesses, the latter two of which are among the most common instructions.
* Static scheduling is limited by the inability to schedule around barriers, such as function calls. Dynamic scheduling can overlap in these scenarios.
At the end of the day, the idea that you can rip out all of this fancy OoO-execution hardware and make it the compiler's problem just turns out worse than having the OoO hardware, with the smarter compiler having to manage the instruction stream to maximize the ability of the OoO hardware to get good performance.
The HP/Intel people who designed the Itanium did have an answer for this one: Stop bits, which allow the software to indicate which opcodes can run in parallel even across words such that processors with more parallelism could run instructions from multiple words in parallel. Itanium was designed around EPIC, or Explicitly Parallel Instruction Computing, which was designed as a next-generation VLIW that took into account lessons learned from previous VLIW designs:
https://en.wikipedia.org/wiki/Explicitly_parallel_instructio...
> Each group of multiple software instructions is called a bundle. Each of the bundles has a stop bit indicating if this set of operations is depended upon by the subsequent bundle. With this capability, future implementations can be built to issue multiple bundles in parallel.
Also:
> Several instructions have data-dependent execution time, and are very difficult to statically schedule. Dynamic scheduling can handle it much better.
They tried to handle this with prefetching and speculative loads, but, you're right, they didn't handle it well enough.
> Static scheduling is limited by the inability to schedule around barriers, such as function calls. Dynamic scheduling can overlap in these scenarios.
Itanium had speculative execution and delayed exceptions to try to get around this. Again, though, not good enough.
Itanium was an interesting design, but it seems that VLIW plus modern superscalar techniques isn't as good as superscalar techniques alone.
With Itanium and then NetBurst, it seemed to me that Intel had really bad tunnel vision with the overall designs. It's like they got over focused one a couple use cases and then designed towards those architecturally.
As an example I remember the marketing for Itanium and then NetBurst really hammered on tasks like media encoding. The chips could tear through MPEG macroblocks! Wow! Of all workloads what percent are tearing through MPEG encoding or other highly ordered cache-friendly things? Most real world code is cache-hostile super branchy pointer chasing.
The philosophy of making the compiler do all the instruction scheduling serves the cache-friendly predictable data stream model. The only way it can serve the real world model of code is to explode memory requirements by having the compiler emit hundreds of variants of routines and use some sort of function multi-versioning to select one appropriate for the current data shape.
This is an unscientific observation on my part but it's how the situation seemed to me two decades ago. When Intel didn't have marketeers demanding "moar media encoding" they ended up with genuinely good chips like the Tualatin (which became the Core line).
Current approach of CPU being basically an optimizing compiler that takes assembly in and generates uops to drive whatever execution units the CPU has might be wasteful silicon wise but... caches are much more transistors than this anyway and it allows CPU design to be separated from incoming assembly and so any improvements there are instantly visible to most existing code.
Intel figured they had control over that by not releasing a 64-bit x86. AMD ruined that for Intel.
I worked there at the time, and this was not the thinking in the company at all. The thinking was that if you needed 64-bit computing, you needed to buy an Itanium, full stop. Otherwise, 32 bits (with PAE) was all you needed. This is what Intel told all its customers. They really thought that IA64 was going to eventually replace x86 altogether.
Or some pressure, because there is a strong business need to run new 64-bit code side-by-side with old x86 code (which is very slow on Itanium).
And 32 bit + PAE might have been fine in 2000, but by 2010 the de-facto limit for 4GB memory space per process would've become a huge issue.
x86 code was slow on Itanic because it was just a bolt-on PentiumPro IIRC. It was only there for compatibility mode, and was never meant to be high performance. Code needing performance was going to be compiled for IA64 using the magic compiler that didn't exist yet.
Of course, that was before the DEC Hostile Giveaway.
- the promised compilers are _impossible_ as we can't predict all branches and all cache misses in _general_ (works better for floating point heavy code),
- the failure to clock higher was IMO largely due to a ridiculously bloated and over-complicated ISA. In other words, EPIC was doomed from birth.
I've written about IA-64 many times; it actually had many neat ideas, but in the end Intel yet again failed spectacularly in moving away from their archaic legacy.
Perhaps mostly due to having too many architected registers AND making most of them rotating/part of register windows... that means more work to do per instruction. Wide superscalar => you need lots and lots of forwarding paths between the ALUs. Combined with no out of order execution to allow for spreading that work out a bit (longer latencies per instruction but largely hidden by other work) and you get hard limits on the clock speed.
The ISA wasn't actually that bloated.