Porting OpenVMS to the Itanium Processor Family (2003)[pdf]
de.openvms.org
de.openvms.org
VLIW made sense when Intel wanted to win the FP-heavy workstation market. But while it was in development, integer-heavy web workloads became dominant and that was basically the ballgame.
Now we have RISC-V reinventing the wheel. Not the worst outcome, but we could have had it so much better...
The main issue people tend to bring up with Alpha is the very loose memory model, of the "things happen, but different processors may not really agree on the order they happened in" type of thing (plus, isn't it rude to want to know what other cores have in their cache?). Which would be a pain in our modern multicore world.
Of course we don't know how things would've evolved over time, ARM (at least on big cores[1]) shifted towards the forgiving model for unaligned access, it's possible over time Alpha would've similarly moved to a more forgiving environment for low level programmers.
[1] On embedded stuff, you're going to the Hardfault handler.
StrongArm Punches Up ARM Performance by Jim Turley:
https://websrv.cecs.uci.edu/~papers/mpr/MPR/ARTICLES/091504....
StrongARM: A High-Performance ARM Processor by Rich Witek and JamesMontanaro
But Tim Berners-Lee had a NeXT and some good ideas…
Itanium 2 dropped a lot of the broken ideas from Itanium 1 in favor of .... slightly better than the worst performance in the industry.
re: numerical workloads, they achieved high performance on SPECfp exactly the same way as late-generation HPPA chips -- just throw more L3 cache at it until the number looks good. Not exactly engineering genius.
It wouldn't surprise me if Itanium actually had pretty compelling SPECint numbers. But a lot of those compelling numbers would have come from massive overtuning of the compiler to the benchmark specifically. Something that's going to be especially painful for I/O-heavy workloads is that the gargantuan register files make any context switch painfully slow.
This was one of the more confusing things about Itanium. It's specifications on paper were really impressive. It had crazy high benchmark results (SPEC, Linpack, etc) compared to similarly clocked competing chips. If real world code behaved like those benchmarks Itanium would have blown other chips out of the water.
I don't know if I've ever seen real-world tests showing the Itanium getting anywhere near the performance of its benchmarks. Intel was also charging a huge premium for the privilege.
By 2004 the only thing Itanium had going for it vs a similarly sized Opteron system was the enormous L3 cache. Real good for SPECfp, I guess.
It's been a while so I probably am misremembering the terminology, but I was always amused with the dynamic profile/feedback system that I always imagined would be more useful for a JIT code generator or JVM style runtime than a traditional compiler.
This is not just a benchmark of the CPUs, but also of the compilers involved. It is well-known that it was very hard to write a compiler that generates programs that could harness the optimization potential of Itanium's instruction set.
https://www.usenix.org/legacy/event/usenix05/tech/general/gr...
and, NonStop on Itanium [PDF]:
https://www.researchgate.net/profile/David_Bernick/publicati...