AMD Killed the Itanium (2005)
utcc.utoronto.ca
utcc.utoronto.ca
Secondly, it's true that AMD hammered the nails in the coffin, but AMD wouldn't have mattered if Itanic had been faster, cheap, and on time. Itanic was a disaster partly because of overly complicated design by committee and partly because of the fundamentally flawed assumption (that you don't need dynamic scheduling, AKA OoO processing).
I have an Itanium in the garage, a monument to hubris.
UPDATE: I forgot to mention that from the outside it might seem that Intel had a singular vision, but the reality is that there were massive political battles internally and the company was largely split into IA-64 and x86 camps.
UPDATE2: Itanium was massively successful in one thing: it killed off Alpha and a few other cpus, just based on what Intel claimed.
It was trying to optimize for instruction-level parallelism when power efficiency and thread level parallelism were coming into vogue. Arguably companies like Sun overoptimized for the latter too soon but it was the direction things were going.
A senior exec at Intel told me at the time the focus on frequency in the case of Netburst was driven by Microsoft being uncomfortable with highly multi-core designs--and I have no reason to doubt that was one of the drivers. There was a lot of discussion around the challenges of parallelism, especially on the desktop, at the time. It generally wasn't the problem the hand wringing suggested it would be.
Even on a modern ultralight laptop, I can run two chrome profiles, three instances of vscode running different projects, docker and a few other things and the CPU never gets pegged. There's a ton of memory pressure from a memory leak somewhere that I haven't bothered tracking down yet- I suspect the SWC compiler (thanks, rust) but haven't proven it yet. All that and I'm still getting 8+ hours of battery life.
I'm really tempted by the idea of a Framework style laptop with user serviceable parts, but my work style has me moving a fair amount, so not having battery or thermal issues is such a boon I don't know that I could make the switch. In the last 6 hours I've written and compiled code, run tests, attended video calls, streamed video from websites, browsed the internet for recipes for dinner, chatted on slack and am still on 46% battery remaining. I have yet to hear the fans turn on.
The two areas this thing will fall down on is music and gaming. The speakers are pretty bad, and though it can run light games off of steam, I doubt it would do well with anything super graphically intense (though I haven't actually tried much, to be honest). Also, the built-in webcam sucks, but decent webcams and headphones are cheap, so it's really only games that you'd want something else for.
I've looked at the LG gram before (the 2021 model [0]), but wasn't convinced. And now after seeing a friend's new Lenovo legion 5 I'm even more uncertain of which one should I pick. (That Lenovo has a handy button to set the power envelope, which seems to actually work.)
I also usually move a lot, but I don't want to optimize for that. It's easier to find a power outlet than to cool a throttled laptop.
I don't know if / how it runs linux, but I've got a friend with a legion, and was also happy with it.
If you'd rather have better graphical performance than battery life, definitely go with the legion. If you can't stand the thought of being tethered to a power cord every time you have to do something serious, then you might want to consider the gram.
I've had gaming laptops before, and after putting the battery through the ringer after awhile it was a struggle to get 4 hours unplugged, which I simply didn't want to deal with again.
I always envisioned a day where we'd have devote the small cores to "parasitic load" tasks-- your media player, Slack/Discord/etc., and a thousand OS maintenance threads. They might run at 95% load to do not very much, but it's no big deal-- the actual software you cared about now has the big cores to itself, and context switching (along with cache and branch-prediction losses) are reduced. I could even imagine getting to the point where tasks could requisition cores on a "no disruptions until actually yielded by the main process" basis for real-time or maximum performance tasks.
My logic goes something like "why have 4 small cores and 4 big cores when you could have 8 big cores". Then the argument goes "yes but the small cores use less power", to which my replay is "true, but the big cores finish faster and can spend more time sleeping, I think the power argument evens out ether way". The real gain is to have better power management for the cores, not by having weird small cores.
One can really notice it on dual cores.
One of the other problems with Itanium was that it was supposed to be an "industry standard" 64-bit processor. But Intel and HP were never quite able to square that with a situation where HP at least saw themselves as more equal than others given their role in the design.
I don't recall any issues with the ISA, and we really liked using its early container capabilities (HP Vault).
Perhaps, there was a belief back then that compilers could be easily optimised or uplifted to generate fast and efficient code, and that did not turn out to be the case. Project management and mismanagement certainly did not help either.
I wonder how a VLIW architecture would pan out today given advances in compilers in last three decades, and whether a ML assisted VLIW backend could deliver on the old dream.
The big problem is SMT, since it's hard to share a VLIW core between processes, while a superscalar core shares really well.
Older GPUs seemed more VLIW-like because they were descended from fixed-function rendering pipelines and essentially just exposed the control signals via instruction encodings. Over time, shader cores have become less VLIW-like, e.g. look at any reverse engineering of recent Nvidia architectures.
This makes sense for the same reason you give for SMT: if you're trying to execute from multiple instruction streams on the same execution units, it makes more sense to use small individual instructions rather than large puzzle pieces.
A modern CPU has the same goal of extracting ILP but accomplishes it in a very different way. The instruction stream is short instructions, each of which specifies a simple operation, and these are reassembled using sophisticated dynamic logic into micro-ops, which then get executed in a fairly similar fashion as a VLIW machine - there are a large number of ports (8 is fairly typical these days), each of which performs a separate operation such as arithmetic, load/store, branch, etc.
A GPU has a similar goal of extracting lots of parallelism but does it in a very different way to both VLIW and modern superscalar CPUs. Each instruction operates over a large SIMD vector - 32 is typical, but this varies from 8 (Intel SIMD-8) to 128 (Imagination & optionally Adreno). The instruction specifies many copies of the same operation, so doesn't have to be that big. On RDNA3 for example[1], the basic instruction size is 32 bits, but 64 bits is also common (see section 6.1 for a summary of scalar and 7.1 for a summary of vector encodings).
These instruction sizes are a bit bigger than typical for CPU, for two main reasons. First, there are a lot of registers (256 vector registers), so that needs a lot of bits to encode. Second, it's common to add extra operations such as negation or absolute value in the same operation. But these operations are generally fairly inexpensive modifications on existing data, not completely separate as in VLIW.
In general, execution on a GPU is in-order, so all the reorder buffers and other techniques of superscalar CPUs are not used. Instead of trying to extract as much parallelism as possible from a single thread, a GPU will use that transistor budget to splat more execution units (and thus more threads) on the chip.
[1]: https://developer.amd.com/wp-content/resources/RDNA3_Shader_...
[1] https://en.wikipedia.org/wiki/TeraScale_(microarchitecture)
[2] https://www.anandtech.com/show/4455/amds-graphics-core-next-...
I'm still unconvinced this is fundamental. It certainly was flawed back then, but compiler theory has improved a LOT since then, we have polyhedral optimization, e.g. that we didn't have access to... You could probably optimize delay line technology that way.
If you don't know whether some value will hit in the L1 or the L3 cache there's wild variance on how long it'll take so you have to do something else in the meantime. On x64, that's the pipeline and speculation. On a GPU, you swap to another fibre/warp until the memory op finished.
Fundamentally the hardware knows how long the memory access took, the software can only guess how long it will take. That kills effective ahead of time scheduling on most architectures.
By the 2nd or 3rd generation, Intel had made some changes that really improved performance a lot. Was the final generation of Itanium the 4th generation?
Sure.
But it's very easy to level down. And almost impossible to level up.
You can find out by disabling the stride prefetcher and observing the 90+% performance regression in all your software!
On x86 SMT gets a bad rap because its never been tremendously effective. It's much more effective in IBM's implementation on Power. It wouldn't be hard to imagine some Itanium SMT monster.
With 2 unstalled threads you effectively halve it (in terms of throughput), 4 unstalled threads you effectively quarter it, etc.
Itanium did have SMT
Perhaps as part of the trend towards increasingly heterogeneous architectures, we'll see big VLIW coprocessors for power efficiency in certain serial workloads (GPUs are massively parallel but little VLIW coprocessors; however, they do have dynamic scheduling).
Which would be fine if there was a time machine available to send back what we know now to the past. But at the time having sucky compilers for what you want to (try to) accomplish was a bad idea.
The same decision now may be good, but then it was a mistake. It'd be like trying to have a Moon landing project in the 1920s: we got there eventually, but certain things are only possible when the time 'is right'.
Current way, while arguably pretty wasteful on all the micro-optimizing CPU does on incoming bytecode, allows designers to nearly freely expand hardware to meet the needs without having compilers to produce different code.
OoO execution does an end run around this by examining the program as it runs and adapting to the current reality rather than a simulated reality. This is the same reason a JIT can do optimizations that a compiler cannot do. The ability to look 500-700 instruction into the future to bundle them together into a kind of VLIW dynamically is a very powerful feature.
As to compiler theory, it really isn't that advanced. Our "cutting edge" compilers are doing glorified find-and-replace for their final optimizations (peephole optimization).
Look at SIMD and auto-vectorization. There are so many potential questions the compiler can't prove the answer to that even trivial vectorization that any programmer would identify can't be used by the compiler to the point where the entire area of research has resulted in basically zero improvements in real-world code.
The thing is, on regular, array-based codes, it's great. And it was then, too -- the polyhedral approach was all being developed at exactly that time, but maybe it's not clear in hindsight because the terminology hadn't settled yet. Ancourt and Irigoin's "Scanning Polyhedra With Do Loops" was published in 1991. Lots of the unification of loop optimization was labeled affine optimization or described as based on Presburger arithmetic. But that is the technology that they were depending on to make it work.
But most code we run is not Fortran in another guise. The dependences aren't known at compile time. That issue hasn't changed much.
The one change now is that workloads that were called "scientific computing" then are now being run for machine learning. But now it doesn't make sense to run regular, FP-intensive codes on a CPU at all, because of GPUs and ML accelerators. So what's left for CPUs that excel on that workload? I'm not sure there is a niche there.
Also the Alpha was the utter opposite of design by committee. A quick skim of the alpha ISA would show you that.
And it wasn't like the Alpha was some embodiment of perfection either. E.g. that mindbogglingly crazy memory consistency model.
I don't at all believe this was true.
> that mindbogglingly crazy memory consistency model
I guess the utterly competent designers were actually stupid eejits then? The memory consistency was AFAI could determine to reduce to the utmost the hardware guarantees and therefore hardware complexity. It was done for speed.
I have the greatest respect for the Alpha design team, they designed a thing of elegance and even beauty. You could learn a lot from it - I did.
No, I don't think that. But I don't think they had some superhuman foresight either, and they made some decisions that in retrospect were not correct. And with the memory consistency model, they made the classic RISC mistake of encoding an idiosyncrasy of their early implementation into the ISA (similar to delay slots on many early RISC's).
> You could learn a lot from it - I did.
I used Alpha workstations, servers and supercomputers for my work for several years back in the day. They were good, but not magical, and even back then it was quite clear there was no long term future for Alpha.
I never suggested alphas were magical, but they did seem extremely good and they were designed for future expansion, it seemed to me they were killed off very much not by competitor supremacy.
Good answr, thanks
These barriers also meant that correct multi-threaded alpha code was no longer particularly fast, because you have to insert expensive memory barriers basically everywhere.
Had Alpha not died early, they would absolutely have eventually moved towards a more strict memory model. As it was, it was essentially an irrelevant architecture by the time people started really hitting all the pitfalls.
> These barriers also meant that correct multi-threaded alpha code was no longer particularly fast, because you have to insert expensive memory barriers basically everywhere.
I don't buy it. MBs are for multi-core code, and in such code you typically do much work on a single core then have a quick chat with another core. So the MBs are there for the inter-core chatter only. In that case having fast monocore code is a big win.
The various memory barriers and locking primitives are arch-specific code, and at least smp_read_barrier_depends() is a no-op on all architectures except Alpha. Apparently around the 4.15-4.16 kernels there was a bit of de-Alphafication going on which entailed removing much Alpha-specific code from core kernel code. Further in 5.9 {smp_,}read_barrier_depends() were removed from the core barriers, at the cost of making some of the remaining memory barriers on Alpha needlessly strong.
For more info, search for Alpha e.g. in
https://www.kernel.org/doc/Documentation/memory-barriers.txt
https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin...
https://open-std.org/JTC1/SC22/WG21/docs/papers/2020/p0124r7...
https://mirrors.edge.kernel.org/pub/linux/kernel/people/paul...
The biggest reason Alpha was a "workstation" chip was margin and R&D issues. It was fast, but they couldn't manufacture in high volume which drove per-chip costs much higher than they could have been if paired with a company like Intel. Meanwhile, complete dependence on manual layout for everything pushed development cost and time to market too far out. Once again, Intel's design tools could have helped reduce this overhead.
Could this have worked better with JIT-compiled applications, e.g. Java given a sufficiently clever JVM, where assumptions can be dynamically adjusted at runtime?
(Edit: As opposed to an AOT compiler.)
The other point in favor of this approach now is that far more code is using high-level libraries. Back then there was still the assumption that distributing packages was hard and open source was distrusted in many organizations so you had many codebases with hand rolled or obsolete snapshots of things we’d get from a package manager now. It’s hard to imagine that wouldn’t make a difference now if Intel was smart enough to contribute optimizations upstream.
With respect to ARM, Intel pretty much blew it, especially on mobile. They were so determined to exploit their x86 beachhead. I remember at an IDF, they were even trying to make a case for how it was important to run x86 everywhere so that Flash would run consistently.
One concern I remember was correctness: one company I worked at didn’t find a benefit worth dealing with a second compiler’s quirks and IIRC some scientists I supported evaluated it but never used it because some of their model output varied (classic floating point drift).
With all due respect, this is simply not true. Especially for the time, GCC was the most sophisticated compiler out there, and the only one that could be easily retargeted to a new or another platform due to the use of the intermediate representation language (IRL). A new code generation backend could be boostrapped within days using the IRL. Cross-compilation for any supported target platform whilst running on the same host was also only possible with GCC. There were no other known precedents at the time (I am not counting pcc, the portable C compiler, due not being comparable to gcc).
In terms of the code generation, GCC was quite up there as well albeit performance and quality of the generated code varied across platforms, sometimes wildly. E.g. the native Sun C compiler generated consistently faster code for SPARC (although their C++ compiler that they had acquired from a third party was buggy as hell).
For Itanium, the GCC Itanium backend was not efficient, and it was a well known problem. On 32-bit x86, GCC generated faster and better code than most commercially available C/C++ compilers with a few exceptions being Intel and Metaware High C compilers (both of which not being widely available and being exorbitantly expensive for a average developer) and being comparable or faster to the Watcom C compiler (Watcom did have an edge over GCC on producing much faster floating point code and the default C struct alignment rules, and GCC had an edge due to allowing control over how many CPU registers could to pass input parameters into a function). Id Software, with the release of Quake 1, had ditched the Watcom C compiler and replaced it with GCC (DJGPP) due to GCC generating better code (I think John Carmack wrote about it at some point). GCC did not support Windows well, though.
GCC and LLVM have both been very sophisticated compilers albeit pursuing different objectives. LLVM appeared due to disagreements over the GCC licensing that RMS was insisting on – to preclude GCC from becoming extensible and allow 3rd parties to produce closed source plugins. So, LLVM was conceived as a modularised and extensible design being more conducive to research, experimentation and extensibility at the expense of supporting fewer platforms. GCC and LLVM have both eventually caught up and have now largely reached the feature parity with each other (with GCC still supporting a larger number of platforms and being a go to choice for embedded development).
I love monuments to hubris so I too have one in my garage-- a four-node SGI Altix 350.
So I don't really buy the "Itanium is bad architecture" story.
It's fate was probably decided around 1999-2000. At that point Itanium was still pretty good against Pentium 3 and Pentium 4. And name "IA-64" indicates Intel didn't plan to make 64-bit Pentiums. So eventually Pentiums would fill low-end segment while the rest would be occupied by 64-bit Itaniums.
AMD killed that plan by releasing AMD64 architecture. It was an obvious upgrade to x86, so it would clearly do better in the market than IA-64. So Intel decided to go for x86-64 too, and Itanium was doomed at that point. They didn't even bother making Itaniums with same clock speed as Xeons.
So it's definitely possible that if AMD decided to stick to 32 bits at that time, Intel would have pushed optimized IA-64. Also AMD64 could be worse than it is. E.g. if they decided to increase only register size but keep the number of registers the same, IA-64 could still come on top.
One of the best aspects of the Opteron was it also happened to be a fantastic 32-bit CPU in addition to AMD64. This was a period where a lot of software, even FOSS wasn't 64-bit clean. There was a lot of pointer arithmetic hiding deep in libraries that were assuming pointers would always be 32-bits.
The Opteron running a 32-bit OS at least as well as a 32-bit Athlon was a huge point in its favor. So your existing system running on new Opteron hardware ran fine and you could mix and match Xeon and Opterons in a fleet. Then switch over to 64-bit on the Opterons for (hopefully) better performance.
The other big problem was that the x86 compatibility story was worse than the earlier hopes. That meant that it not only wasn’t competitive with the current generation competition but often even the previous or worse - note losing to the original Pentium or even a 486 here:
https://tweakers.net/reviews/204/8/intel-itanium-sneak-previ...
Now, they could have improved that but statistically nobody was going to pay considerably more for lower performance in the hopes that a future update would improve matters.
The Athlon and Opteron weren’t just fast, they also had flawless 32-bit support so even if your 64-bit software update never happened you could justify the purchase based on their price/performance.
Intel kept the X86 price at a point where no bean counter would favor investing in new architectures. Fortunately AMD broke the headlock on x86.
I dunno about that.
My own personal opinion is that Intel has never been able to re-architect their way out of the fact that the cornerstone of their success is that they are selling x86, and their customers mostly don't care about the theoretical advantages of the bright shiny new thing. They just want to run their software just like they always have. There's a reason why IA-64 has joined iAPX432, i860, StrongARM and i960 as Intel footnotes (outside of the embedded market).
When they were philosophizing about what Itanic should look like, the only thing that x86 obviously needed from a market perspective was a bigger address space. And AMD was smart enough to deliver on that, and here we are.
Many of the computer nerds watched in awe as vendor after vendor dropped their hugely expensive and engineering heavy custom CPU architectures and lined up behind Intel. IBM was the only big player who didn't swallow the bait. "Even if they fail, that's a huge success" was a common observation at the time.
And sure enough. I don't think they failed on purpose, but business wise it was a win-win situation. The x86 architecture would have won anyway, because of the sheer scale, but the Itanium wreckage hastened it. Everyone needed to move, so why not move to x86/Linux directly?
> In some ways Itanium was the most successful bluff every played in the tech industry. In much the same way that Reagan's Star Wars bankrupted the Soviet Union got almost every single competitor to fold. Back at the beginning of the project, Intel was nowhere in high-end & 64-bit computing. There was HP (PA- RISC), Sun (Sparc), Dec (Alpha), IBM (Power), MIPS (SGI). Intel wisely picked the partner with the stupidest management (Carly) to give up their competitive edge and announce to analysts that Intel's vision/roadmap is so awesome that RISC is dead and that they're going to follow the bidding of their master Intel for their 64-bit plan. Wall Street bought in to the story so much that almost everyone else with competitive chips folded their strong hands to Itanium's bluff - SGI spun off MIPS and MIPS decided to leave the high-end space. Compaq undervalued Alpha and let it die. Sun tried to become a software company and if it weren't for Fujitsu making modern sparcs, sparc would be dead.
> Basically, with nothing but PR and Carly's stupidity, Intel wiped out over half of the high-end computing processor market.
> Thankfully AMD had the vision to see through the bluff, and saw the opportunity for 64-bit computing that worked; and thankfully IBM didn't have someone like Carly around so they saw the value in retaining competitive advantages; or the computing world would be pretty bleak place right now..
An important part of the PR machinery was that by picking up a hot topic from academia, they really got absolutely everyone to talk about VLIW as the next genreration RISC. And everyone already knew that RISC was superior and x86 was a toy, but which also was mostly true at time.
In the end, what won was huge caches and huge OoO pipelines. Linus Torvalds had some strong opinions and well known opinions on this, which turned out to be mostly right.
HP was legendary for culture, like "management by walking around, and talking to the people on the ground", which was different from Fiorina's style.
Compaq was the most noteworthy IBM-compatible PC company, before Dell's dorm room dirt-cheap generic PC clones business skyrocketed into an empire.
DEC was the maker of the PDPs and VAX-based minicomputers on which much of the field of Computer Science was arguably developed, and later MIPS- and then Alpha-based workstations and servers, while also still developing VAXen (the plural form of the word).
All those proprietary CPU ISAs listed (PA-RISC, SPARC, Alpha, POWER), when they were introduced on engineering/graphics workstation computers, were especially exciting, because -- separate from the technical architecture itself -- they would briefly probably be the fastest workstation in your shop. All of these made MS-DOS/Windows PCs and Macs look like toys by comparison (though, eventually, Windows NT 3.51 started to be semi-credible if you just needed to run a single big-ticket application program). And you didn't know what exciting new development would be next.
Maybe it was like if, today, several makers of top-end gaming GPUs resulted in a leapfrogging on a cadence of every few/several months. And if they had different strengths, and, incidentally, curious exclusive game software to explore. Or like the very recent succession of Stable Diffusion, ChatGPT, etc., and wondering what the next big wow will be, what they've done with it, and what you can do with it.
When I knew some Linux developers working on Itanium, some were already calling it "Itanic". (I didn't read much into the name at the time, because were a lot of joke derogatory names for brands and technologies.) Later, I thought "Itanic" was because it was a huge expensive thing that was doomed to sink. The theory in the TFA sounds like most competing ship companies gave up on their own engine designs when they heard how great the Titanic would be.
Then a few years later, Compaq is bought by HP, which does nothing with the remaining DEC/Alpha IP, the same tech that helped AMD build the Athlons.
AMD was also hiring away many Alpha developers IIRC.
[1] https://www.mercurynews.com/2010/04/20/analysts-carly-fiorin...
Whatever her other faults as a CEO, that’s just not what happened with the Itanium. The writing was already on the wall for high-end Unix in the mid-1990s.
HP teamed up with Intel and had them take over the bulk of R&D expense with HP continuing to extract profits from the shrinking market for over a decade. Meanwhile the competitors DEC and SGI and Sun basically went out of business. (IBM of course retained its niche as the only choice for those who only buy IBM.)
Nobody talks about Sun’s contemporary leadership using phrases like “that dumb hockey jock Scott ruined Sparc.” Somehow it’s ok when the CEO was a woman.
You don’t hang around with enough ex-Sun people if you haven’t heard derogatory comments about McNealy. But his ultimate failure at Sun wasn’t the same scale, and Sun was never as well managed or universally revered as HP.
One thing McNealy did get right is that Sun was pretty much the only one of the large Unix vendors that wasn't at least preparing for the possibility of an all-Microsoft future with NT. (IBM was arguably placing more of a small side bet that execs like Mills didn't really believe in but almost no one besides Sun dismissed NT out of hand.
I have an acquaintance who has been at HP for ages and his characterization is more that she was left holding the bag.
Hurd did seem to right the ship when he took over. But, to the degree many of us didn't really recognize at the time, a lot of that was financial engineering and eating of seed corn.
Instead Intel/HP nuked the entire mid/high-end of the industry including their own project and set computing back by a decade or so.
She was also a notoriously terrible CEO for other reasons. And then tried to jump-start a political career with one of the worst campaign videos ever made.
I’m particularly interested in the Alpha because it seems like the thing was designed with many of today’s CPU performance challenges in mind. E.g. simple stuff like 64-bit but also things like caches and multiprocessing (cf the very weak concurrent memory model). See also [2]
[1] eg see this from ’97 claiming that 0.5GHz Alpha could run x86 at comparable speed to a 2GHz PII https://www.usenix.org/legacy/publications/library/proceedin... but I’ve also seen claims that the Pentium Pro of ’95 would have been faster than this
[2] https://dl.acm.org/doi/10.1145/151220.151226 and I think the same here: https://www.hpl.hp.com/hpjournal/dtj/vol4num4/vol4num4art1.p...
Dan Dobberpuhl - founded PA Semi, acquired by Apple in 2008 and kicked off Apple's custom processor strategy.
Jim Keller - arguably the most famous processor architect alive today (Athlon 64, AMD Zen, Apple A4/5).
Computing was in no way set back by a decade. The alternative to Itanium was, as Gelsinger has said publicly, an enhanced x86 Xeon--which is what Intel ended up doing (and which HP subsequently adopted to run HP-UX and it's other enterprise OSs).
Obviously ARM has won out over x86 on mobile and--in a limited way--on the desktop. ARM's footprint will probably increase. We'll see. Then there's RISC-V. But that's all basically RISC.
Typed it on a phone and noticed the error only after the edit timer expired.
HP was the only casualty to Itanium, but that was self-inflicted.
Sun was in a somewhat similar boat with the SPARC v8 architecture, and they were rather late with UltraSPARC (SPARC v9 ISA). Yet, they managed to hold out longer due to having a switched memory controller and a very wide memory bus, which allowed them to become the best hardware appliance to run the Oracle database (despite being less performant), and divert the cash flow into the UltraSPARC development. UltraSPARC I was underwhelming, and with UltraSPARC II they finally caught up with other RISC vendors and gradually started outperforming some (e.g. MIPS) in some areas.
Amusingly, the 512-bit wide memory bus has made a comeback in Apple M1 Max laptops (laptops!), and M1 Ultra has a 1024-bit wide memory bus.
Why HP went all in with ditching their own perfectly fine PA-RISC 2.0 architecture is an enigma to me tho.
Again, thanks Rick Belluzzo.
SPEC would like to dissagree with you.
I was comparing: a) 32-bit MIPS CPU's with 64-bit MIPS CPU's, b) 32-bit SPARC v8 and 64-bit SPARC v9 (UltraSPARC) CPU's, and c) performance of 32-bit RISC CPU's comparatively to each other. 32-bit MIPS and SPARC v8 CPU's were slow, with MIPS32 being one of the slowest across the entire board.
I was not comparing MIPS64 to UltraSPARC II or III because MIPS64 implementations (especially R10k and R12k) were exceptionally highly performant, especially in numeric computations that UltraSPARC CPU's were not known for at the time. UltraSPARC II/III systems were renowned for very high, sustained overall system throughput, and nor for high CPU computational performance.
At the time, if one wanted a number crunching beast, they had a choice of either MIPS64, or PA-RISC 2.0, or POWER CPU's. Mostly either MIPS64 or PA-RISC 2.0 (I am not including DEC Alpha – another early performance contender – because it perished too prematurely in the acrid belly of Compaq/HP acquisition shenanigans and did not get a chance to advance past 21264).
R8000 (MIPS IV) was fast and later MIPS64 CPU's were very fast, especially on floating point operations, and consistently outperformed competing 64-bit RISC and x86 CPU's because the MIPS64 was an ISA redesign that addressed and fixed many of the problems of the 32-bit version of the ISA.
When the Itanium shipped years late and slower than expected it was too late for any of the competition (except Power) to recover. Granted the x86-64s were ramping up and they would have all had tough competition, even without Itanium.
So long as one puts big, fat "giga-money-losing" and "humiliating" disclaimers on "success", then yes.
Vs. - what if, instead of Itanium, Intel had more-quietly designed and delivered good, high-performance x86-64 CPU's? I'm thinking that, by bottom-line metrics, would have been a vastly more successful business strategy.
Meanwhile, ARM was designing little low-power RISC toys that were obviously no danger to Intel at all.
"AMD originally announced AMD64 in 1999[14] with a full specification available in August 2000"
vs.
"In June 1994 Intel and HP announced their joint effort to make a new ISA..."
This is why Itanium got traction: everyone knew that you needed volume to stay in the game. IBM had a strategy to get that with Apple & Motorola (PowerPC started in 1991), but HP did not have anything like that for PA-RISC. DEC might have gotten there if they’d had a more aggressive partner for the lower-end Alpha strategy but the merger killed any chance of that.
Since x86 was rising so fast, it might not be clear why Intel got involved. That goes back to the licensing rights: they couldn’t prevent companies like AMD from competing directly with them. Itanium was the attempt to close off that line of competition legally and they were willing to attack their own product margins to do it.
Back then, the whole professional world had switched to 64 bit, both from a performance and memory size perspective. That is why the dotcom time basically was based on Sparc Suns. The Itanium was way late, it still was Intels only offering in the domain. Until x86-64 came and very quickly entered the professional compute centers. The performance race in the consumer space then sealed the deal by providing faster CPUs than the classic RISC processors of the time, including Itanium.
It is a bit sad to see it go, I wonder how well the architecture would have performed in modern processes. After all, an iPhone has a much larger transistor count than those "large" and "hot" Itaniums.
What you also have to remember is that Itanic was a very weird architecture. It's hard to write compilers for it, and it made the cardinal error of baking microarchitectural decisions into the ISA.
Or maybe, just maybe one of those other vendors would have gotten their head out of the sand and made a 64bit processor that ran windows. But I think Wintel was set too deep to anyone challenge that
The selling point for Itanium was compatibility but when they failed so badly at that it leveled out the field since you were going to have to recompile anyway.
Windows NT originally shipped with support for all major CPUs targeted by UNIX workstations, yet all of them faded away until Itanium.
Sure, and as any Linux advocate will tell you, most folks on Windows are stuck there due to the proprietary applications that only run there. They don't care much about the OS, but they need the apps that they know and which have their data locked away.
These apps didn't run on the other processors, so Windows on other arches was mainly a curiosity.
https://en.m.wikipedia.org/wiki/FX!32
https://learn.microsoft.com/en-us/windows/arm/apps-on-arm-x8...
Had it not been the case and Intel alongside its OS partners would have managed to push it no matter what.
Microsoft and HP were already on Intel side, also Microsoft already had experience with JIT compiling x86 thanks to their collaboration for Windows NT on Alpha.
That came after AMD64. It was Intel trying to prevent brand damage by saying their AMD compatible cpus were not knock-offs of a competitor that had plagued them with knock-offs.
Back around 2002 or 2003, I sent away to AMD for a set of reference manuals, got back a nice five-volume set which said "x86-64" everywhere and a little note in the box saying "where it says x86-64, read AMD64"
Itanium was already in trouble. It was hot (really hot) and underperforming. It wasn't selling really well, since it was too expensive. Of the UNIX vendors, only HP was left standing behind Itanium. IBM had already long pulled out of Itanium and also out of Monterey, the UNIX that would unify all unixes.
AMD64 was the light that suddenly came shining and everybody knew that was where everybody was going.
Apple did reinvent the smartphone. But, like the iPod, it didn't fully hit its stride for a few years.
I always found the idea behind the VLIW processor architecture to be a quite good one to be honest, but I read many engineers in many places saying that it's a bad one and it was doomed from the beginning.
The article says that the death of Itanium is mostly due to the disinvestment in the IA-64 caused by the threat of AMD overtaking the x86 market, even though the competition for the x86 market probably benefited all of us, I still find it a bit sad that there was this loss of architectural diversity and sometimes I wonder how well the Itanium would perform today if it weren't killed.
I read an article several years ago from an engineer at...ooh, I think it was DEC or IBM. He said that during development of the Itanium, the Intel guys had talked to them and they advised most strongly that Intel drop the project, because they had been down that road and thought it was a dead end.
It's not clear how Intel thought it could possibly work.
If we are talking about JIT, yes, it does, for it instruments the runtime, gathers the information about hot code paths and performs the in-place optimisation. Think of the profile guide compile time optimisation having been carried over into the runtime.
Predication places the burden of creating optimal instruction bundles AND the correct hinting via the use of predicates on the compiler. If stars aligned, the code could perform blazingly fast. It turned out that aligning the stars in an optimal space time sequence was an arduous task due to the actual hints only being available at the runtime.
Which is where JIT has delivered well (and cheaper!) without requiring a radically different VLIW design.
Even the M1 could be argued to be close, it’s a very wide machine.
Of course actual VLIW still are around as DSPs.
VLIW is considered: Multiple Instruction Multiple Data, in each line of assembly you can send out something like 4 (or 8) instruction each with a different target, and it will work as long as there aren't dependency issues.
GPUs are still Single Instruction Multiple Data (SIMD), for every vector operation you are doing operation: adding vectors, taking a dot production you are only executing a single op at a time.
SIMDs are really close to the RISC/CISC paradigm, and there's various extensions for other types of SIMD processing in different ISAs used today. VLIW is a much different set of assumptions, requiring the compiler to program in the same instruction level parallelism that a superscalar chip will parallelize via it's architectural features (pipelines/branch prediction/et cetera).
Sort of. It's both, really. On nvidia at least, the threads in the warp are simd, but between warps it's mimd. And that's before we get into SMs.
VLIW seems quite suited to DSP / SDR applications.
More links: https://en.wikichip.org/wiki/qualcomm/microarchitectures/hex...
https://pages.cs.wisc.edu/~danav/pubs/qcom/hexagon_micro2014...
https://blog.tensorflow.org/2019/12/accelerating-tensorflow-...
Ironically, they are also running head first into the compiler issues with almost nothing around taking advantage of their dual issue potential.
https://en.wikipedia.org/wiki/Itanium#/media/File:Itanium_Sa...
I have itanium on a shelf somewhere-- while I was using it I got to do some assembly level debugging to track down a report a GCC bug I hit. Well worth in entertainment the $50 or whatever I paid for it.
That's more than a decade after AMD releasing their first x86 compatible CPU. Intel was very aware of the threat before they started the work on Itanium. They even tried to hold back AMD with a lawsuit which they lost and ultimately allowed AMD to release their 80386 compatible CPU.
I find it more likely that it failed for similar reasons why iAPX 432 and i860 failed. There just wasn't a market.
In the modern landscape, it'd be closer to a v86 mode, where the hypervisor is RISC-V, but user x86 applications can run full speed.
All the legacy PC platform crap can be done away, replaced by RISC-V standard OS-A profile.
I don't think AMD is currently limited by their instruction set. Even if they are, there may be an argument to move to ARM instead of RISC-V to take advantage of the software already ported because of Apple's transition and the Graviton chips. Windows already runs on ARM but hasn't been announced to run on RISC-V, after all.
This is exactly what I see happening. AMD will have to move to RISC-V to stay competitive, and x86 acceleration is a compelling feature they can offer.
>Windows already runs on ARM but hasn't been announced to run on RISC-V, after all.
I doubt this one will be an issue for long. During last summit, in talks by the RISC-V foundation itself (specifically, the technical ones about ongoing ISA work), Windows was mentioned a few times as the reasoning for some new specifications.
This strongly implies Microsoft is working on Windows for RISC-V, even if Microsoft themselves haven't said a word about it.
The year is 2259. The name of the place is Babylon 5.
It's also not clear that would be a gain for AMD. x86 has a lot of lock-in and AMD is one of two viable suppliers of it. Helping along a RISC-V transition would create opportunities for attackers that don't need to license x86. AMD doesn't seem to have a reason for that right now. But maybe that's an Innovator's Dilemma kind of situation and they should be cannibalizing the present to setup their future.
> I won't be surprised at all if the whole multithreading idea turns out to be a
> flop, worse than the "Itanium" approach that was supposed to be so
> terrific - until it turned out that the wished-for compilers were basically
> impossible to write.
>
> - Donald Knuth 4/25/2008 in InformIT [0]
[0] https://www.informit.com/articles/article.aspx?p=1193856The other part however is complete miss
>Let me put it this way: During the past 50 years, I’ve written well over a thousand programs, many of which have substantial size. I can’t think of even five of those programs that would have been enhanced noticeably by parallelism or multithreading. Surely, for example, multiple processors are no help to TeX.[1]
Threadripper got killed off because it was starting to compete with AMD's own products, performance increases started to diminish, AMD very much took a short break when were at the top again. The only difference is that Intel managed to stay on top with their Core architecture for so incredibly long.
That's not necessarily bad, of course. This behaviour has led to competition and competition is almost always good for the consumer.
It's good to have competition in the market.
Given how well competition has worked for consumers, I do wonder if there’s some regulatory option here to reduce those back room deals. Given the way the world runs on microprocessors now there’s a decent argument that maintaining a robust market is like what we used to do to prevent one railroad or steel company from getting too much control.
I've owned AMD CPUs before. 386DX-40 and Athlon II and such.
But Threadripper was unreal. And seemed like it was going to be the norm for an era of HEDT.
But AMD got greedy on the promise of corporate desktop dollars, over-stratified, and let a flagship technology fade into "what about Threadripper? No, it's for Lenovo customers."
The fact that UNIX vendors (SGI, SUN) preferred AMD as a second platform shows how good the AMD was at the low end of the spectrum compared to other offerings.
Intel is well-known for their decade-long roadmaps, and if Itanium had succeeded and consumer-class machines had been in their sights, they would have needed a part on such a roadmap. But I don’t remember ever hearing about it.
They must have been planning to segment the market, possibly more aware than most that if they didn’t, fast, low-margin x86 had a good chance of winning everywhere.
I guess they did alright with Xeon.
That left little room for Itanium to be cheap, because it was competing only in the high end, with less economy of scale, more expensive peripherals, etc.
When this was conceived, Intel’s view was somewhat consistent with Microsoft were others. Remember that at one point Windows NT ran on other architectures (eg DEC Alpha, MIPS).
At this point Intel wanted to solve the compatibility issue by booting on an x86 chip essentially. The article mentions this part.
In the 90s AMD (and Cyrix) had competitive 486 parts due to due to some old licensing agreements. Intel wanted to end this with a new architecture.
But then Itanium was delayed and really expensive. At the same time AMD invented the x86-64 extensions and released Athlon, which was much cheaper and has a much easier path to 64 bit. Due to those same licensing agreements Intel was entitled to those extensions, copied x86-64 and the HP-Intel Itanium died.
In the early 2000s Intel was in a tough spot. They’d defeated the 486 competitors with Pentium. This was the last clock speed race. But Intel hit the 3GHz barrier.
The only thing that saved Intel was their mobile chips (ie Core Duo). They didn’t have clock speed but they had efficiency and performance. At the time you could find articles where enthusiasts found ways to build desktops out of Core Duo chips. They were that good. But Intel tried to milk Pentium just a little too long.
Athlon forced their hand, particularly when Opteron started gaining server market share. Opteron destroyed the last hope Intel had of forcing the proprietary EPIC architecture down people’s throats.
The Core Duo was so successful it is still on the DNA of today’s Intel’s CPUs. It’s where the branding “Core” originated. It saved Intel.
And Intel was showing off an around 10GHz, probably water-cooled, Netburst architecture CPU at the Intel Developer Forum at one point.
As I mentioned in another comment, a senior Intel tech exec told me at the time folks like IBM's Bernie Meyerson were basically going around making fun of Intel for not understanding the issues with leakage current and so forth something like: Of course, we know this stuff. But Microsoft really wants frequency, rather than multi-core.
My conclusion is that Intel leadership really thought they could turn the frequency crank once or twice more. And they really couldn't.
What were the supposed advantages of itanium over x86?
Intel probably would have left the CPU market by now, and one of the promising workstation RISC ISAs of the 90's would have replaced x86 instead (like Alpha).
> What were the supposed advantages of itanium over x86?
Itanium went full in on VLIW (https://en.wikipedia.org/wiki/Very_long_instruction_word), but that bet didn't work out in the real world.
Did they ever ship anything? They pop up from time to time, i haven't heard their company name in a while, i wonder what they're up to as of today.
But just as the PC blew up the minicomputer market, x86 Linux servers and LAMP changed the game at the time Itanium uptake was supposed to explode.
AMD64 made sure there was a path forward on x86, but the pre-dot-com-bust deployments were actually just fine in i686-32-bit memory space.
I do wonder what x86 would've looked like if Intel was the one to extend it to 64-bit, given that 32-bit x86 already had a few reserved extension points that appear to have been put there precisely for this purpose.
It was also incredibly expensive.
It was never "alive" and would have died on its own.
Windows was ported to quite a few platforms, but none of them mattered except x86