> I think the key difference was that in the early days, one could only afford a single-pass compiler. Then, double-pass (but it was slow). Itanium was just at the time when compilers could _practically_ do deeper program analysis.
> There is no individual piece in making a good compiler for Itanium which I can't solve. At the time, most interesting CS problems that smart people were working on are what we called "microprogramming" in my college jargon -- optimizing individual algorithms, optimizing an assembly loop, etc. What made an Itanium compiler hard was (again in local jargon) "macroprogramming" -- making all those things work together.
Both these circle back to a point I made earlier in the thread - who’s going to do that for a processor nobody is using? There is an alternate timeline where a combination of hardware and compiler/software iteration make Itanium competitive at a performance level, but intel’s non-technical decisions made that impossible at a practical level. It was never made cheap or available enough where anybody could tinker at the lower end, and at the high end they suffered from the chicken/egg “nobody bought it because x86 was faster today”.
The Itanium did to very well in some supercomputer deployments where the code could handle the architecture’s good parallelism. But that was likely net-new code for a niche product.
> The hypothetical difference was the other way around. CISC had complex instructions (which might take many cycles to run, like a string copy). RISC has simple instructions. Ergo, "reduced instruction set." The technical difference in early processors was RISC was pipelined and CISC was microcode (where all instructions took many cycles).
Hypothetically, yes. In the real world a lot of work was done to minimize that over x86’s lifetime. In the early days RISC did do a lot of what was promised, but clever (some would say hacky in some situations) updates to x86 CPUs and compilers made these advantages less (although one could argue the x86 microarchitectures that showed up in the mid to late 1990s were more RISC like). Faster and better caching (variable length x86 instructions meant that common instructions can have a shorter encoding and take up less space in the instruction cache, vastly reducing expensive cache misses) also minimized a lot of these issues. The main issue was a lot of these “enhancements” kept the performance per watt ratios at very poor levels, which caused no end of headaches as laptops became more popular and left intel (and AMD with x86) with no competitive alternative to ARM in phones.
> The reason for CISC was largely so programs could be smaller. This made a huge difference if a computer has e.g. 8k of RAM. CISC, circa Pentium days, was hard to make fast because: > - CISC variable length instructions were hard to decode, and RISC fixed-length ones were easy. This mostly disappeared as decode units became smaller perhaps circa 2005. > - CISC was hard to pipeline. Again, circa 2005, it became easy to (1) avoid annoying instructions in code and treat them as very slow backwards-compatibility special cases in hardware; or (2) do a translation.
I agree with all these points, but even before the decode enhancements ~2005, lots of work was done to mitigate these. But the most obvious thing that x86 caught up with was clock speed. A central argument in favor of RISC was that it allowed for an ability to jack up the the clock speeds of processors, that could then iterate over the reduced instructions faster and provide better performance in most computing cases. This clock speed advantage (in practice) was eliminated, though not due to issues with RISC itself, but more because it was only the large volumes of x86 chips could justify the higher costs of staying near the bleeding edge of transistor manufacturing allowing smaller transistor sizes (something that ARM would eventually come to lead, though).
The CISC instructions were often heavily improved upon over hardware iterations or moved over to new ones (MMX being a famous example) that were more in line with how code was used (hilariously MMX became kind of redundant soon after as GPUs took over those functions). There were also many cases where x86 could do things in fewer instructions than RISC, which helped even more at higher clock speeds as x86 caught up.
Again, I’m not necessarily defending CISC/x86. But it had so much engineering heft thrown at it due to its install base that it often brute forced its way to performance and it was only when performance per watt metrics started to matter that an alternative came on the scene (and it was not something that Itanium would have been better at - ARM would probably be causing the same issues to intel in the data center had Itanium taken off as it was). This was always going to be a hindrance to any replacement. The fact there was genuine competition in x86 kept prices lower and development cycles active, too.
A lot of what you state is correct, but in practice x86 chips were still faster for most workloads out there. In a perfect world all this time and effort would have been heaped on a far better CPU architecture (CISC or RISC), but alas…
Even today x86 still outperforms ARM on most server chips, but our AWS loads are using their ARM chips because it’s cheaper per unite of compute, which is fine for most workloads. This may even get better as more focus is put on ARM via compilers or architecture iterations.
> True. Although that's more of a business distinction than a technical one.
It’s a technical one. They provided immediate technical enhancements, but didn’t mean all your current code had to be rewritten. It was optional (until video games got so sophisticated that they were required then) and you didn’t need to rewrite/recompile your OS to use it. However, GPUs are a lot more niche, so fundamental architecture changes can be done more easily, especially as most code out there is done via higher level APIs.
> I'm betting 50% on convergence, as number of CPU cores grows, and GPU cores become increasingly complex. I think the M1 may be the track we eventually converge on, with diverse cores optimized for diverse tasks.
I agree on this for most consumer products. SoCs have taken over the embedded and mobile market and that can continue to other markets. But we’ll see what happens if AI continues its current trajectory. There it’s the GPUs that matter more than the CPUs and we could see something different emerge.
> I am too -- although that's a business rather than technical argument -- but I'm not sure that's exactly what happened here; I think it's about half-right. I think Intel simply overestimated the amount of time it would stay dominant in the market, underestimated the competition, and underestimated the time Itanium would take to develop.
It’s all about business, though. If the world was purely technical, Amiga, DEC, or Sun would be on top of the world. Intel did everything you just said, but also tried to do too much. A lot of what intel does is over-engineered in a sense they do things in a more complex way that necessary (some cynical people would say on purpose to make larger margins on hardware/chipsets eg USB, which at a low level is very complex even for the original spec).
> I don't think they were intentionally making the CPU either easy or hard to copy.
At an engineering level, no. But higher up they made sure the way it worked with IP, etc would make this difficult. The fact that x86 clones exist at all was a legal miracle and a quirk of history (pretty much IBM demanding it for the original PC and intel was not big enough yet to say no): https://jolt.law.harvard.edu/digest/intel-and-the-x86-archit...