https://en.wikipedia.org/wiki/TOP500
> As of November 2020, all supercomputers on TOP500 are 64-bit, mostly based on CPUs using the x86-64 instruction set architecture (of which 459 are Intel EMT64-based and 22 are AMD AMD64-based. The few exceptions are all based on RISC architectures). Thirteen supercomputers, including the no 2. and no. 3 are based on the Power ISA used by IBM POWER microprocessors, three on Fujitsu-designed SPARC64 chips. One computer uses another non-US design, the Japanese PEZY-SC (based on the British ARM[8]) as an accelerator paired with Intel's Xeon.
There are non-x86 architectures in the TOP500, including ones which have less cruft than x86, but the x86 chips keep on being used in some of the fastest machines on the planet. My hypothesis is that x86 cruft doesn't really matter, and you'd need to go to a cruft level that was orders-of-magnitude worse for the ISA choice to dominate performance.
Intel's Pentium processors (the original ones) were doing pretty badly because they made the pipeline too deep at the expense of other things.
You are probably thinking of the Pentium 4 which was designed as a speed demon with a very deep pipeline and failed to reach its target frequency.
Word and Autocad also run on the Mac!
It's always at the expense of something. Transistor budget is fixed.
Also, take into account the CPUs are not always the more expensive part of the compute node - GPUs, HBM, lots of DDR4, and fast networking gear are also pretty expensive and will be more or less constant as you change CPU architectures.
And while counting architectural registers is something, in reality modern out-of-order processors do something called "register renaming" so that more registers can be used on-the-fly as it dynamically creates a data flow dependency graph. Yes, inside each processor.
Also, if it’s in the silicon and there’s a bug in the algorithm, a microcode update can fix it for everyone. If the bug was in the compiler, you’d need to recompile everything to fix it.
It's only noteworthy "feature" was that it was able to ship ARMv8 support before ARM had a proper ARMv8 CPU design. Time to market for a new ISA was fast, but that's about it.
The vast majority of CPUs shipped are not x86 compatible.
Statista claims 23.5 billion microcontrollers are shipped annually.
I know microchip (the PIC people) made press releases roughly annually as they shipped another billion flash microcontrollers. Google found the one from 2011 when they shipped their tenth billion PIC chip.
I find it difficult to get x86 sales figures. Intel gross revenue is high because they have their fingers in everything. AMD financial statements claim about $2B/quarter total revenue, so if you figure the average shipped price of a AMD cpu is $200 and they made all their revenue off CPUs, that would be 40 million CPUs shipped per year, which seems both ridiculously high AND about a 25th the quantity of microchip PICs shipped.
One way to look at the number of ARM CPUs shipped is the licensing / holding company has made enough licensing fees to pay for about "a hundred and fifty billion" ARM chips in its lifetime.
There exists both 16 bit protected mode (available since the 80286) and 32 bit protected mode (available since the 80386).
Removing the 16-bit and 32-bit modes don't actually remove any instructions from the platform (save the binary-coded decimal instructions)--you're largely saving only a few bits of decoder table entries at best. Furthermore, processors reset into 16-bit mode on startup for compatibility reasons, so killing 16-bit and 32-bit mode would introduce major compatibility headaches.
ISA extensions can be more easily removed since there's already a CPUID bit that tells operating systems and applications whether or not they are used. The MPX extension for bounds checking is now regarded as a mistake, and Intel has already confirmed that they are removing it from future processor generations. The TSX extension for transactional memory is apparently on the hit list because of Spectre, and was removed from some processor generations.
The only significant processor execution unit space that is truly obsolete I can think is the x87 floating-point execution unit logic, with the concomitant MMX execution unit logic--SSE is just strictly better for everything here, except if you're trying to actually get the 80-bit precision. But the existence of 80-bit floating point in the 64-bit ABI (i.e., long double) means you'd have a hard ABI break that would potentially break software even written today, and the pain of breaking that ABI is probably not worth whatever savings you get out of it.
Didn't the switch to UEFI effectively reduce the scope of this problem to only apply to motherboard firmware? Operating systems no longer need 16-bit code to boot.
Once you remove a single opcode, it's not really x86 any more, and you need emulation at OS level. But then, once you've done that, why not remove more instructions? Why not remove all of them and start again on a much more power-efficient platform? Why not remove the memory model?
One of the brilliant ideas in the A1 is a flag for whether the current process insists on the slower but more comprehensible x86 memory ordering.
You can e.g. remove all MMX instructions and signal this fact through CPUID. Very few apps require MMX (as in can't work without it).
Another option is to remove the instructions from hardware and emulate them in microcode.
That said, the MMX instructions in particular are so problematic to use (and SSE ubiquitous and strictly better) that I suspect you could introduce a processor that lacks MMX support and break almost nobody, certainly far fewer people than removing x87. I don't know if there is an implicit or explicit actual dependency on MMX anywhere.
I had similar issues with PPC software that just blindly assumed I had Altivec on my G3. It wasn't fun.