What really killed 68k was the thing no one here is qualified to talk about: Motorola simply fell off the cutting edge as a semiconductor manufacturer. The 68k was groundbreaking and way ahead of its time (shipped in 1978!), the 68020 was market leading, the '030 was still very competitive but starting to fall behind the newer RISC designs, leading its target market to switch. The 68040 was late and slow. The 68060 pretty much never shipped at all (it eventually had some success as an embedded device).
It's just that posters here are software people and so we want to talk about ISA all the time as if that's the most important thing. But it's not and never has been. Apple is winning now not because of "ARMness" but because TSMC pulled ahead of Intel on density and power/performance.
Maybe the 68000 -> 88000 transition was the problem?
The Tier 1 Unix system companies (many of whom had been Moto 68K customers) already had their own RISC designs and a lot of the second and certainly third tier companies were getting acquired or going out of business. So by the time there was really a solid 88K product--at least for the server market--almost no one was lined up to design systems around the chip.
Data General did for a while. Forget who else did. But it just never got critical mass.
Sun and DEC and IBM made CPUs for their own computers too - but not to compete for basic PCs. Motorola made a lot of phones at one point but not to the degree that they could lock in top of the line fabs.
It's not that the 68000 family was necessarily impossible to use in lower priced PCs by the way. Philips built the 68070 for use in CD-I and other consumer machines. And Apple and Amiga made it work for a while with more mainstream parts.
Arguably x86 and arm are the "RISCiest CISC" and "CISCiest RISC" architectures, and have succeeded due to ISA pragmatism (and having the flexibility to be pragmatic without breaking compatibility) as much as anything else.
And MIPS failed for the same reason ARM pulled ahead: in the late 90's Intel took a huge (really, huge) lead over the rest of the industry in process and MIPS failed along with basically every other CPU architecture of the era
Amusingly the reason ARM survived this bottleneck is because it was an "embedded" architecture in a market Intel wasn't targetting. But there's absolutely nothing technical that would prevent us from running very performant MIPS Macs or whatever in the modern world.
Had it not been the case, the market would have been driven to Itanium no matter what.
The performance (or complete lack thereof) of x86 compatibility mode was one problem. But VLIW's reliance on software to do the right thing (with compilers etc.) was another big thing even for native IA64 code. And a decent-performing Itanium was delayed enough to arrive during dot-bomb vs. dot-com.
If Itanium had really been the only 64-bit chip available for use in systems from multiple vendors maybe it would have succeeded by dint of an Intel monopoly position. But once x86_64 arrived from AMD and Intel ended up following suit, it was pretty much game over for Itanium.
The point was that ia64 was certainly not an "inferior" ISA, It did just fine by any circuit-design measure you want. Like every other ISA, it failed in the market for reasons other than logic design.
But logic design is irrelevant beyond a certain point. Intel certainly had the process chops at the time and knew how to design microprocessors given a certain set of parameters. It's just the parameters were wrong (see Intel's late step-down from frequency on x86 above all else as well--driven I was told at a very high level by Microsoft's nervousness around multicore).
And I doubt the exceedingly-advanced compiler that the Itanium ISA required contributed in any way to its fp performance.
No, that's pretty much exactly what happened. Were you there?
Beating everyone else at synthetic benchmarks is about as impressive as blowing away paper targets at the firing range. Real applications (and their operating systems) shoot back.
Itanium on the other hand was killed by x86_64, yes.
Digging through Usenet archives, I found a post from John Mashey discussing the RISC-CISC spectrum. I found a copy here:
https://userpages.umbc.edu/~vijay/mashey.on.risc.html
The M68K gets called out as having some really complicated instructions in it,
> We thought that adding more addressing modes was the way you made a machine more powerful
https://www.computerhistory.org/collections/catalog/10265816...
People get caught up in intuitive notions of “complex” or “reduced” when talking about RISC and CISC, so it’s very helpful to have specific ideas about what of problems you’re creating for yourself years down the line, when you’re designing an architecture in the 1970s or 1980s. The M68K has instructions that can do these complicated copies from memory to memory, through layers of indirection, and that creates a lot of implementation complexity even though the instruction encoding itself is orthogonal. Meanwhile, the x86 instructions are less orthogonal, but the instructions that operate on memory only operate on one memory location, and the other operands must be in registers. That turned out to be the better tradeoff, long-term, IMO.
Of course nowadays we have the transistor budgets to make even complicated instructions work.
It is interesting to see what Motorola cut from the instruction set when they defined the Coldfire subset (for those who don't know, the coldfire family of CPU used the same instruction set as 68000 but radically simplified, the indirect addressing methods, the BCD mode, a lot of RMW instructions, etc. were removed. The first coldfire after 68060 which was limited to 75Mhz, ran at up to 300MHz).
I guess it was sorta fashionable back then.
Whatever ISA pragmatism can mean whatever you happen to define it as.
AMD’s processors are also fabricated by TSMC; why aren’t they competitive with Apple?
Again, everything comes down to process. ISA isn't important.
The 3nm parts have just barely dropped in phones due to process delays and poor yields (N3B will miss targets and only N3E will meet the original N3 targets a year or 18 months late). So they are a non-factor in the pc/laptop market.
Happy to see efficiency numbers (something other than cinebench please) but there is no node deficit anymore. Apple laptops and zen laptops are on even footing now.
If there is still an efficiency deficit then the goalposts will have to be moved to something like “os advantage” (but of course Asahi exists) or “designed for different goals!” (yeah no shit, doesn’t mean it’s not more efficient).
Again, I am fairly sure that apple creams x86 efficiency in stuff like gcc (chrome compiles) or JVM / JavaScript interpreting (IntelliJ/VS Code) even without “the apple software”, but, everyone treats cinebench like it’s the end-all of efficiency benchmarking.
I absolutely know the apple stuff creams x86 by multiples of efficiency (possibly 10x or more) in openFOAM - a Xeon-w 28c (or even an epyc) pulls like, 10x the power for half the performance of a m1 ultra Mac Studio.
(And yes, “but that’s bandwidth-limited!” and yes, so are workloads like gcc too! Cinebench doesn’t work like 90% of the core, it doesn’t work cache at all, it doesn’t care about latency. That’s my whole point, treating this one microbenchmark as the sole metric of efficiency is misleading when other workloads don’t work the processor in the same way.)
https://www.extremetech.com/extreme/334856-the-apple-m1-ultr...
And that is without getting into the high-end segment, where Apple doesn't have anything that can compete with the Threadrippers and Epycs.
How does it do in JVM efficiency (IDE) or GCC (chrome compile) efficiency? How does it do in openFOAM efficiency?
But it’s always cinebench.
That's a name I don't see every day. Good to know more people are using it. :)
Doing that required a very large amount of area and transistors in its early days. So much that very smart people thought that the extra area requirements would kill that approach. It still does take a large amount of area, but less and less relative to the available die. Moore's law basically blew past any concerns there.
But it wasn't always obvious that that would be the case.
It's not the sequencing that's the issue, it's the exception correctness. Rolling back correctly and describing to the OS exactly which of the many memory accesses actually faulted in an indirect access is very complex. X86 doesn't have indirect addressing modes and never had to deal with that.
It's certainly not correct to say that the m68060 "pretty much never shipped at all". I have several m68060 systems that would disagree with you. The chip even went through six revisions, with the last two often overclocked to 200% of its official speed. It actually competed well with the Pentium on integer and mixed code, although the Pentium's FPU was faster. Considering the popularity of m68k and x86 at the time, that was pretty darned impressive.
I'd love to get one of these to help do more m68k NetBSD pkgsrc package building:
https://amiga68k.com/?product=bfg9060-68060-cpu-accelerator-...
max instruction size: 486, 12; 040, 22. (x86 has since grown to 15)
number of addressing modes: 486, 15; 040, 44
indirect addressing? 486, no; 040, yes
max number of MMU lookups: 486, 4 (but usually 1); 040, 8 (but frequently 2)
So it wasn’t just manufacturing, it was also the difficulty of the task. 680x0 was second only to the VAX in terms of the complexity of its instruction set.
Anything but simpler and completely irrational to someone coming from the orthogonal ISA design persepective.
For a start, on x86, the loading of an address is a separate instruction with its own, dedicated opcode.
On m68k (and its spiritual predecessor PDP-11), it is «mov» (00ssssss) – the same instruction is used to move the data around and to load addresses. Logically, there is no distinction between the two as an address and a numeric constant are the same thing for the CPU (the execution context defines the semantics of the number loaded into the CPU register), so why bother with making an explicit distinction?
Having 2x separate instruction for loading addresses and moving the data around would have made more sense if data and address registers were 2x distinct register files, which m68k had and x86 did not, and speaking of the registers x86 was completely starved of general purpose registers anyway effectively having five of them (index registers are semi-general purpose anyway so they do not count). Even x86-64 today has 16x kinda general purpose registers which is very poor. AMD29k, as another extreme, could have 256 general purpose registers and 32x has been the sweet spot for many ISA's for a long time.
Secondly, there was the explicitly segmented memory model with near and far addresses. Intel unceremoniously threw the programmer under the bus with having to explicitly manage segments and offsets within each segment to calculate the actual address, and the address could not cross a 64kB segment. Memory segments have been a commonplace and predate x86, yet the complexity of handling them is typically hidden in the supervisor (kernel) level. m68k, on other hand, has had a flat memory space since day 1 that only really, for all practical reasons, took off with Windows 2000 on x86 – almost 2 decades later after m68k got it.
Lastly, comparing max instruction sizes for m68k and x86 is a bit cheeky. m68k has fixed size instruction encodings that allow the CPU to use a simple lookup table to route the processing flow as well as the extracting addressing mode(s) from the opcode could instantly give an indication of the total instruction length
Whereas x86 has had the variadic ones requiring a state machine within a CPU to decode them, especially as the x86 ISA grew in size, often stalling the opcode decoder pipeline due to the non-deterministic nature of the x86 opcode encoding.
For example, if you had an array of 32 bit integers on the stack, an instruction like "MOV EAX,[ESP+offset+ESI*4]" would load the value of the element indexed by ESI. Change that MOV to LEA, and it would instead give you a pointer to that element that can be passed around to another function. Without LEA, this operation would require two extra additions and one shift instruction.
Microsoft Assembler syntax confused this issue by being designed to make some operations more convenient, "type safe", and familiar to high-level programmers, while clumsier at getting to the raw memory addresses.
That led to people using "LEA reg,MyVar" when it wasn't necessary, simply because it was shorter to type than "MOV reg,OFFSET MyVar" :)
LEA is also present on the 68K and has the same uses.
Great answer but the root cause is Motorola failed to find enough customers to drive volume. Increased volume => increased profitability => money to invest in solving scaling problems. Intel otoh gained the customers and was able to scale.
Sun was a big costumer pushing them to make aggressive new chips and they were so disappointed with Motorola that they developed SPARC and released their first SPARC machine at the same time as he 68030 and most costumers preferred the SPARC unless they had software comparability issues.
Apple Mac did get a fair amount of volume to. And there were many other uses as well.
Sure it wasn't the wealth that Intel had, but they were hardly struggling. They clearly sold enough chips to pay a design team.
> '030 was still very competitive
Questionable. A simple ARM low power chip beat it. And the RISC designs destroyed it.
ARM2 even in 1886 basically doubled up 68020 not to mention the bigger RISC designs.
Lots of companies were taping out RISC designs by 1985 and all of them basically beat them.
ColdFire can be made entirely 68k compatible with simple low-overhead emulation software. It could be have been a viable path forward for 68k, but they rolled it out almost 10 years too late, after they'd totally given up on 68k.
Motorola abandoned 68k, is what happened.
In favor of PPC.
RISC was, and still is, the way forward.
x86 was just a distraction along the way, and a dead end.