Not as nice as a PDP-11 processor, but damned close!
Not as nice as a PDP-11 processor, but damned close!
Edit: well, 24-bit physically, but flat linear address space and an instruction set that permitted 32-bit addressing.
Intel did try to do RISC unsuccessfully with the i860--which also had some sort of VLIW mode, which makes their later Itanium even harder to understand.
Intel was making enough coin from x86 that they were reluctant to EOL it in favour of something RISC. Not that they didn't try. Multiple times (Itanium being the worst, though much later). But each time their replacement failed to ignite, and Intel pulled back from the brink and they pushed something that was a noticeable incremental improvement while keeping backward compatibility.
And with the Pentium & the ISAs that followed they just seemed to pull a rabbit from a hat. They made the x86 CISC architecture scale in a way that pundits at least? didn't seem to think was possible.
Motorola didn't invest in time in doing for 68k what Intel did for x86 with the Pentium. And they screwed over their customers as a result. They basically told everyone (Apple, NeXT, various workstation makers, etc.) to go to 88k. 88k was an architectural failure. So then they did PowerPC along with IBM and others. And people who used 68k previously either went PowerPC (Apple) or rolled their own RISC (Sun with SPARC, HP with PA-RISC, etc.) or just left the market entirely (Apple, Commodore).
ColdFire is roughly to the 68k what Pentium II or so was to the x86 line. It's actually quite nice. They dropped a couple things from 68k but it's about 90% compatible, and can do a very effective software emulation to make itself 100%. But it got pushed basically only as a microcontroller or specialist (routers, ethernet switches, etc.) architecture. You can still buy them from NXP I believe.
These days I don't like big-endian. And the separate address vs data registers seems weird now. But the 68k instruction set was overally very nice. I do imagine sometimes an alternate timeline where Motorola had rolled ColdFire out sooner as an MPU line for consumer stuff and Apple, etc. had used that instead of going PowerPC. I wonder what a 64-bit 68k would have looked like.
IMHO Apple wasted a whole pile of engineering years getting MacOS classic to port over natively to PowerPC. It was extremely unstable and large parts were emulated 68k code for years. It might have saved them some unprofitable years, who knows? They seem to have finally learned the lesson and they own their own CPU now.
The Pentium was not great, not terrible. IIRC 2 ways superscalar in-order, quite faster than the 486 (also because of caches, better memory, faster and wider FSB, way better latencies for tons of instructions) but it would not have been enough if Intel had tried to just make it evolve slowly.
Of course at the time and taking price and competition into account, it was quite good to make faster PCs. But by itself it only showed a limited capability for scaling x86.
The Pentium Pro was the uarch that showed that scaling was possible. Even if they have been deeply refined across the years and generations (putting NetBurst appart, of course), most of the technics it introduced in the x86 world are still a foundation of how the current generation x86 cores work. The rabbit that came out of the hat was the result of quite a number of years of work, most of it parallel to the P5 dev. The Pentium II consolidated that with better support for consumer software of the time (with still copious amount of 16-bit code), and, logically, marketing for consumer hardware (with manufacturing progress and tricks allowing a price not too high).
Eh, I don't agree. Look at the other superscalar microprocessors available around the same time:
* POWER (1990)
* Alpha 21064 (1992)
* PA-RISC 7100 (1992)
* Pentium (1993)
* POWER2 (1993)
* MIPS R8000 (1994)
* PA-RISC 7200 (1995)
* UltraSparc (1995)
* Alpha 21164 (1995)
So, first of all, all of these are in-order processors. Out of order execution didn't come until later ('95 for PA-8000, '96 for R10k and Pentium Pro and '98 for Alpha 21264 and Power3).
Second, all of the RISC superscalar processors prior to the Pentium are also (technically) dual issue.
But Pentium did something unique: It can issue two integer instructions in the same cycle. Before that, all the others could issue only one instruction of a given type (load/store, int, and fp) per cycle. In other words, while they may be "dual issue", your instruction stream needs a mix of those types to get any benefit from multi-issue on those early implementations.
In terms of having 2x execution units of the same type: POWER2 got there later in 1993, but was not a single-chip processor, followed by R8000 (also multi-chip), PA-7200 and 21164.
That's why Pentium had pretty good integer benchmarks, even while it was getting crushed on floating point compared to RISC.
Worth noting that had Commodore's astonishingly out of touch and schizophrenic corporate management not driven the company into the grave, the Amiga engineering team had already settled on PA-RISC as their next generation architecture as well, and I believe even had reached prototype stage on the vapors of cash renained in their final months of existence.
The 90s was the era of Dell JIT delivery and grey-box clones. The margins fell out of the whole market. There was no room for the likes of an Atari or Commodore.
Apple only turned things around for themselves in the next millennium by becoming a luxury brand maker of consumer appliances, and "computers" as we know them have become a much smaller market than handheld devices.
Maybe the "Amiga" brand could have continued as a high end graphics card line for PCs. (But it's not like there was any special sauce left in Commodore eng talent that could have made anything competitive). But Commodore as a maker of home computers or workstations? Nothing could have saved them.
Sometimes I wonder if there was a point in time - where real history switches to a fantasy history - in which a non-pc clone paradigm ends up winning that war. If so what was the earliest point this could have happened, and would it even matter since market forces would ultimately pull any alternate success in the direction of an informally standards-federated clone paradigm anyway.
It seems as if VLIW with a sufficiently smart compiler to handle the instruction packing was the "fetch" Intel kept trying to make happen, but it wasn't going to happen (people kept stumbling on the "sufficiently smart compiler" bit). First with the iAPX 432, then the i860, then Itanium.
Unfortunately Apple crippled the memory bus on most of its products so as not to compete with its more expensive models like the Mac IIfx. So by the time the 486 DX4 came out, there was simply nothing comparable on the Mac side, and then the Pentium was the final nail in the coffin and Apple had to transition to PowerPC. But modern processors are so microcode-heavy, with advanced pipelining, branch prediction and instruction interleaving, that they can't be programmed by humans anymore, at least not anywhere near the level of optimizing compilers. GPUs also killed the golden age of innermost loop optimization in the late 1990s.
If I could have one wish, it would be to revert all processor "progress" since the 68040 or possibly PowerPC 601 around 1995, and make something like RISC-V with under 1 million transistors, little or no cache, and 68k assembly or simpler. Then arrange them on a 2D grid. The 50 billion transistors of an Nvidia RTX 4080 would make for 50,000 cores, and even if we lose 1-2 orders of magnitude for interconnect, that's still 500 cores. I'd put at least a 16 GB memory directly above that, connected by vias, with a content-addressable caching scheme for in-memory processing. A scalable chip like that would give us 3 orders of magnitude more performance than we have today, and pull away at perhaps 100x improvement each decade. Maybe Apple's M1 will eventually do that stuff, but right now it's too mired in domain-specific hardware for video and AI. I view that with the same skepticism as DSLs and mourn the loss of symmetric multiprocessing. But I digress!
I agree we shouldn't have completely abandoned these machines in the 90s. the massive resurgence we see of vector and simd today proves that there was substantial value - and maybe the software story would be a lot better if we hadn't have taken a 20 year piss break.
Though of course I know it's been done and either failed technically, or failed in the market, several times. E.g. fairly recently the Parallela Epiphany stuff (https://parallella.org/2016/10/05/epiphany-v-a-1024-core-64-...) which sounded super-gee-whiz-bang-neat but went nowhere.
But maybe the application domains and compilers just weren't there for it yet.
I like what e.g. Parallax has done with their Propeller 2, in the microcontroller domain. 8 cores round-robining on a shared memory, but with their own smaller localized workspace RAM, and a pile of I/O. No need for interrupts to manage concurrency, just assign a core to it. https://www.parallax.com/propeller-2/
The complexity is sadly mostly inherent in the problems being solved.
I also wouldn't hold my breath while waiting for the 68000 to do anything... all the instructions are enormous, and they take ages to execute.
There are many computational instructions that don't make much sense on addresses, especially if you have rich address modes to begin with. And if you'd need to, there was the "movea" instruction for copying between the register files.
You could, in one single instruction,
struct B {
int x, y, z;
};
struct A {
int a, b, c;
struct B *ptr;
};
void func(struct A *array, int i) {
array[i].ptr->z = 3;
}
Unless I am remembering how it works incorrectly.This is a little closer: https://godbolt.org/z/5K4MEYe31
and IIRC there was also the "LEA" instruction ("Load Effective Address") which stored the computed memory address instead of fetching/storing to that address. And I believe some compilers would (ab)use that to do math as well (though this was complicated by the aX/dX register split)
EDIT: Yeah, so if I'm not misreading MC68020UM the memory indirect mode is slower
move.l #3,([6,a0,d0.l*8],4) ; 9 cycles best case
vs. move.l 6(a0,d0.l*8),a0 ;4
move.l #3,4(a0) ; 3