A Brief History of Microprogramming
people.cs.clemson.edu
people.cs.clemson.edu
The TRW-530 was supposed to compete with the IBM 1401, but the 530 was so unreliable that only two were made. Case Tech (now CWRU) got stuck with one years after it was discontinued. It was used only to copy UNIVAC tapes to IBM tapes as part of a migration to new UNIVAC hardware. The machine at least had decent tape drives. A custom interface connected it to a UNIVAC 1107, as a total slave of the reliable 1107. The machine was so cost-reduced that they didn't have buttons on the front panel; the operator used a "magic wand" (a logic probe) to touch contacts to set bits in registers.
Everybody involved in keeping it running was happy when all tapes had been converted and it could be sent off to the scrap heap.
The microcode in main memory idea came from an era when CPUs were slower than memory. Today, we need three levels of caches to deal with CPUs orders of magnitude faster than main memory accesses. But it wasn't until the 1980s that mainstream CPUs decisively pulled ahead of main memory on speed. In the 1960s, memory was waiting for the CPU, not the other way round.
This led to some strange dependencies on memory. The IBM 1620 had a base 10 multiplication table in main memory, loaded at boot time. That was how multiply worked.
Many architectural decisions in the history of computing came from the speed and cost of main memory. Virtual memory was developed as a cost-saving measure, to allow faking more memory by paging stuff out to disk. Paging out to disk is obsolete today (unknown on mobile, painful on desktops, dumb on servers) although most major OSs still support it.
RISC is one of those too, and it's not a coincidence that the idea started in the early 80s when CPUs were still slower than memory.
Now RISC was behind. Superscalar RISC machines were built. But the lower cost and design simplicity were gone. The motivation for RISC went with them.
On top of that, most RISC machines padded all instructions out to the same length. Compared to x86, programs took 2x as much memory, which not only ran up cost, but required more memory bandwidth, one of the scarcest resources now that CPUs were pumping so many instructions.
A lot of people think about data accesses when talking about memory speeds, but instruction fetches are equally if not far more important to consider --- an OoO/superscalar design can "execute around" slow reads/writes if there are other non-dependent instructions, but literally can't do anything if it runs out of instructions to execute and has to wait for more to come from memory.