How the 8086 processor's microcode engine works
righto.com
righto.com
V33 was used in PC-98 DO+, and in SOC form (NEC V53) in legendary Akai MPC3000, probably because Roger Linn's original LM-1 and LinnDrum were also full x86 PCs with custom cards.
Sadly I cant find any place that sells V33/V53, maybe could be sourced in Japan.
I had absolutely no idea re: the LinnDrum being a PC. That's really cool. I've never seen one in the flesh, but now I'm wondering if anybody has done conservation of the firmware. It would be a lot of fun to look at the code.
8086 didn't according to https://retrocomputing.stackexchange.com/questions/11327/wha..., but V30 was more like an 80186 which did trap, but then 8080 emulation mode might have worked differently again.
Also I don't recall how I worked out that the extended Z80 instructions were causing the problem. Debugging was a lot more complicated in those days, although programs were also much smaller.
https://www.ceibo.com/eng/datasheets/NEC-V20-V30-Users-Manua...
according to MAME V40 has it as standard exception 6? https://github.com/lesbird/MAME4apple-037b5/blob/master/MAME... Then again https://forum.vcfed.org/index.php?threads/nec-v20-and-cp-m.7... says no.
Did you mean during development? LM-1 does not look like a PC, though I think I do see the NEC CPU in these screemshots: https://www.matrixsynth.com/2014/06/vintage-linndrum-lm-1-sn... - of course there's more to a vintage PC than an 8086.
https://www.elektronauts.com/uploads/default/original/3X/6/8...
https://www.elektronauts.com/t/the-seminal-groovebox-linn-90...
https://www.vintage-radio.net/forum/showthread.php?t=87739
http://studiorepair.com/gallery/Linn/9000/index.html
LM-1 and LM-2 were Z-80 based. Damn, cant edit my first post :( Thank you for correction.
https://octopart.com/upd70136al-12-nec-10428216
If you're complaining about finding politics "here" (i.e. on Hacker News), you should complain to the person posting the article, not to Ken Shirriff. Ken Shirriff can't control what you're presented here.
This is so fucked up in multiple ways I do not know what to say. Even if person is trying to flee Putin's regime - the West is actively denying entry to those people. I would understand if you were Ukrainian but it appears you are not.
Are our inherited sins worse than our committed ones? In other words, is it more disgusting to be born Russian, or to choose to work at Google?
I had a few questions:
Does a CPU's decoder circuit connect directly to this ROM then? If so is it connection to the translation ROM?
The second question I had was regarding the the image used to demonstrate the CORD routine. What are the microinstructions that don't have an address or a source and destination such as NCY, LRCY and RCY?
Lastly the author states:
>"Surprisingly, there's no specific "microcode decoder" circuit. Instead, the logic is scattered across the chip, looking for various microcode bit patterns to generate control signals where they are needed."
Could you elaborate on this? What is meant by "looking for various microcode bit patterns to generate control signals where they are needed"? What exactly is the thing that is doing the looking here? I feel like this is an important point but I am unable to understand it.
The microcode address register loads the instruction opcode directly as the microcode address. For a microcode long jump or subroutine call, the microcode address is loaded from the Translation ROM. So you can think of the Translation ROM as a table of addresses for special cases.
> The second question I had was regarding the the image used to demonstrate the CORD routine. What are the microinstructions that don't have an address or a source and destination such as NCY, LRCY and RCY?
I wanted to focus on the hardware for this post, so I didn't go into a lot of detail on the micro-instructions. NCY is "jump if not carry" (i.e. the carry flag is not set). "NCY 7" does a jump to line 7. "LRCY tmpc" is "left rotate through carry the value in temporary register c". RCY is "reset carry". You can sort of see how it's performing a divide through a loop of shifts and subtracts.
> What is meant by "looking for various microcode bit patterns to generate control signals where they are needed"? What exactly is the thing that is doing the looking here? I feel like this is an important point but I am unable to understand it.
Well, I'm not sure if it is important, but I found it interesting :-) My point was that you might expect the micro-instructions to go through a block of decoding circuitry that generates the control signals for the processor. Instead, various bits from the micro-instruction go all over the processor. There are gates scattered all over that match on patterns such as if bit 3 is set AND bit 4 is clear AND bit 5 is set then do something like load a register. So the decoding logic is ad hoc and extremely distributed. There are some decoding blocks that are more structured, such a turning the micro-instruction's ALU bits into ALU control lines, but these blocks are also scattered across the chip.
Hopefully this all makes sense. Let me know if you'd like me to elaborate.
>"I wanted to focus on the hardware for this post, so I didn't go into a lot of detail on the micro-instructions."
Understood. I hope this means we can look forward to a part 2 then! There is really a lot fascinating things here.
However some of the address bits are ignored, for example anything that matches the bit pattern 00xxx0xx is some kind of ALU operation, either between two registers, or a register and memory. The 3 bits in the middle select the operation, the low bit the operand size (byte or word), and the second one the direction (which operand is in memory, if any - with two register operands both encodings can be used).
All of this extra information is decoded by random logic. The microcode for this class of opcodes looks like this (in one representation anyway, not necessarily the most easy to understand but based on the patent):
M -> tmpa 1 XI tmpa
R -> tmpb 4 none WB,NX
SIGMA -> M 4 none RNI F
6 W DD,P0
On the left are the register moves: M is the first/destination operand, and R the second one. SIGMA (Σ in the patent) is the output from the ALU.The first line of code on the right sets up the ALU ("type 1 instruction" XI: select ALU operation from opcode bits; first operand in tmpa, second one is always tmpb). The next one prepares to write back the result and decode the next instruction. If the result goes to a register, the third line will be the last one executed ("type 4", RNI: run next instruction; F: update flag register). If it goes to memory, the store happens on the fourth line ("type 6", write, default data segment, no auto-increment).
You can find more information in the linked patent (US4449184) and blog post (https://www.reenigne.org/blog/8086-microcode-disassembled)
(This is implemented in hardware with a multiplexer, not as sequential if-then-else instructions.)
There's no decoding logic between the opcode in the instruction register and the microcode address. So yes, the opcode bits are the index into the microcode ROM.
rep_lodsb covered a lot of the instruction handling in the sibling comment.
I hope to write a part 2 before too long.
All these piled-on complications are, first and foremost, seriously wacky, and are not found in well-conceived designs as illustrated by ARM and 6502 that perform much better per unit die area than coeval Intel products.
The designers must have been congratulating themselves on how clever they were being, while painting themselves into corners that ultimately made performance worse than a simpler, more regular design could have been.
This gratuitous complexity spilled out into the ISA, imposing a tax on programmers and, later, compiler writers, that we are still paying today.
Well, I mention that the 8086 is convoluted, complex, and full of corner cases, so I think there's enough editorial :-)
As far as "seriously wacky", there are much stranger systems. For instance, the Honeywell 1800 mainframe has two program counters, and they can go forward or backward. The Intel iAPX 432 has instructions with arbitrary bit lengths, so they have no connection with byte boundaries. The AGC doesn't have shift instructions but special memory locations that shift values you store there. In comparison, the 8086 is a pretty middle-of-the-road processor.
Some small minicomputers like the PDP-8 in the 60s/70s were hardwired logic; they were so simple that microcode wasn't necessary to cheaply implement them. Along the same lines, some early microprocessors like the 6502 also weren't microcoded (though they still tended to use large ROM lookup tables driven by a custom state machine so sort of half-microcoded).
Some very large early supercomputers were also hardwired; the CDC 6000 series, Cray-1, or IBM's largest System/370 models, for example. Microcode was avoided to minimize the logic delay during a single cycle and get things running as fast as possible. The same philosophy would resurface in the 1980s with RISC. But then RISC chips started adding microcode back in to handle complex multi-step instructions, etc.
Nearly everything else has been microcoded: everything from 8-bit microcontrollers to the Motorola 68K to the VAX, later RISC designs like POWER and ARM, every iteration of x86, etc.
The same thing happened with microprocessors. Early microprocessors didn't have enough space for microcode so they all had hard-wired control logic. It wasn't until the 8086 generation that you could put enough transistors on a chip to make microcode practical in a microprocessor. The 8087 co-processor is an interesting example. In order to fit the microcode onto the 8087 chip, Intel had to use weird semi-analog multi-sized transistors. This let them store 2 bits per transistor in the microcode ROJM.
There’s two types of microcode: vertical (the kind people think of when hearing “microcode”) and horizontal (a large PLA). So, the 6502 is not “sort of” microcoded, it is microcoded, just horizontal microcode.
A better example of horizontal microcode (and what I learned on) are the PDP-11s. A clear micro instruction address counter and the ability to jump around the micro instruction ROM, but very wide control words where it's clear 'these three signals go into this mux, this signal is an enable over here, etc'.
[*] For standard definitions of microcode.
You could use a PLA to produce exactly the same outputs from the same inputs, but it wouldn't be microcode. The PLA is "just" an optimization. But with a PLA, you can't change a micro-instruction without ripping out some of the logic and replacing it. You have a hardware implementation rather than a software implementation, which is the hallmark of microcode. In other words, changing a control signal value in the 8086's microcode is trivial, but changing a control signal in the 6502 is considerably harder.
As a concrete example, the design of the Apollo Guidance Computer is clearly microcoded, with specific microinstructions. However, the implementation is gate logic that completely obscures the microcode design. (Not a ROM implemented with NOR gates, but highly-optimized logic.) So I consider the implementation of the AGC to not be microcode.
(This is kind of splitting hairs, and I'm not really into subtle linguistic arguments.)
Even today, I would be shocked if ARM designs weren't partially microcoded as well. Something like the ERET instruction is begging to be.
In short the ARFM has decode logic that makes a bunch of clocks worth of ALU controls directly, while the x86 has decode logic that makes an index into a ROM from where that information is read/generated
"In 1951, Maurice Wilkes came up with the idea of microcode: instead of building the control circuitry from complex logic gates, the control logic could be replaced with another layer of code (i. e. microcode) stored in a special memory called a control store. To execute a machine instruction, the computer internally executes several simpler micro-instructions, specified by the microcode. In other words, microcode forms another layer between the machine instructions and the hardware."
Note the "1951".
What microcode really does is make it possible to share hardware resources inside the execution of the same instruction. Hardware is really (really) expensive, and a naive implementation of a CPU needs a lot of it that you can't really afford.
Take adders. You need adders everywhere in a CPU. You need to compute the result of an ADD instruction, sure, but you also need to compute the destination of a relative branch, or for that matter just compute the address of the next instruction to fetch. You might have segment registers like the 8086 that require an addition be performed implicitly. You might have a stack engine for interrupts that needs to pop stuff off during return. That's a lot of adders!
But wait, how about if you put just one adder on the chip, and share it serially by having each instruction run a little state machine program in "micro" code! Now you can have a cheap CPU, at the cost of an extra ROM. ROMs are, technically, a lot of transistors; but they're dense and cheap (both in dollars for discrete logic designs and in chip area for VLSI CPUs).
That is, microcode started as a size optimization for severely constrained devices. It was only much later that it was used to implement features that couldn't be pipelines. Only the last bit was a mistake.
(As an aside, the 8086 has a separate adder for address computations, independent of the ALU. You can see this in the upper left of the die photo.)
Only with some other kind of state machine, though. I was maybe being a little loose with my definition for "microcode" vs. other state paradigms.
Maybe the converse point makes more sense: "RISC", as a design philosophy, only makes sense once there's enough hardware on the chip to execute and retire every piece of an entire instruction in one cycle (not all in the same cycle, of course, and there were always instructions that broke the rules and required a stall). Lots and lots of successful CPUs (including the 8086!) were shipped without this property.
My point was just that that having 50k+ transistor budgets was a comparatively late innovation and that given the constraints of the time, microcode made a ton of sense. It's not a mistake even if it seems like it in hindsight.
I believe every other pdp-11 was microcoded.
ALSO: these days, assy lang makes it easy for compiler writers to express programs written in a high-level language!
It's software, all the way down..... almost.
If you can find a copy, reading the prior chapter on hardwired control units might be reasonable as it is somewhat "background" for following the microcode chapter.