I've always been curious about the XLAT instruction. It performs an operation like `MOV AL, [BX + AL]` (if that were a valid instruction). I assume this instruction was designed to do EBCDIC <-> ASCII conversions using a simple loop consisting of `LODSB / XLAT / STOSB` which would explain why the input and output are both hardcoded to AL. But it seems like it would have been fairly easy to allow the programmer to specify different register(s) and the instruction probably would have been a lot more useful generally if that was allowed.
It is the only place in the ISA that I know of where an 8-bit value is zero-extended to 16 bits before being added to another value. Short jumps use a sign-extended 8-bit value (and are a special case anyway since they deal with the instruction pointer). Other additions are always 8+8 or 16+16. Is there some special case in the processor to allow just AL to be zero-extended in this way? If so, that would go some way toward explaining why this instruction isn't more flexible. I realize this may be very difficult to determine working from a die photo, at least without completely reverse engineering the entire microcode ROM.
[Internally the CPU uses the "TMP" register as a temporary storage location during execution of the XCHG operation, and so the NOP will clobber whatever value was previously held there.]
One can observe two of the internal registers ("TMP" and "IND") using an invalid JMP FAR m32 opcode. Normally this is encoded as opcode FF/5 with a mod-rm byte that specifies a memory location (mod=00, mod=01, or mod=10) from which to fetch a doubleword CS:IP pointer. Specifying a mod-rm byte with mod=11 (register) - which is an undefined encoding - skips the "load doubleword from memory" microcode subroutine and causes the CPU to jump directly to the location IND:TMP instead, allowing one to observe the values held in these internal registers.
At some point I'll get around to writing this up in more detail, as I don't think anyone else has ever described the behaviour of this particular invalid 8086 opcode.
But I'll probably go back and re-test LDS/LES more thoroughly at some point just to make sure I haven't missed something [or more likely, that I'm not misinterpreting my scattered notes].
By the way, I really appreciated your detailed bus sniffer logs of the 8088 executing various instructions. It was enlightening to read through the traces and helped me understand what was going on "under the hood" of the CPU.
Glad the sniffer logs were useful! I have been using them pretty regularly for debugging and profiling things.
I'd love to see more about that --- there has been plenty of documentation on "first-byte" undocumented opcodes, but far less research and information on undocumented operand combinations/groups and such. In particular, the infamous FF FF (that gets disassembled by DEBUG and a few other assemblers as "??? DI" ) falls in that area, as does LEA with a register operand ("LEA AX, AX") and several others that don't immediately come to mind.
If I remember correctly, on the Z80 some of the hidden architectural registers' state was also visible as undefined flags. What you've described sounds very similar to that.
Yep, reading about this (from a post that I can't find at the moment) was one of the things that inspired me to try illegal/undefined mod-rm byte encodings.
I had also seen a comment (which I cannot find at the moment) describing the behaviour of LEA r16, m16 when it's given an r16 source operand (so the mod-rm byte has mod=11): the value loaded into the destination register is the value of the internal "IND" register, i.e. whatever memory address the CPU had most recently calculated.
So I tried all of the undefined mod-rm encodings (especially the instructions where the 3 'reg' bits encode a sub-opcode) and mostly just found a bunch of aliases for existing instructions. But the FE and FF opcodes -- which encode JMP and CALL -- would randomly crash my test code in strange and beautiful ways.
Single-stepping the execution of the opcode would crash the DOS 2.1 version of DEBUG.COM, so I ended up writing my own minimal INT1 handler to capture the value of CS:IP immediately after the malformed JMP or CALL, dump it to the screen, and then restore a sane CS:IP before exiting single-step mode.
My hypothesis (from a close reading of the 8086 patent) is that the group decode ROM marks which opcodes have a mod-rm byte and the microcode has a sort of subroutine call (triggered by mod != 11) that calculates the EA and leaves the result in the hidden temporary registers before falling through to the instruction-specific microcode. Skip that subroutine (using mod=11) and interesting things happen.
[1] http://www.os2museum.com/wp/undocumented-8086-opcodes-part-i...
For all practical purposes 90h is indeed a NOP, despite the fact that the CPU actually executes the opcode.
As far as XLAT, I haven't seen anything yet to answer your question. I hope to figure out the microcode, which should provide more details.
What portion of the chip handled the translation?
In any case, look elsewhere in these comments for a discussion of the NEC V33 processor, which replaced the 8086's microcode with hardwired control.
As for the 8086's microcode, each micro-instruction is 21 bits: 5 bits for a source register, 5 bits for a destination register, 3 bits for a "type" field, 7 bits for two more fields, and 1 bit to control the flags.
The microcode ROM is the big rectangle in the lower-right of the die photos.
One thing that I would love to see eventually is a series of comparisons with the 8088.
Looking at the dies, the 8086 and 8088 are completely identical over most of the die. For the most part, I can look at the 8088 die if something is unclear in my 8086 photo.
The biggest difference is the bus control logic, which is entirely different (as you'd expect). This is in the upper-right corner.
The 8088 has a 4-byte instruction queue instead of a 6-byte instruction queue, so there's a bit of wasted space in the register file there. The queue registers and control circuitry are slightly different because the 8088 needs to access half of the word at a time.
I saw a couple of microcode differences, but I don't know if they are related to 8086 vs 8088 or are bug fixes.
Some of the pin driving circuitry is different, as you'd expect, since only 8 pins are used for data.
To summarize: is the 8088 a completely redone chip? No. Are they the same chip with a jumper? No. Most parts of the 8086 and the 8088 are the same, but some parts were completely redesigned.
http://static.righto.com/images/8086-alu/die-labeled-alu.jpg
Those silicon shapes are pretty fascinating and each seem unique with respect to the other? How do they come up with these? Are these just the product of signal routing and minimal surface area needed for the circuit? Almost like a reverse puzzle?
My second question was about the right side of this photo:
http://static.righto.com/images/8086-alu/inverter-diagram.jp...
What are the purple lines and green regions here?