The microcode and hardware in the 8086 processor that perform string operations
righto.com
righto.com
This is true even for later x86 processors that move entire cachelines at a time, and also leads to one sneaky technique for detection of VMs or being debugged/emulated:
https://repzret.org/p/rep-prefix-and-detecting-valgrind/
https://silviocesare.wordpress.com/2009/02/02/anti-debugging...
The following should load AL with 0 on an 8088, 1 on an 8086:
mov bx,test_instr+1
mov byte [bx],1 ;make sure immediate operand is set to one
mov cl,255 ;max shift count to give time for prefetch
cli ;no interrupts allowed
align 2
shl byte [bx],cl ;clears byte at [bx] after wasting lots of time
nop ;first byte in queue
nop ;2nd
nop ;3rd
sti ;4th (intr. recognized after next instruction)
test_instr:
mov al,0ffh ;opcode is 5th, operand is 6th
The 255x left shift is done in an internal register, during which the 8086 has more than enough time to fetch all six following bytes. I didn't test this code but am 99% certain it works.Sort of a self-deprecating name I chose. But it's also an interesting instruction, in that Intel had to actually write microcode specifically to support it in every x86 compatible chip going forward.
It loads CX bytes from memory at DS:SI into the accumulator. Overwriting whatever was there previously every time.
The only uses I can think of:
- test memory speed
- if CX is known to be 0 or 1, conditionally read one byte
- trigger some I/O device where the value read isn't importantSCAS is the one that scans for e.g. a nul byte.
ldy #len
loop:
lda source-1,y ; memory accesses: loads opcode, address, value
sta dest-1,y ; memory accesses: loads opcode, address, stores value
dey ; memory accesses: loads opcode
bne loop ; memory accesses: loads opcode, offset
So 4 instructions and a minimum of 9 memory accesses per byte copied, more if not in the zero page. Even unrolling the loop gets you down to 6.Compare this to the 8086:
mov cx, len
mov si, source
mov di, dest
rep movsb ; memory accesses: read the opcodes, then 1 load and 1 store per byte
Even forgetting about word moves, you're down to 2 memory accesses per byte. No instruction opcode or branching overhead.[Don't forget the decimal mode bit in the 6502's status register :-)]
The 8086 D flag is the direction for string instructions, whether to increment or decrement the index registers after each repetition. One use was for moving memory between overlapping regions by starting at the end and working backwards.
The 6502 D flag is for decimal mode. When set, the instructions ADC and SBC do BCD arithmetic instead of binary, e.g. 0x09 + 0x01 = 0x10 instead of 0x0A as you'd expect. The 8086 also has support for BCD, but took a different approach, using separate instructions DAA/DAS for the same purpose.
Both can lead to bugs when you assume the flag is in the "normal" state but it somehow got flipped.
Transfer Alternate Increment (TAI), Transfer Increment Alternate (TIA), Transfer Decrement Decrement (TDD), Transfer Increment Increment (TII)
For contrast 5 years older 80286 already did 'rep movsw' at afaik 2 cycles per byte. 6 years later Pentium did 'rep movsd' at 4 bytes per cycle. Nowadays Cannonlake can do 'rep movsb' full cachelines at a time at full cache/memory controller speed.
When moving data PUSH/POP were slightly slower than using LDIR/LDDR, though.
So out of the 24 cycles per iteration, there's 2 useful ones: 1 read, and 1 write. You need to unroll this sort of loop a few times to get anything near 9 cycles per byte, and ~80% of the cycles are still going to be kind of wasted.
(Still annoys me to this day if I find myself writing code to copy stuff around. Why isn't it in the right place already?!)
What is 'special' for the 8086 is that this was a low cost microprocessor using microcode, as opposed to a multi-million dollar mainframe CPU.
[1] https://en.wikipedia.org/wiki/Microcode#History
edit: fix typo
You can view microcode as one of the many technologies that started in mainframes, moved into minicomputers, and then moved into microprocessors. In a sense, the idea behind RISC was the recognition that this natural progression was not necessarily the best way to build microprocessors.
The concept of the ISA as a standard implemented by hardware, as opposed to the ISA being a document of what a specific hardware design does, goes back to the IBM System/360 family of computers from 1964. That was the first range of computers, from large to (let's be honest) somewhat less large, that could all run most of the same software, modulo some specialized instructions and peripheral hardware like disk drives not present on the cheapest systems, at different levels of performance. As others have said, microcode is plenty old, but the idea of architectural families is one of the things which really gets us to where we are now.
See chapter 8: http://www.bitsavers.org/components/ti/TMS340xx/TMS34082_Des...
And that standard PC's DMA controller was originally intended for use with the 8-bit 8080 (eighty-eighty, not eighty-eight) processor. It was slow and only supported 16 bit addresses. IBM added a latch that would store the high address bits, but they weren't incremented when the lower bits wrapped around. So buffers had to be allocated in a way that they didn't cross those boundaries, and the BIOS disk functions would return an error if your code didn't do this properly.
A marvelous machine. High resolution (if monochrome), Hercules-like graphics, for example.
But the kicker was that it actually had an MMU[1], built from discrete logic chips, in order to run SINIX (a Xenix 8086 derivate, but with added MMU support among other things) and enjoy full memory protection. I reverse engineered that MMU, it's quite fun. Paging wasn't possible though.
[1] Early on as a variant called "PC-X", but they apparently did away with the early board revisions, and all machines I've seen say "PC-D/PC-X" on the back, so they all have the MMU and can be both.
Nevertheless, due to their high cost for those times none of these 2 CPUs had seen much use before two years and a half later, when IBM launched the IBM PC/AT in August 1984.
Most programmers have encountered INS and OUTS for the first time while using PC/AT clones with 80286. Embedded computers with 80186 became common only some years later, when the price of 80186 dropped a lot.
However they have been launched only in 1984 and computers made with them have become available only some years after the PC/AT with 80286, most likely even later than the launch of 80386, so despite that they implemented all the 80186 extensions they do not count when discussing the timeline of the INS/OUTS introduction.
[1] I can't take much credit. The code was Bresenham's algorithm, as explained to me by the OS/2 group. At the time, OS/2 was the cool operating system that was the wave of the future and Windows was this obscure operating system that nobody cared about. The impression I got was that the OS/2 people didn't really want to waste their time talking to us.
But, what if the overlap had been just 10 bits, or just 8, leaving a much larger functional address range before we got clobbered by the 20-bit limit; what if it was 22 or 24 bits of useful range? Can you speculate what effect such a decision would've had, and why it wasn't taken at the time? I understand in 1978 even 20 bits was a lot, but was that optimal?
They already multiplex the 16-bit data bus on top of 16 of the 20 pins of the address bus. But with only 40 pins, that 20-bit address bus was already using half of the total chip pinout.
And, at the time the 8086 was being designed (circa. 1976-1978) with other microprocessors of the time having 16-bit or smaller address buses, the jump to 1M of possible total address space was likely seen as huge. We look back now, comfortable with 12+GB of RAM as a common size, and see 1M as small. But when your common size is 16-64k, having 1M as a possibility would likely seem huge.
For the longest time, Intel was fixated on 16-pin chips, which is why the Intel 4004 processor was crammed into a 16-pin package. The 8008 designers were lucky that they were allowed 18 pins. The Oral History of Federico Faggin [1] describes how 16-pin packages were a completely silly requirement, but the "God-given 16 pins" was like a religion at Intel. He hated this requirement because it was throwing away performance. When Intel was forced to 18 pins by the 1103 memory chip, it "was like the sky had dropped from heaven" and he had "never seen so many long faces at Intel."
[1] pages 55-56 of http://archive.computerhistory.org/resources/text/Oral_Histo...
I'm sure there must be architectures that use a multiplexed low/high address bus, like a latch signal that says "the 16 bits on the address bus right now are the segment", then a moment later "okay here's the offset", and leave it to the decoding circuitry on the motherboard, to determine how much to overlap them, or not at all. Doing it this way, you could scale the same chip from 64kB to 4GB, and the decision would be part of the system architecture, rather than the processor. (Could even have mode bits, like the infamous A20 Gate, that would vary the offset and thus the addressable space...)
But, yeah, it was surely seen as unnecessary at the time. Nobody was expecting the x86 to spawn forty-plus years of descendants, and even though Moore's Law was over a decade old at the time, it seems like nobody was wrapping their head around its full implications.
Most modern architectures do that, see https://en.wikipedia.org/wiki/SDRAM#Control_signals for an older example (DDR4 and DDR5 are more complicated). The high address bits are sent first (RAS is active), then the low address bits are sent (CAS is active).
The hard part is figuring out how to avoid slowing down the chip while you're waiting for both parts of the address. Mostek figured out how to implement the timing so the chip can access the row of memory cells while the column address is getting loaded. This required a bunch of clock generating circuitry on the chip, so it wasn't trivial.
I discuss this in more detail in one of my blog posts: https://www.righto.com/2020/11/reverse-engineering-classic-m...
"The processor was to be assembly-language-level-compatible with the 8080 so that existing 8080 software could be reassembled and correctly executed on the 8086. To allow for this, the 8080 register set and instruction set appear as logical subsets of the 8086 registers and instructions."
The segment registers provide a way to address more than a max of 64k, while also maintaining this "assembly-language-level-compatibl[ity]" with existing 8080 programs.
"Various alternatives for extending the 8080 address space were considered. One such alternative consisted of appending 8 rather than 4 low-order zero bits to the contents of a segment register, thereby providing a 24-bit physical address capable of addressing up to 16 megabytes of memory. This was rejected for the following reasons:
Segments would be forced to start on 256-byte boundaries, resulting in excessive memory fragmentation.
The 4 additional pins that would be required on the chip were not available.
It was felt that a 1-megabyte address space was sufficient. "
Ref: page 16 of https://www.stevemorse.org/8086history/8086history.pdf
Credit to mschaef for finding this.
If only there was a famous Bill Gates quote, about how 16MB (the limit of 24-bit physical addresses) ought to be more than enough...
Well, they were dead wrong about /that/.
With the REP prefix (another single byte), an entire loop could run in microcode without any additional instruction fetches. Remember that each memory access took 4 clock cycles and there was no cache yet.
--
Eliminating opcode fetches might still speed things up today in some situations, but modern x86 cores are optimized to decode and dispatch several simple operations each cycle without having to go through a microcode ROM (and thus have to use a slower path for the complex instructions).
Also the fact that compilers didn't emit most of the more complex/specialized instructions led to Intel not spending much effort on optimizing those.
I believe some compilers will still emit rep movs for memcpy, rep stos for memset, rep cmps for memcmp, and sometimes even rep scas for memchr/strlen if given the appropriate options.
(Microcode also makes it possible to design in bug-patching features. Possibly including patches for bugs in your direct-implemented instructions.)
SCAS compares AX against [ES:DI] while CMPS compares [DS:SI] against [ES:DI], so the paragraph before is slightly incorrect (should be "SCAS compares against the character in the accumulator, while CMPS reads the comparison character from SI").
sed "s/4-/1-/", perhaps?
The only thing that comes to mind is fetching vectors from the IDT. I guess it simply output CS in those cycles?
It's kind of surprising to me that it's the address adder that increments/decrements the index registers. Since it's a direct operation on a register, naïvely I would have assumed it's the ALU doing that. Is it busy during that time?
EDIT: Having read a bit further, I guess it's because of the microcode's simplicity... so, with the microcode being as it is, using the BL flag to instruct the logic to use the address adder to increment/decrement the index register prevents needing additional ALU uops, which would make the instructions slower? So then I guess it would not have worked to make that BL flag instruct the ALU instead? Does anyone know any details?
> IND → SI JMPS X0 8 test instruction bit 3: jump if LODS
> DI → IND W DA,BL MOVS path: write to DI
Is "W DA,BL" a typo, should it be "W DS,BL"?
The names are those used in a patent filed by Intel and come from before the official documentation for the 8086 was written. It might make the microcode a bit more readable to invent a different notation that is more in line with normal x86 assembly, but then again this is a very niche topic :)
Andrew Jenner's microcode disassembly used the original names. I've been changing the register names to the modern names, but missed that one. So it should probably be "W ES,BL", although I don't really like "BL" either. (Maybe "PM" would make more sense for plus/minus?)
A bit more on the original register names. The original 8080-inspired registers XA, BC, DE, HL were renamed to AX, CX, DX, BX. The original registers MP, IJ, IK were renamed to BP, SI, DI.
There's a diagram showing the original registers in the patent: https://patents.google.com/patent/US4449184A
Another thing I just noticed:
> Apparently a comprehensive solution (e.g. counting the prefixes or providing a buffer to hold prefixed during an interrupt) was impractical for the 8086.
Did you have a particular kind of buffer in mind? The buffer could not be internal (i.e. not exposed to the programmer), because the interrupt handler itself could use string instructions. Or, it could just switch over to some other code entirely, and switch back much later.
I guess since interrupts already push not just the usual CS:IP on the stack like a far call would, but also the FLAGS register, so that the special IRET could restore it, the x86 could also have pushed additional data describing the "prefix state"?
Or maybe even try to cram it into the FLAGS bits that were still unused originally? With the 7 spare bits of the 8086, the 4 segment prefixes (of which only can be active at any time, so 2 bits), plus LOCK, would fit and leave 4 bits to spare.
It would be interesting to find out what the real solution was in later x86 processors. I'm sure someone has looked into this.
** 0.1010101x STOS
0 IK -> IND 0 F1 6 ;check count if repeated
1 M -> OPR 6 W DA,BL
2 1 DEC tmpa
3 IND -> IK 0 NF1 7 ;exit if no repeat
4 SIGMA -> BC 5 INT RPTI ;handle interrupt
5 BC -> tmpa 0 NZ 1 ;redundant check???
6 BC -> tmpa 0 NBCZ 1 ;loop while Not BC Zero
7 4 none RNI
** 1.10100100 string instr interrupted (RPTI:)
0 PC -> tmpb 4 none SUSP
1 1 DEC tmpb
2 SIGMA -> tmpb 0 F1 5
3 0 PFX 6
4 SIGMA -> PC 4 FLUSH RNI ;PC corrected
5 SIGMA -> tmpb 0 UNC 3 ;has REP
6 SIGMA -> tmpb 0 UNC 4 ;has seg pfx
7 4 none none
edit- Posting code for STOS since it's simple, even though segment prefixes don't affect its function.
- No idea if line 5 could be removed, maybe needed for timing?
- MOVS, INS and OUTS use a special memory-to-memory/IO operation encoded as 'W IRQ DD,BL'. It doesn't use the IND register at all, maybe the internal DMA controller instead?
The 8086 patent also only included a few examples, and the naming conventions for microcode registers and instructions. The full microcode for that one was disassembled by Andrew Jenner.
It would be interesting to know how they sped up multiplication and division between the 8086 and 80186!
At least now the solution is to simply pass the full PC of the instruction thorough the pipeline. Or at the very least mark the PC for each ROB entry, and offsets for each instruction in the ROB entry.
As an aside do you know yet what you're going to work on once you've mined out the 8086?
It also uses microcode and Ken has not surprisingly written about it too: http://www.righto.com/2016/02/reverse-engineering-arm1-proce...
Yeah, it literally has RISC in the name, but the SuperFX is called RISC by it's creators as well, but it's about as far aawy as you can get. RISC was a buzzword to the point that it's still a bit muddled to this day 40 years later.