Reverse-engineering the conditional jump circuitry in the 8086 processor
righto.com
righto.com
I think a large part of it was pin count optimization; Intel had weird beliefs about keeping the pin count down. (E.g. Intel was really resistant about using even 18 pins for the 8008, which was way too few.) Another issues was the 8086 supported two different bus protocols: "minimum" and "maximum" mode, where "minimum" was straightforward and easy to use, while "maximum" provided much more information but needed another chip to decode the signals. (I don't know if anyone used "minimum" mode.) Finally, the 8087 needed a lot of information about what was going on inside the 8086's prefetch queue, so that needed more bus signals.
They might seem weird by today's standards, but by late 1970's standards, a '40 pin' package was significantly cheaper for others to incorporate into their designs than another size. There were a lot fewer "standard socket sizes" available back then, and anything that wasn't already available entailed a lot of expense to manufacture, resulting in a very high initial cost (comparatively) to recover the capital expenditure to bring the new size to production.
http://archive.computerhistory.org/resources/text/Oral_Histo...
Well, look at that. It’s this ad: https://www.flickr.com/photos/mwichary/3236356042 and I was 11 at the time.
Alternatively, the 'rep' prefix would be interesting.
So basically:
repeat: rep
cs:
scasb
jcxz done
jmp short repeat
done:Still clever, though. I guess it's to make it easier and quicker to index over words or multiples of words?
By using the same register for base and index you can also multiply one register by 3, 5 or 9.
Earlier (16 bit) x86 chips did not have the scaling feature and were limited to certain combinations of base and index (BX/BP as base, SI/DI as index), so LEA was less useful. If the registers are carefully assigned, it could still be used to do an addition and put the result into another register. Normal ALU operations always use one of the operands as their destination.
LEA also gets the "register + offset" thing done in a single instruction instead of two (MOV + ADD). It's also really easy for both assembler programmers and (dumb) compilers.
The "run in parallel" stuff is you looking at modern(ish) CPUs and thinking the original 8086 looked anything like that inside.
If one ignores the extra instructions, one gets the impression that most of the differences will be in the BIU, and some in the address adder block added to the system. But possibly that additional adder ALU can be ignored?
However the microcode format remained essentially the same[1], so I don't think there was a fundamental redesign in either the EU or BIU.
The '286 is different, and not just in the BIU (which now has to enforce segment limits etc). From what I've pieced together looking at die shots, and US patent 4442484:
There are three 6-bit fields to select registers for each micro-instruction. ALU operations can apparently take any register as operand, only immediate values have to be first loaded into a temporary register[2]. The microcode is also organized more like a conventional ROM instead of being addressed directly by opcodes.
Bytes from the prefetch queue first go through a separate decoding stage. An "entry point PLA" translates the opcode (with additional inputs for 0Fh and REP prefixes, real/protected mode, and "modr/m extended opcodes") into a microcode address. That address, any operands including a 16 bit immediate and 17(?) bit displacement field, and other flags are placed into a "decoded instruction queue" holding up to three instructions.
From what I've read the 386 was actually very similar despite adding 32 bit registers and paging. The next major changes to the microarchitecture came in the 486 and Pentium.
[1] https://news.ycombinator.com/item?id=34334799
[2] https://rep-lodsb.mataroa.blog/blog/the-286s-internal-regist...
There's no clock gating; everything gets the clock. There's not a lot to say about HALT. It stops memory operations and the prefetch queue, so the processor stops processing until it gets an interrupt.
Not hit, hlt instruction.
> There's no clock gating; everything gets the clock. There's not a lot to say about HALT. It stops memory operations and the prefetch queue, so the processor stops processing until it gets an interrupt.
Oh, ok. So ucode is just sitting in that state waiting for Q to fill in for an extra long time is all? Is the prefetch queue flushed or are the next bytes sitting in Q prefetched, simply not presented to ucode?
But then you need circuitry to process the microcode. Why is that desirable?
Is microcode basically an abstraction layer to allow a larger machine code instruction set be supported on a smaller native instruction set - and if it is why is that better than just using the smaller instruction set directly?
But CISC complexity and microcode was significantly driven by the lack of good optimizing compilers. A lot of performance critical code had to be written in assembly, and writing assembly was difficult and time consuming so it was nice to have more expressive instructions.
Early CPUs also had little or no cache and instruction fetch bandwidth was a very significant bottleneck. This was another significant driver for CISC.
So it wasn't just the case that CPU designers were idiots from the start, they did have reasonable reasons for the choices they made at the time. RISC required a certain confluence of hardware and software advancement to happen before it became the obvious or better alternative.
Interestingly things swung back the other way a decade or so later, as CPUs got vastly more complicated and capable, the ISA became relatively less important and CISCs were able to mostly catch back up to RISCs.
Also, for the original question, you really don't want to program in microcode directly, as it's kind of a mess. Micro-instructions expose a lot of ugly hardware details. Moreover, it locks you into a fixed architecture and you can't upgrade, because the microcode will change.