Inside the 8086 processor's instruction prefetch circuitry
righto.com
righto.com
In fact the Apple II was designed for 4k DRAMs with a whole-cycle time of 500ns, which was the low end of the range in 1977. The idea is that the machine double pumps a 2MHz memory bus, with cycles alternating between the 1 MHz 6502 and the video scan circuitry (which doubles as the refresh controller). By the advent of 4164 chips, DRAM cycle rates were pushing 4MHz. But Woz was stuck with more limited hardware.
Similarly the 8086 bus is multiplexed between address and data, a memory access is 4 cycles. So a 4.77 MHz IBM PC was actually cycling the DRAM at just 1.2MHz (though with no clever tricks, as MDA/CGA had physically separate framebuffers; there the extra bandwidth really was wasted).
Which isn't to say that the footnote is wrong, just that "DRAM was faster than CPUs" is only part of the story.
> the 8088 has a 4-byte prefetch queue instead of a 6-byte prefetch queue
My whole professional life, I've been led to believe that an 8088 is just an 8086 with a tiny state machine bolted in front of the bus to split multiword accesses. That it has design differences is really surprising. If the prefetch queue is smaller, then... what are they putting in that die area? Or did they redo the layout in other ways?
It was easy to measure the length of the prefetch queue by a self-modifying program, which replaced the instruction placed 4 bytes after the current instruction pointer.
On an 8088, the new instruction was executed, while on an 8086 the old instruction was executed.
Various other such obscure differences (e.g. the value of a pushed stack pointer or the result of a shift operation with a shift count greater than the register size) were used to identify NEC V20 and V30, Intel 80186, 80188, 80286 and 80386, because architectural identification means, like the CPUID instruction, have been introduced only in Intel Pentium and in the late 486 models that were launched after Pentium.
It became hard on the 386, as linked in the article: http://www.rcollins.org/secrets/PrefetchQueue.html
For detecting 8086(/88) and NEC V30(/20) vs. 80186 and anything later, I'd recommend using a shift instruction with a count of 33. The 186 and later processors will mask it with 0x1F, resulting in a shift by 1 instead. A shift count of 32 would also behave differently, but not set any flags because a shift by 0 effectively becomes a NOP.
This is true even on x86-64, only for 64 bit operands is the count masked with 0x3F.
Coincidentally a topic I did some reseach into recently. It appears to be in both the Pentium(P5/54) and PPro(P6), but not the 486.
The 8086 didn't have that problem, because it only started a new prefetch when there were two bytes free in the queue, so there was enough time for the jump microcode to stop it.
edit: already mentioned in the article
It's definitely not the case on the 486, the linked HN post written by me was incorrect about that. Tried running self-modifying code on a 486 and it required a jump to make it work.
In the US there has been a push towards teaching using modern, practical languages. It was previously more common for universities to "teach the principles." Figuring out how to apply that to the industrial zeitgeist was on the student. In 2002, the uni I went to prided itself on being allied with industry: they introduced programming with Java. The uni down the road introduced students to Gofer (Haskell circa 1992), then Pascal.
Modern x86-64 though is quite different, and knowing the 8086 ISA won't help you with following a modern assembly listing. The course material you've found is probably just really old.
No, from within the past few years; here's a random example:
https://sanjayvidhyadharan.in/Downloads/Microprocessors/Lect...
Optimized code tends to heavily use vector instructions which didn't exist in older generations, but everything else stayed the same, more or less.
It was ARM that became radically different when it moved to 64 bit.
Always enjoy your posts!
> The much-reviled solution was to create a 4-megabyte (20-bit) address space consisting of 64K segments,
Did you mean 1-megabyte here? (and again later in the same “paragraph”, pun intended)
The reason being that separating the segments into their own address spaces would create something more like a Harvard architecture, which isn't really wanted in a general-purpose computer.
It would be even more of a nightmare to program than the already janky segmented memory model we ended up with, so I'm happy IBM wisely decided to avoid it.
I ask because I'm interested (for no good reason :P) in automating 808x bus traces like http://www.jagregory.com/abrash-zen-of-asm/images/fig5.1aRT....
I have a homebrew8088 board with a V30 in it, which I've just started playing with. I've run into two issues:
(1) It runs the processor in "minimum mode", which doesn't expose as much internal/debug state on the pins. (Vs "maximum mode", where the CPU consolidates many control signals into 3 pins, which drive an Intel 8288 companion chip, which drives the control signals)
(2) To plug an intel processor in, I think I need to use a 3.3V <-> 5V shifter, and https://czh-labs.com/products/rpi-33v-to-5v-26-i-o-bidirecti... doesn't expose GPIO0/1, which the board uses for the AD5 / INTA pins.
The board is open source and I've never had a PCB made, so this might be a fun yak to shave. The Raspberry Pi doesn't have enough GPIOs to run the CPU in master mode, but I don't need to be IBM compatible, so I figure I can use an 8088 and simply steal some of the higher address pins.