The Intel 8088 processor's instruction prefetch circuitry: a look inside
righto.com
righto.com
From one DRAM generation to the next the cycle time might be different for chips with the same access time. This allowed me to make the 512KB Macintosh faster than the 128KB one (http://www.merlintec.com/lsi/mac512.html).
Same thing of processor and memory speed bins.
I sincerely doubt I was the first to work that out; but I remember being so incredibly happy when I figured that one out, when it solved a problem i had.
Cannot now recall why the difference was significant, something about installing different routines for bashing serial ports i think.
If you were very lucky, some magazine might have mentioned it. Another way out was to just use disassembler if some other software package performed the same thing.
Many things are immeasurably easier than what I remember as a middle-schooler with an utterly anachronistic 286 in post-Soviet early 2000s Moscow, so that’s nice. It doesn’t make the blasted loop work, though.
(Many others are also worse. Today’s me could work the motherboard design of the 286 by looking at it, even without the manuals; my current laptop’s manufacturer’s refusal to release the schematics annoys me enough that I’ve half a mind to ask some physicists if they have a CT machine they could run the board through.)
I remember coding a game on C64 in the eighties. Just to figure out how to print the players score so that it is sufficiently fast was a challenge. Dividing by 10 with modulo to convert numbers to digits was just way too slow.
My method was not to use normal math, but to directly manipulate screen RAM characters when the score increased.
That was a very cheap way to increase the players score by say 1000 – you didn't even have to care about 3 lowest digits, just inc thousands place by 1, if it overflowed past 9, increase next position left, etc.
You know, setup a test harness. With timers and such, then walk through the cases.
Of course, that is exactly what the manufacturer should have done! I always wondered at the high errata metrics associated with some catalog parts. It is just not enough to work through the circuit and hope for the best!
Nevertheless, there have been some late variants of 486 that have been introduced after the first Pentium, in 1994 or later, and which had CPUID, e.g. the Intel 486DX4 (100 MHz).
AMD had 2 generations of 486DX4 (and of 486DX2), the first did not have CPUID (and it had a write-through cache memory), while the second had CPUID (and it had a write-back cache memory).
Some Cyrix CPUs with properties intermediate between 486 and Pentium had CPUID, but it was disabled by default and it could be enabled in the BIOS.
Measuring the length of the prefetch queue was the standard method to identify 8088 vs. 8086 and this was available in several commercial CPU detection utilities that were available for MS-DOS, e.g. in Norton Utilities or the like.
At that time I have discovered this by disassembling such a utility program.
This has forced the introduction of the snooping workaround, otherwise the stores into the data cache would not have influenced the content of the instruction cache.
> However, the 8-bit bus enabled cheaper computer hardware.
I wonder if that can be expanded on. An old (now departed) friend once said to me that this was a mirage, because the only thing that the narrower bus really bought them was that it made the 16kB configuration possible, which no-one actually bought (the minimum for using floppies was 32kB!). He claimed that the narrower bus didn't actually make the configurations with more RAM cheaper because it made them require more support chips. Is there any truth to this?
Another thing is that even if almost nobody bought the minimal RAM configuration, having a low-cost configuration can be very important from a marketing standpoint. (By the way, it's kind of amazing that the base RAM for an IBM PC was just 16 kilobytes, and now 16 gigabytes is a base RAM configuration.)
https://www.tech-insider.org/personal-computers/research/199...
On the 8086, the bottom 16 address pins (of 20) are multiplexed with data, on the 8088, only the bottom 8 address pins are multiplexed.
So you had to demultiplex the bus no matter which chip you chose. I think the cost savings mostly come from being able to configure systems with only 8 DRAM chips per bank (instead of 16 DRAM chips per bank with the 8086). A 16 bit bus also requires that your ROM chips are in pairs, and it probably increases the motherboard routing complexity. And a small bit of extra decoding logic.
A 16 bit bus would have also required IBM to skip straight over the 8 bit ISA standard with it's smaller sockets and forced up the complexity of all PC expansion cards (which would also require double the ROM chips, double the RAM chips and extra logic)
Edit: And now that I think about it, the fact that the upper bit of the address bus aren't multiplexed might actually allow you to simplify the DRAM row/column addressing logic... But only if you put the column into the upper bits of the address.
Most of the savings were not from memory but from the reuse of the existing 8-bit peripherals without additional hardware and without software changes.
On a 16-bit bus, you could connect an 8-bit peripheral to one of its halves, but then the internal registers that previously were at consecutive addresses now were spread at multiples of 2.
When doing only 8-bit transfers, you could rewrite the software drivers to use the modified addresses. However you could not do 16-bit transfers, because the bytes were no longer in the same word. This could be fixed with additional buffers, decoders and registers, to convert 16-bit transfers into pairs of 8-bit transfers on the same bus half (like 8088 did internally), but that would increase the cost.
Even if these compatibility problems were not too difficult to solve, most preferred to avoid them, in the quest for minimum cost.
"ARM has a lot more market share in people's minds than in actual numbers"
haha, that's true, especially on HN
Hell, smartphones are by far the most personal computers we have ever had.
Note the presence of the sentence: "Of course, mobile phones are almost entirely ARM."
I mean, I agree with your point; but the distinction is there. We need some way to differentiate between computers for hacking and computers that are vending machines for the hacks of others
If I write simple programs on a iPhone using Shortcuts, does it become a general purpose computer? It's programming, just with a UI and its own graphical language. How complex a program do you have to be able to "compile" in order for it to count as writing a program on device? Because there are tons of little programs being written using Shortcuts, (and also Pythonista), so you'll have to be more specific.
My install of Termux from F-Droid on my Android phone disagrees with that. While more limited than a PC, most smartphones can still be considered "user programmable computers".
If you ignore the majority of devices out there that you can write code for then that's not really a sensible definition of market share.
Let’s say even if it is possible to do it, would the resulting saving of real estate and power would be worth the effort?
1. The queue is 4 bytes long, which fits neatly in a two bit counter. You would have to switch to a three bit counter to store the 5th state, which increases the area and power usage by ~50%
2. The MT signal is explictly needed to stall the execution unit when the prefetch queue is empty. If you replaced the flag register with a 5th state of the queue counter, then you would still need combination logic to generate the MT signal (queue[0] != 1 || queue[1] != 1 || queue[2] != 1)
I'm guessing this two bit counter + MT flag scheme is actually optimal from a transistor count perspective.
The queue itself though can be in one of five states - its length can be 0, 1, 2, 3, or 4.
The difference between the position of the read and write counters (which is always available through the hardcoded XOR subtraction circuit detailed in the article) is either 0, 1, 2, or 3.
The flag allows you to tell whether the 0 result of that subtraction is a zero length queue or a full queue.