1984, the Year of the 32-bit Microprocessor (1984)
archive.org
archive.org
Most interesting sentence: "Currently, it is not likely that such techniques would be used for domestic surveillance."
Philosophical Issues
The specter of Big Brother may not be of concern in
Western society today, but the evolution of distributed
intelligence among machines with speech-recognition
capability certainly provides the technical base for
monitoring our activities. In fact, the U.S. National
Security Agency has developed what may be the world's most
advanced speech-recognition algorithms. This system spots
keywords in intercepted verbal transmissions from
"unfriendly" nations. Currently, it is not likely that
such techniques would be used for domestic surveillance.
But speech technologists as well as the public must be
aware of the potential loss of privacy.
[0] - https://archive.org/details/byte-magazine-1984-01/page/n213I admire the foresight of Motorola in making the ISA itself 32bit from the get go.
Huh? The second paragraph states: "First, let's define our terms. A 32-bit microprocessor has a full 32-bit architcture, a full 32-bit implementation, and a 32-bit data path (bus) to memory." (And then the author spells out what each of those three things means individually.)
IIRC, the 68000 missed both the 32-bit implementation and the data path.
Interesting to note though, that Intel's later 80386SX product would NOT quality for the article, even though it's a compatible reduction of a chip that does. (The SX version of the product had a 16-bit bus and was physically closer to a 80286 in terms of external hardware interface.)
POWER9 has an absurd setup which is 192 bits wide at the CPU (2 controllers each with 4 24 bit DMI interfaces), but each DMI has a 256 bit path t memory.
68010 had a 1 instruction loop cache, which accelerated tight loops considerably, as it wouldn't need to fetch the instruction.
Think e.g. a move with a builtin decrement then a branch if not zero.
What I do remember is that while 32-bit ISA, the data bus was 16-bit. We had both hardware and software debuggers. The hardware debugger was basically flipping a switch from going from a regular clock to a single step with set of hex displays on the front of the computer that allowed you to see what was on the databus, no internal CPU state. The software debugger allowed you to see CPU registers, display memory. No stack trace.
One of our projects was to write a resident monitor, maybe more analogous to a super micro kernel. Because the monitor ran in supervisor mode, while normal code ran in user mode, and our software debugger also ran in supervisor mode, we couldn't use the software debugger for our monitor. We could only use the hardware debugger. I got really good at decoding 68K instructions from reading the hex display while single-stepping.
A silly aside from that last comment; We built our monitor on then modern PCs, specifically, I had a laptop running FreeBSD, because it had a 68K emulator available. To get our assembled code to the micro computer, we had to transfer it over a 1200 baud serial connection, and it took around 30 minutes to transfer our monitor before we could even test. So, we heavily utilized the emulator which was nearly instantaneous. We had run all of our tests against the emulator and everything looked good. We were still in lab at around 2AM, and our professor walks in and asks how we were doing. I responded we were doing great, and had just uploaded our latest code to test. Our test on the actual hardware failed spectacularly. Started using the hardware debugger to find out why, and it turned out to be an addressing issue (the instruction was using the default 16-bit addressing instead of the intended 32-bit addressing). It was literally a 1-bit bug. I manually edited the memory and reran and it was fine. There might have also been a jokingly light back-hand across the face of the team member that wrote the routine that failed.
When System 7 came out, 32-bit addressing became available but not all software was “32-bit clean”. So you could enable or disable 32-bit addressing in the memory control panel. This is similar to the A20 gate on IBM PCs.
By 1993 or so the hardware was no longer built to support 24-bit addressing. PowerMacs followed soon after.
I recall many people running non 64-bit clean apps complaining about Catalina, which I found curious.
"Abusing" unused address bits is then just a matter of ensuring that any virtual addresses you use are restricted to the appropriate range. "x32" binaries do this in order to ensure that any pointers will fit within 32 bits even on x86-64, for example.
So it never goes onto the address bus. So it never causes an addressing issue.
The 24-bit issue of the 68000 is fundamentally different. Those tagged 32-bit pointers were used directly, but truncated by the width of the actual physical interface so they never touched the outside world.
...until Motorola released CPUs with full 32-bit address pins on the package. Oops.
Tagged pointers that don’t make it to the bus are still a compatibility problem. It means malloc() can’t return an address overlapping the tagged bits.
What really killed it though was perceived slowness. Apple insisted on shipping generations of Macs with crippled busses running at half the bit width of the CPU and/or at reduced frequencies.
I remember the very fastest Macs in the early 90s not being able to even run a scrolling 2D game at 640x480 in 8 bit color because it just wasn't a priority. Meanwhile a 486 could run DOOM in 320x240. We had to wait until the 60 MHz PPC 601 arrived, and even it struggled with 640x480 graphics (the lowest resolution available). On top of that, most software was still emulated 68000, so the perceived slowness of Macs continued until Steve Jobs returned and introduced the colored iMacs.
I eventually came to love x86 assembly for the brief time I used it though. It has so many easter egg instructions for moving bytes around wider registers and getting free side effects with memory access that it felt like there was always another way to gain a bit more performance out of hand-rolled loops.
Of course that's all gone today, because instruction sets are so.. byzantine that a good compiler will usually beat hand-rolled assembly. That's because the real processing happens as RISC beneath microcode so a human can't really know the optimal way to string instructions together to keep the pipeline full or avoid cache misses. Also the SIMD stuff expanded the solution space to such a degree that it's almost pointless to do anything directly. Better to use a vector language like MATLAB and compile to SIMD accelerated C or go to GPU processing instead IMHO. Otherwise you're perpetually struggling with premature optimization and can't work at a productive level of abstraction.
https://ia600609.us.archive.org/BookReader/BookReaderImages....
Turbo Pascal - IBM Pascal - Pascal MT+
Price: 49.95 | 300.00 | 595.00
Compile and Link Speed: 1 sec | 97 sec | 90 sec
Execution speed: 2.2 sec | 9 sec | 3 sec
Disk Space 16 bit: 33K with editor! | 300K + editor | 225K + editor
Disk space 8 bit: 28K with editor! | Not Available | 158K + editor
...
Locates Run Time errors directly in source code: YES | NO | NO
----
Extended Pascal for your IBM PC, APPLE CP/M, MSDOS, CP/M 86, CCP/M 86 or CP/M 80 computer features:
- Full screen interactive editor providing a complete menu driven program development environment.
- 11 significant digits in floating point arithmetic.
- Built-in transcendental functions.
- Dynamic strings with full set of string handling features
- Program chaining with common variables.
- Random access data files.
- Full support of operating system facilities.
- And much more.
Which unfortunately hasn't resulted in a more responsive user experience.
Faster CPUs just meant that optimising code got less important. Now, we're stuck with a bloated stack, top to bottom.
[1] https://en.wikipedia.org/wiki/Whitechapel_Computer_Works
See http://chrisacorns.computinghistory.org.uk/Computers/ACW.htm... or wikipedia.
I know of machines designed and built with the z8000, but I've only heard of the z80000 being used for embedded devices (possibly one main customer?).
Are you able to provide a reference to an actual general purpose computer built around the z80000?
Basically, the original 68000 isn't capable of atomically restarting an instruction interrupted during a memory access cycle, and so there's no way to implement a standard MMU.
So Apollo just chucked in a slave CPU that would detect the interruption of the master CPU, halt it, deal with any remapping or what have you, and then just completely reset the master CPU.
On that regard, it's a shame the Amiga released (1985) with a 68000 cpu. Particularly, move from SR became privileged 010+ and thus caused problems when Amiga finally moved past 68000. Those could have been easily avoided by releasing Amiga on 68010 to begin with, which is also a CPU with slightly higher IPC.
They compounded the issue by using 68000 again on A500/A2000 (1987), for negligible savings.
Of course, this does pale next to the gross mismanagement Commodore did of the Amiga thereon, which ultimately led to Commodore's own demise.