An invalid 68030 instruction accidentally allowed the Mac Classic II to boot
downtowndougbrown.com
downtowndougbrown.com
Rather than a "real" instruction that CPU designers consciously created and which was meant to do something useful but wasn't documented, it could just be that this is an illegal instruction and the logic in the CPU is doing whatever it happens to do when given don't-care inputs. (Maybe this is what the author meant, and I'm just catching up.)
Normally the CPU would detect illegal instructions and cause an exception. This would mean there are certain situations where it doesn't.
I found a manual at https://www.nxp.com/docs/en/reference-manual/MC68030UM.pdf. On page 8-9 of the manual (which is page 276 in the PDF file), it says:
> An illegal instruction is an instruction that contains any bit pattern in its first word that does not correspond to the bit pattern of the first word of a valid MC68030 instruction or is a MOVEC instruction with an undefined register specification field in the first extension word.
Note "in its first word". According to the write-up, the instruction is 3 words long. The first word is normal, and the weird bits occur in the second word. So quite possibly the 68030 doesn't validate this second word, just plows forward with the logic that implements the CAS instruction, and lets whatever happens happen.
(Great write-up and amazing dedication, by the way!)
I was also taking a logic course at the time and dreaming in terms of truth tables, so it was quite a formative time in my early understanding of processor architecture.
But maybe on the 68030 in this case, the bits must be zero even if they have no documented use, because there is hardwired logic for another instruction that is activated by those bits being set, somewhat like the 6502 illegal opcodes?
http://goldencrystal.free.fr/M68kOpcodes-v2.3.pdf
It's reminiscent of ARM, but the relevant part is that the CAS instruction's second word bits 5:0 look like a "modrm" (to use the x86 terminology) where the officially documented values select only Dn, but the undocumented variant would correspond to (d16,An). At least, that's my theory for why A1 gets modified.
This seems to put something into A1 [...]If you clear bit 3 from the second word of the instruction this stops happening.
Incidentally, I remember another old "bug" in King of Fighters that "incorrectly" checked the carry flag of the SBCD instruction, which it used to decrement the round timer and end the current round. Completely undocumented of course, but if you don't emulate the arithmetic status flags when doing binary coded decimal operations, the round timer in KOF will just keep on going forever, cycling from 00 to 99 :P
SNK were really the gods of the 68000 chip.
As a fan of retro arcade machines and the 68k, I'd love to hear more about why SNK were godly in how they maximized the 68K.
You saw similar kinds of shenanigans with the 6502 chip and its "hidden opcodes" that eventually became semi-official.
The mere idea of using binary coded decimal to implement a time counter that can trivially be converted to decimal numeric sprite indices is brilliant. If SBCD had operated as described in the manual, it would have been completely inappropriate due to the extra time SBCD takes to run. But since the instruction already alters the status register, there's no need for an additional CMP, and so SBCD is faster overall, and then a simple shift-then-index is wayyyy faster than a full binary-to-decimal conversion when displaying the timer on screen (DIV performance was abysmal on the 68000).
This is how they managed to squeeze so much juice out of the chip, as evidenced in the Metal Slug series, and SNK vs Capcom https://en.wikipedia.org/wiki/SNK_vs._Capcom:_SVC_Chaos
An OS vendor wouldn't do these things, because there'd be no guarantee that a future chip (68010, 68020, 68030 etc) mightn't actually behave as the manual states, and then the code would break.
> An OS vendor wouldn't do these things...
Good observation. Writing for a closed, custom-built platform had its advantages.
While obviously a subjective judgement, a lot of people who hand coded assembler on 68k processors regard the ISA as especially elegant, powerful and fun to develop for. In many ways I think of it as peak CISC, thanks to its orthogonal instruction set and wildly flexible addressing modes. And of course the platforms which used it are legendary, from consumer (Mac/Lisa, Amiga, Atari ST, Sinclair QL) to workstations (SUN, Apollo, Quantel) to gaming (Sega Genesis, Neo Geo, Capcom, Atari, Namco, Sega, Taito, Konami) to embedded (automation, print/network controllers, synthesizers, appliances). I'm certainly biased but to this day the 68k (and its 8-bit little brother the 6809) are the only CPUs I still enjoy writing assembler on.
Or simply be an emergent but unintended behaviour of the implementation, as is the case for most of the undocumented 6502, Z80, and x86 instructions I know of.
I definitely agree...but I'd say Motorola really got carried away with those wildly flexible addressing modes. Which lead them into implementation, power draw, and gate-delay hells by the late 1980's and the 68040. The future was ever-rising transistor counts and clock speeds - and their 68k architecture just couldn't go there.
Yeah, while they could certainly be extremely powerful, I'll admit the edges of my 68000 programmer's reference card quickly got dog-eared from how often I'd need to remind myself exactly how some program-counter relative indexed redirection+offset instruction worked. Almost made me miss the days of simple 8-bit loads, stores, compares and branches being all we had.
> The future was ever-rising transistor counts and clock speeds - and their 68k architecture just couldn't go there
I've always wanted to understand more about why Motorola abandoned the 68k architecture. I understand the broad factors cited in the Wikipedia article and on RetroStackExchange but I don't recall anyone citing supporting the addressing modes specifically (though it makes sense). I never programmed x86 assembler but my sense was that ISA also had its own oddball complexities. I never understood if there was some fundamental conceptual difference between the 68k and x86 ISAs that prevented one from being able to scale into the future while the other could. Would love any more info or links if you have them handy.
https://userpages.umbc.edu/~vijay/mashey.on.risc.html
Another way to view it: Motorola did not have a senior 68k implementation engineer, who could look down the road and push back against cool- or easy-sounding ideas for making 68k programmers happy.
(I once heard that, with virtual memory, a single 68040 instruction could generate 16 page faults. No, that'll never happen in the real world - but once the spec' is final, the CPU implementation team has to lay out a chip that can handle every situation correctly. And if you're pipelining a sequence of "tough" instructions against corner-case data - yeah, that can be factorial hell.)
> Motorola did not have a senior 68k implementation engineer, who could look down the road...
Yeah, this makes sense. Having started out as a complete newbie user and fanboy on 8-bit micros and then leaping to the brand new Amiga 1000 as my first 68k (because it was just so awesome), I've realized the perspective I had back then on Moto and the 68k line wasn't very complete. Reading some of the first-hand oral history from insiders in recent years shows that Motorola management made key strategic mistakes like not realizing the potential of what they had at various inflection points. The 68k was created by a new team with very little experience but which had some unusually brilliant people on it. That yielded a bold and expansive design with lots of deeply powerful aspects (like that addressing) but it may have been "too expansive" (or maybe over-complete) for what would be the first part in a long product line.
Moto also didn't seem to realize they were in an all-out, high-stakes drag race to advance silicon fabrication faster and farther than their competitors. Moto was quite the laggard both in leading edge process technology and in perfecting their leading edge manufacturing reliability/predictability. This may have just been due to Moto being a huge conglomerate with lots of divergent businesses, like selling millions of 8-bit 6803 derivatives to General Motors every year. Whereas in the 1980s, Intel could still adopt a startup mindset and singular focus when they decided it was crucial. Maybe that's the over-arching meta here. Intel bet the company on figuring out some way to scale the x86 into the future and Moto treated the 68k CPUs like they were just another line of business.
Personally, I now think of the 68k line as a beautiful swan born to distracted, mediocre parents who never really understood its potential, while the x86 was a bit of an ugly duckling born to committed, passionate, smart parents who were determined to find a way to make it successful. Maybe things would have turned out differently if the Moto board of directors knew that in 30 years the most valuable corporations in the world would all be based around silicon fabrication and IP. :-)
http://ref.x86asm.net/geek.html
https://gist.github.com/seanjensengrey/f971c20d05d4d0efc0781...
We don’t really know the exact details of what this instruction does. With some limited testing, I believe I’ve observed that the resulting value of A1 depends on the original A1 value, the value of A7, and the program counter. But I’m not sure. Maybe someone can make a program that tries out a bunch of different register values and memory contents, and attempt to deduce what exactly the instruction does so that it can be emulated accurately. Until someone decides that it’s worth trying to figure out, MAME is patching this bug out of the ROM in order to allow the Classic II to boot.
IMHO this is definitely worth figuring out for accurate emulation. I'm not familiar with 68k but the bits in the instruction offer a good clue - my theory is that bits 5:3 of the 2nd word seem like another mode field, and instead of selecting one of the Dn registers via mode 000, 101 is selecting (d16, An) again and the Dc field, containing 001, is being interpreted as A1.
We were still bothering with demoscene stuff in 8 bit home computers, and those of us busy with 16 bit home systems were focused on Atari and Amiga systems.
PC and x86 at home only took off, meany really taking off among demoscene and other home users, was when VGA and sound cards became part of a standard PC.
Amazing work. Thanks for the exposition.
(I miss the 68000 line. Those were such great chips...)
MIPS carried the flag, and now RISC-V is even more readable.
The dusk of x86 era is nigh.
<https://github.com/jjuran/metamage_1/blob/master/mac/toolcha...>
Though I'm not denying that some of the newer macOS APIs are very poorly documented. As in, you know you've stumbled upon the cool shit when you end up on one of those old pages with a blue gradient in the header that says "Apple documentation archive".
Reading from a jump table with an index that's too big is a realistic sort of bug to have, so I could see that part making it into modern shipped software. But I would expect the process to fall over when it happens, not keep on trucking like it did here.
FWIW, WebAssembly is an environment where bugs of this sort are more possible, since it has a single linear address space where every address is both readable and writable. So if your garbage address is within range you can do an erroneous read, write or CAS and get away with it. But then invalid instructions like in the post will cause the WASM module to fail to load, so it's still not 1:1 comparable with this issue in the mac's ROM.
(ARM and RISC-V do, anyway)
It did cause a system error the first time I stepped through the instruction with MacsBug on my LC 475, but then it was fine after that.
For a quick test/code loop nothing beats the Asm-* family - remember to save often and make backups though :)
Another modern/maintained one, focusing on low end 68000/68010 machines, is asmtwo.
AsmOne was a very popular assembler at the time, and has seen a few derivatives. I remember an old one called trash'm'one.
XC68LC040RC25B
02E23G QEDP9348D MALAYSIA
Love the work you do Ken. Thank you.
I think that to be a perfect article, they should wrote :
By the magic of buying a Classic II and hacking the ROM...
Another possibility is that is a special institution in the chip specifically for Apple that again was used as a copy write detection or protection scheme.
The system booted in spite of that undocumented instruction. When things work, you don't go looking for undocumented things that are contributing to the working state.
Millions of C programs work accidentally, in spite of undefined behavior. Nothing gets investigated until a compiler change triggers something.
Of course they would have fixed it if it had prevented the system from booting, I even said that in the article. I still think the odds of what happened here were pretty small. That's what I meant by miraculous.