IN/MSX: Running 4 Copies of an Operating System at Once (2008)
jeff-barr.com
jeff-barr.com
At Imagen (a Stanford TeX project spin-off started by Knuth's sidekick Luis Trabb Pardo, building the first typesetting-capable commercial laser printers using, at first, wet-process Canon imaging engines (LBP-10)) in the early 80's, we used the same Sun board (Andy Bechtolsheim, the designer, was a consultant for us while he started up Sun).
I wrote our own "real-time OS" on the bare 68K Sun hardware (first time I'd ever written a full (if simple) OS from scratch), and remember fairly vividly the hard-knocks learning experience about race conditions just like the one he describes here. Running for hours or days without error and then crashing randomly--nightmare time.
Luckily, we also had an ace hardware guy, Kok Chen, from the Stanford SETI project, and he and I and the logic analyzer would run test setup and lie in wait for the condition to show up, then look back in time at all the (Multi)bus transactions to see what actually happened. (Kok later moved to Apple and became a distinguished engineer, one of very few folks who could work on whatever they wanted.)
I was 19 back then and I didn't know much (I learned a lot during that assignment) about OS design and concurrency, I remember spending a few nights just wondering why my code wasn't working at times. I mean, after all it's just two instructions one next to the other, right? It's too unlikely that a context switch were to happen right between these two lines of code, right? How wrong I was.
Awesome read, by the way.
Oh boy, I hope to use this quote at every possible opportunity from now on.
Also, ROS worked within a flat address space and had no awareness of the memory management features of the SUN board. Unix, on the other hand, took full advantage of the MMU and I had no way to control what it did.
I also seem to remember that there were 68000 boards that "solved" this by actually including two 68000's on the board and running them in lock-step 2 cycles apart, or something, and then triggering a trap on both of them if the leading one triggered the MMU - that way the trailing CPU would be halted at a point where sufficient state was available to reset state on both of them..
Do you remember if you had to deal with anything like this? Or did it have simple enough requirements to get away without it (or actually run a 68010)?
I love the 68k family - after my C64, I got into Amiga's, and I'm still disappointed Motorola didn't manage to keep up. It was so much cleaner to work with than the x86...
The 68K ISA was a joy to program. With some prior experience on the 6502, I was productive on the 68K within days.