A Great Old-Timey Game-Programming Hack
blog.moertel.com
blog.moertel.com
The problem was that this was terribly slow, it flickered like crazy and it was unplayable. I was very sad because my game was working but unplayable for anybody so I tried to engineer a way to make it stop flickering. The solution came when I found out about a couple of functions in pascal that let you clear a specific character in the console at a specific X,Y coordinate and write another character that that coordinate. What I ended up doing was keep track of all the changes in the game for each frame (snake movements, food position) and just re-draw only the portions of screen that had changed.
This was great, no more flickering and the game was playable. (Nobody really played it because nobody cared but I was really proud of it).
Found out years later that this approach is pretty much what Carmack did in his old games: Adaptive Tile Refresh[1]
Bill Gates: "... We were also fascinated by dedicated word processors from Wang, because we believed that general-purpose machines could do that just as well. That's why, when it came time to design the keyboard for the IBM PC, we put the funny Wang character set into the machine—you know, smiley faces and boxes and triangles and stuff. We were thinking we'd like to do a clone of Wang word-processing software someday."[1]
Like it sounds, the game went through a lot of changes like this in structure through the course of development, where the nature of the game had entirely changed. I had a number of friends with my model of calculator, which was not the more-common TI-86 at my school, so I was the only, err, game in town. Because it was in-effect several games over its lifetime, and different people wanted different versions of it, I learnt a lot about maintainability and version management. "You have the six-month-old version of this with food, and it has a bug that I fixed five months ago in a different game. And the fix doesn't apply here". Or "I solved the slowness you're seeing in this case by removing 3 instructions from inside of this loop here, but that cut the key response time in half, is that alright?"
It was a major learning experience for me. My calculator is long-dead and I don't have a version of the game anymore, but I bet that I could still remember how to make changes to it if I saw the code again.
I also recall writing minesweeper (this time on a TI voyage 200), also in TI-basic. The main difficulty, of course, was that when you hit an empty square you have to recursively uncover all adjacent empty squares. So I wrote this recursively, but on this TI there were local variables to recursive calls which also meant that the calculator would sometimes run out of memory when doing its recursive calls (at vaguely unpredictable intervals). Fortunately, I discovered that this could be caught with the rudimentary exception-handling mechanism provided, so I would use this to explore not everything but a small radius around the selected position. The puzzling thing, however, was that the radius was uneven, and got smaller and smaller (as if some memory was being leaked by the interpreter), and more surprisingly, in the graphics view that I used, whenever the exception handling would trigger the TI would display a small warning "differentiating an equation may yield a false equation". To this day I have no idea why.
It all sounds silly now but I learnt a lot typing naive code for a slow, buggy, and limited interpreted, on a cramped calculator keyboard, during boring classes. If smartphones had existed back then, I wonder what it would have changed.
What Carmack achieved (by using the hardware features of the graphic cards of that day) is properly described in the link you give: smooth side scrolling on the PC. The kind of scrolling seen on the consoles of that day, which had a hardware support, or in the famous http://en.wikipedia.org/wiki/Super_Mario_Bros (a copy of which is the unreleased "Dangerous Dave in Copyright Infringement" made by Carmack). That has nothing to do with drawing only the changed characters with simple print at the given screen coordinates, what you did.
To achieve that, it used exactly hardware feature of changing the origin of the area in memory that is going to be picked up by the graphic chip to be presented on the screen. That changing was made in one pixel increments. Without it, no smooth scrolling. Second, he redraw the tiles, he didn't print the characters.
So no, that is absolutely not "pretty much what Carmack did in his old games." What you did at the end was done by absolutely every game at these times (using the commands of cursor movements to position the characters then printing them, in Turbo Pascal using gotoXY from the console library), what you started with (reprinting the whole screen just because) was simply as "pessimized" approach as it can be made at all. You didn't even optimize anything, you did what everybody did, you just started with the "pessimization." I'm old enough to remember. There is also at least somebody here who can provide a link to any computer magazine which printed such games in a few lines of code (at that time sources were made small enough to be printed in the magazines).
For the reference, at least the terminals in 1973 already had "Direct Cursor Addressing" -- the technology behind the gotoxy you used:
http://vt100.net/docs/vt05-rm/chapter3.html#S3.8
(It was already standardized anyway!) That terminal had 1440 characters, all directly addressable, and the maximum communication speed was 2400 bps which gives 300 characters per second or almost 5 seconds to push the whole screen to it. If you moved for example 10 characters, each with a 5 byte sequence, you needed only 50 bytes, so you can change their position 6 times per second. Everybody at that time knew that stuff. It was slow enough to see.
And yes, I did that while I was in highschool, which is still years after Carmack did his deed. It was around uh... 2005?
I thought I'd just share an interesting "hack" I did when I was a teenager which, to me, felt like a real genius trick to smooth gameplay. Back then my code was literally a lot of messy spaghetti goto hell, so you can understand how much of a novice I was.
I recommend re-reading what I said, I never claimed I "pretty much" did what Carmack did, just that the idea behind incremental upgrades was "pretty much" similar to Carmack's genius. No more, no less.
Let me rephrase in less words, it is absolutely not pretty much what Carmack did. You just "rediscovered" cursor positioning feature introduced eons ago on the character-based terminals. Carmack did something else and with the bit-mapped screens using the specifics of the graphics adapters of that time. The tracking you try to refer to he did on the scrollable pixel addresses, you just did simple character prints for $deity sake!
For the explanation of my opinion, see my previous post.
That's all I did, really. Not pretending I did anything more complex than that, I didn't even have access to a framebuffer or anything... Let's say the "principle" was the same, the delivery far from it.
Your school book is the equivalent to the Turbo Pascal on-line help (or TP book if you used TP 3). The similarity of the formulas is the similarity of your implementation and Carmack's "being first." Having your post the top of HN discussion of one real hack makes me deeply sad as it obviously reflects there are some readers who support your view of "the same principle." Or just don't understand, in my view.
My claim is that anybody who thinks moving the text cursor is "the same principle" as Carmack's one pixel at the time smooth scrolling doesn't have elementary understanding of both. And I explained the difference.
I have nothing against morgawr writing about his discovering cursor positioning. I have against supporting him equating this with something completely different. If a five year old goes out of the go cart and says "ma I can ride the plane" nobody cares. If he's 20 something and claims the same after the same "feat" on the pilot forum, some pilot can try to correct him, even if it's the lost time. Even if he also knows that there's the famous "don't argue with .... People might not know the difference."
I still think it was worth writing these comments.
The cunning part of the Carmack games was presumably widening the screen - I don't think I ever saw this done on the BBC. But as a schoolboy in 1989/1990 I'd experimented with reducing the BBC's screen height by one character row (by tweaking R6) to avoid flicker when scrolling vertically. Same principle. It's a fairly obvious thing to do, once you've figured out what's going on.
I think one can come up with better evidence for John Carmack's uniquely amazing skills.
What in the world has that to do with Carmack's algorithm of updating Mario after one pixel hardware scroll (after which he has to redraw the "unmoving" part fully) and adding more background once he spends the hardware moves?
This is someone who found out from Romero exactly what the Keen games did: http://www.vintage-computer.com/vcforum/showthread.php?30619...
Or maybe you're just having a bad day, in which case, hope it improves, chill, try not to get caught up in arguments when you're already pissed off (I know I do that and come off as a jerk sometimes).
Why are you angry at him? I thought the main point of the fine article wasn't just about the technical details, but also the fact that they wouldn't give up. Would you rather that he had seen his flickering game and just say 'Aw screw it, programming is crappy!' and walk away?
Or do you want a signed and printed statement of your 'absolute' rightness, a full retraction from him never to compare himself with Carmack again and pg to come into this thread and strike down his comment from the top and relegate it to the bottom of the page until he spends years learning what people knew 30 years ago?
Most of the people here understood what you were discussing. My first computer was a Sinclair ZX81, and I remember well those magazines you brought up.
But you were really rude the way you were discussing it, and quite frankly, you were calling him names, unless you think that delusional is a nice way to refer to someone.
And yes, my email is in my profile if you do want to email me privately and call me names.
Because right now, that's what's happening. And unless this is a particularly off day, I'd be astonished if you haven't had this reflected to you before by friends, family, and/or co-workers.
You're 100% right on the technical merits. Everybody here understands what you're saying and agrees with you — nobody cares. It seems like you're going out of your way to humiliate the original commenter for no good reason, e.g., by calling him delusional, even after he admitted that he understood what he did in high school in no way rose to the level of what Carmack did.
The OC was describing the sensation vindication and pride he felt when he figured something out on his own and later learned that a more advanced version of same idea already had a name. This is a feeling most programmers can relate to, I think. I only learned what a "trie" was, for example, years after I implemented a shitty version on my own.
Whatever your intentions are, you have to understand the impact your words have. And in this thread, at least, it's not pretty.
clrscr();
draw();
This redraws twice, by first clearing and then overwriting again. The flickering should go away completely if you actually redraw the screen, e.g. gotoxy(1, 1);
draw();
At least in the things I wrote in Turbo Pascal back in the day this was enough to fix performance problems from drawing too. gotoxy() calls cost time too and simply writing a large bunch of characters in one go is often fast enough. Of course, writing directly to video memory is by far the fastest approach.The last time I had fun with this was with FreePascal in my first semester with a friend. We wrote a text-mode UI library for one assignment (it wasn't strictly necessary, but it was a fun learning experience); it had a message loop that handled focus and dispatched input to controls, full-screen "windows", sub-screen dialogs, and buttons, text input boxes, combo boxes, and tabular view. We used gotoxy() and write() calls throughout, which had the benefit that we could test it on my friend's old VT-100 terminal as well (it was slow as hell, though, but you could literally see where to optimize draw calls, similar to using X remotely over a slow network where you can watch it draw each menu half a dozen times), but planned on moving to a more optimized strategy. Alas, we never really had the time for a v2 rewrite so the project just went to die. And by now it's probably of no use to anyone anyway (but it taught me a lot about how windowing systems work or at least can work).
It's definitely one of those tricks that comes in handy now and again :).
I like this line right here. It does seem like we've piled on abstraction after abstraction in these days. Sure this does make things easier, but I think things have gotten so complex that it's much harder to have a complete mental model of what your code is actually doing than in the simpler machines of the past.
And I'd say today there's a tendency to overabstract and overdesign stuff.
In today's world of "everything fits", some modern developers don't try to "bottom align" their code, but rather "top align". No, a hello world app doesn't need an XML config file.
Of course, not being constrained by machine limitation anymore (unless you're doing video/simulations/games etc) is a blessing
Here's one story on it: http://plumecopy.com/bach-picasso-on-the-creative-process/
I'm not entirely sure it's all nostalgia; the c64's sound chip was a bit more flexible than the NES's, and there was a real culture of musicians on the c64, so I may have just gotten used to a higher level of musicianship as a base.
It's only due to the deficiencies within certain abstractions that it frequently becomes necessary to subordinate and go lower down the chain.
[1] I was basically sharing buffers between SPI send/recv interrupts and the code that used them, so I needed a way to lock buffers. I later replaced the code with much simpler, much slower synchronous code.. :-/
I lost myself there, but my main point is: in electronics (embedded systems, mainly) all this beautiful joy of crazy optimizations is still alive :D
The right way to calculate this figure is (t1 - t0)/t0, rather than the author's formula which seems to be (t1 - t0)/t1. For instance: (157 - 98)/98 = 60%, but the actual amount is (157 - 98)/157 = a 38% speed up. A heuristic: 60% of 157 will be much more than 60 (since 60% of 100 = 60), which means a 60% speed up would reduce the speed to below 97 cycles.
It gets even more misleading the more efficient it gets: Adding up the cycles, the total was just 1689. I had shaved almost 1200 cycles off of my friend’s code. That’s a 70 percent speed-up! The author has 1200/1689 = 71%, but the correct numbers yield 1200/(1689+1200) = 42%.
Not that I don't think these are significant gains, but it's just misleading to label them like this. If you've removed less than half the cycles, there's no way you've seen a 70% speed up.
By my calculations (which I'll machine-check now with Maxima), the 30% and 70% speed-up claims are sound. First, let's fire up Maxima:
$ maxima
Maxima 5.30.0 http://maxima.sourceforge.net
using Lisp SBCL 1.1.8-2.fc19
Distributed under the GNU Public License. See the file COPYING.
Dedicated to the memory of William Schelter.
The function bug_report() provides bug reporting information.
Now let's calculate how much faster we made the row-copy code by unrolling its loop: (%i1) orig_loop_speed: (1*row) / (157*cycle)$
(%i2) unrolled_loop_speed: (1*row) / (120*cycle)$
(%i3) unrolling_speedup: unrolled_loop_speed / orig_loop_speed, numer;
(%o3) 1.308333333333333
In other words, the unrolled-loop speed is 1.3 times the original-loop speed. That's 30% faster, right?Likewise, how much faster did shaving those 1200 cycles make the tile-copy code?
(%i4) copy2_speed: (1*tile) / (2893*cycle)$
(%i5) copy3_speed: (1*tile) / (1689*cycle)$
(%i6) copy3_vs_2_speedup: copy3_speed / copy2_speed, numer;
(%o6) 1.712847838957963
I'm willing to believe that I've screwed up here, but I can't see where. Can you show me?Thanks for your help!
However, this can confuse a non-thorough reader because your previous measures are not about "speed"; they are about "cost" (cycles). The parent just did not realize the switch, and thought that you were claiming to have reduced "cost" (cycles) by 30%/70%, which wouldn't be true.
I've updated the story to explain the calculation the first time it happens.
Thanks again!
Also I believe the author is referring to speed improvements, whereas you are referring to time improvements. Let's say I take 200 operations to perform a task, then I optimize code to perform the same task in 100 operations. The program runs at 20 operations per second regardless of how many operations there are.
The task which took 10 seconds before will now take 5 seconds, or a 50% time reduction, which is how you calculate it. However, the program is now running twice as fast, because in 20 seconds, the program can run twice, which is how the author calculated it.
Both of you are correct, just depending on what your context is.
Apple developer support themselves described this idea in Technote #70, http://www.1000bit.it/support/manuali/apple/technotes/iigs/t...
However, we had to leave some space at the edge of backbuffer memory, because if there's an interrupt right at the beginning of the blit, the interrupt handler's stack frame could overflow outside of the backbuffer and corrupt other memory. That one was fun to find. [Edit]: I seem to have missed the second footnote where he already describes this issue.
It had a "barrel shifter" that gave you free shifts of powers of two, so you could calculate screen byte offsets quickly:
// offset = x + y * 320
ADD R0, R1, R2, LSL #8
ADD R0, R0, R2, LSL #5
// = 2 cycles
It also had bulk loads and stores that made reading/writing RAM cheaper. The trick there was to spill as many registers as you possibly could, so that you could transfer as many words as possible per bulk load/store. LDMIA R10!, {R0-R9}
STMIA R11!, {R0-R9}
// Transfers 40 bytes from memory pointed to by R10 to memory pointed to by R11,
// And updates both pointers to the new addresses,
// And only takes (3+10)*2 = 26 cycles to do the lot.
Happy days... ; ARM spells it
ADD R0, R1, R2, LSL #8
; We spell it
CLX R0 ; 1 cycles
ADX R1 ; 1
ADX R2 ; 1
LDX N, 8 ; 3
RSX ; N (const 8 in this case)
Try optimizing THAT!So true that back in the day much of a game programmers mental effort was spent on how to make big ideas fit inside small memory, anemic color palettes, and slow processors.
Also, the quality of tools nowadays is so much better than what you would've used back then. So much so, in fact, that these old machines could even become pretty useful as a learning tool since they're (relatively) simple and allow much easier testing and debugging since they can be used effectively as some kind of VM.
Every 30th of a second, the screen would have to be refreshed. Arcade programmers would perfectly tweak the loops of their assembly programs such that the screen refresh would happen at the right timing. As the CRT scanline would enter "blanks", they would use the borrowed time to process heavier elements of the game. (ie: AIs in Pacman). The heaviest processing would occur on a full-VSync, because you are given more time... as the CRT laser recalibrates from the bottom right corner to the top left corner.
Of course, other games would control the laser perfectly. Asteroids IIRC had extremely sharp graphics because the entire program was not written with "scanlines" as a concept, but instead manually drew every line on the screen by manipulating the CRT laser manually.
Good times... good times...
The Pac-Man Dossier is a great reference for this kind of stuff:
http://home.comcast.net/~jpittman2/pacman/pacmandossier.html
IIRC, Kangaroo used the old scanlines system, but its release date is after Pacman. I guess everyone was still trying to figure out how to do scanlines correctly at that time... so it must be different from arcade-game to arcade game.
Here https://github.com/gorhill/rayoid
So much time spent on counting CPU cycles... Atari ST's pixels layout was really tricky.
There were better games out there, I will always remember Llamatron, my favortie (I wonder if the sources are available somewhere.)
This lead to some neat hacks, like the guys that swapped the vector generator stuff for a laser projector:
In my case, it was on a Mac on a PowerPC CPU. It's a far cry from the limited resources of early personal computers, but this was at a time when 3D was hitting big time - the Playstation had just come out - and I was trying to get performance and effects like a GPU could provide. A hobbyist could get decent rasterization effects from a home-grown 3D engine, but I was working as far forward as I could. All that unrolled code, careful memory access, fixed-point math... I spent a lot of time hand-tuning stuff. It wasn't until I dug into a book on PowerPC architecture that I found some instructions that could perform an approximation of the math quickly, and suddenly I was seeing these beautiful, real-time, true-color, texture-mapped, shaded, transparent triangles floating across the screen at 30fps.
It was about that time that the first 3DFX boards started coming out for Macs, though, and that was the end of that era.
Wikipedia has one more interesting detail about CoCo 3:
http://en.wikipedia.org/wiki/TRS-80_Color_Computer#Color_Com...
"Previous versions of the CoCo ROM had been licensed from Microsoft, but by this time Microsoft was not interested in extending the code further.[citation needed] Instead, Microware provided extensions to Extended Color BASIC to support the new display modes. In order to not violate the spirit of the licensing agreement between Microsoft and Tandy, Microsoft's unmodified BASIC software was loaded in the CoCo 3's ROM. Upon startup, the ROM is copied to RAM and then patched by Microware's code."
Edit: Oh! Just got to footnote 2. Thanks, author!
The C128 had two separate video chips/ports, a C64 compatible chip showing a 40x25 character (320x200 pixel) display, and the "VDC"[1] showing 80x25 characters (640x200, or with interlacing, 640x400 or more), which was output on a separate connector. The VDC had a hideous way to change the display: it had its own video RAM, which the CPU couldn't access directly, instead the video chip had two internal registers (low and high byte) to store the address you wanted to access, and another register to read or write the value at that address. But that wasn't enough, the CPU couldn't access those VDC registers directly either, there was a second indirection on top: the CPU could only access two 2nd-level registers, one in which to store the number of the 'real' register you wanted to access, then you had to poll until the VDC would indicate that it's ready to receive the new value, and you would save the new value for the hidden register in the other 2nd level register. (There's assemply on [1] describing that 2nd level.) Those two registers were the only way of interaction between the CPU and the 80 column display.
[1] http://en.wikipedia.org/wiki/MOS_Technology_VDC
This was extremely slow. Not only because of the amount of instructions, but the VDC would often be slow to issue the readyness flag, thus the CPU would be wasting cycles in a tight loop waiting for the OK.
Now my discovery was that the VDC didn't always react slowly, it had times when the readyness bit would be set on the next CPU cycle. Unsurprisingly, the quick reaction times were during the vertical blanking period (when the ray would travel to the top of the screen, and nothing was displayed). During that time, there wasn't even a need to poll for the VDC's readyness, you could simply feed values to the 2nd level interface as fast as the CPU would allow, without any verification. Thus if you would do your updates to the screen during the vertical screen blank, you would achieve a lot more (more than a magnitude faster, IIRC), and the "impossibly slow" video would actually come into a speed range that might have made it interesting for some kinds of video games. Still too slow to do any real-time hires graphics, and the VDC didn't have any sprites, but it had powerful character based features and quite much internal RAM, plus blitting capabilities, so with enough creativity you might have been able to get away by changing the bitmaps representing selected characters to imitate sprites. And you could run the CPU in its 2 Mhz mode all the time (unlike when using the 40 column video, where you would have to turn it down to 1 Mhz to not interfere with the video chip accessing RAM in parallel, at least during that chip's non-screenblank periods.) My code probably looked something like:
lda #$12 ; VDC Address High Byte register
sta $d600 ; write to control register
lda #$10 ; address hi byte
sta $d601 ; store
ldx #$13 ; VDC Address Low Byte register
ldy #$00 ; address lo byte
loop
cycles
stx $d600 ; select address low byte register 4
sty $d601 ; update address low byte 4
lda #$1f ; VDC Data Register 3 ?
sta $d600 4
lda base,y ; load value from CPU RAM 4 ?
sta $d601 ; store in VDC RAM 4
iny 2
bne loop ; or do some loop unrolling 3 ?
..
(28 cycles per byte, at 2 Mhz, => about 300-400 bytes per frame. Although the C128 could remap the zero page, too (to any page?), and definitely relocate the stack to any page, thus there are a couple ways to optimize this. (Hm, was there also a mode that had the VDC auto-increment the address pointer? Thus pushing data to $d601 repeatedly would be all that was needed? I can't remember.))How would you time your screen updates to the vertical blanking period? There was no way for the VDC to deliver interrupts. It did however have a register that returned the vertical ray position. Also, the C128 had a separate IC holding timers. Thus IIRC I wrote code to reprogram the timer on every frame with updated timing calculations, so that I got an interrupt right when the VDC would enter the vertical blanking area.
As I said, I'm not aware of any production level program that used this; perhaps some did, but at least the behaviour was not documented in the manuals I had.
The VDC felt even more like a waste after I discovered this. The only use I had for it was using some text editor. I wasn't up to writing big programs at the time, either.
PS. sorry if that was a bit long.
Another interesting tidbit that should be obvious, but I miss a lot. The format of the graphics was fixed and not necessarily on the table for things that can be changed to make the code work. All too often it seems I let what I'm wanting to accomplish affect how I plan on storing the data I'm operating on.
Today most of game programmers just ask for a bigger GPU.
I myself, for example, find these clever hacks very rewarding, maybe even more than finishing the stuff I'm working into.
Incidentally that reminded me. How does one get to these positions? As the API's are closed there is no real way for an outsider to learn the ropes.
- Deep understanding of low-level programming issues: assembly, architectures, caches, DMAs
- Experience doing advanced stuff in available platforms like OpenGL, DirectX, Cuda/OpenCL, etc.
- Fluency in Maths is not necessary for all low-level development, but will certainly be for physics and graphics work.
Check out places like http://directtovideo.wordpress.com/, http://www.iquilezles.org/www/index.htm and http://fgiesen.wordpress.com/. If you can show you even understand what they are talking about, you have a foot in the door. The rest is hand-on experience. :)
Thanks about those links. While I'm well acquainted with iquilezles the other two were new for me.
The first link was especially interesting as I'm quite interesting in OpenCL raytracing, so that will save me a lot of research time in the near future, especially about the data structure selection. Interestingly with OpenCL 2.0 and devices that have shared memory anything that is fast to build on CPU and slow/open area of research on GPU can be pretty much ignored as the bottleneck between the two devices do not exist anymore.
http://www.templeos.org/TempleOS.html
Some features:
64-bit ring-0-only single-address-map (identity) multitasking kernel
HolyC programming lanaguage interpreter
Praise God for binds using timer based random number generators
Create comics, hymns, poems as offerings to the Oracle
"The 6809E was featured in the TRS-80 Color Computer (CoCo), the Acorn System 2, 3 and 4 computers (as an optional alternative to their standard 6502), the Fujitsu FM-7, the Welsh-made Dragon 32/64 home computers,"
I know and have/had almost all of them. Maybe more popular in Europe?
Until you come back and decide to buy something, and suddenly .. it mattered which CPU you chose. Z80, 68xx, the rumoured 68k, and .. so on. It was a bewildering time.
(Thankfully, the still-functioning, embedded Video Game Arcades were able to provide some relief. I knew a few arcade owners, lets just say.. ;)
Anyway .. there was a Tandy shop, and I had a CoCo section in my (long-lost) notebook. There were some funky things, no? But still .. in the Amstrad invasion, the Atari megalith, the Commodore death spiral.. Tandy kind of lost the plot. Strangely so. The CoCo were somehow 'tainted' in my mind, back then, because of the TRS80 connection. Things didn't have to be so grey.
Nevertheless, if the connection with the Arcades had been made, I'm sure the CoCo could have kicked some serious ass. What heady times!
So, I urge you to reconsider your opinion; I believe it may be wholly incorrect.
Alas, I realize not everyone thinks the way I do, but I consider among the worlds most functional 'computing experiences' occurred during the Arcade era, where, indeed, one could wonder at the .. sheer .. computing power .. being used to render sprites. With 256 colors.
So, at the risk of perhaps moving the goal-posts, let me just say that while the 6809 might've been a 'games console cpu', at the time, it was also used in a few .. relatively interesting .. micro's.
Anyone returning to that era, in their own way, contributes what they can. In my case, I consider the still-functioning Joust and Defender machines in my neighbourhood, at the very least, accessible through a repository ..
And I can also confirm that they also used the PSHS/PULU hack to do some very fast memory copies and blanks.
Brought back nice memories.
Why not a power of 2 like 16 or 32?
Their fix of drawing in the opposite order works because a push/pop of the stack will then affect the upcoming graphics bytes instead of the previous ones.
Edit: yes, the stack pointers were 16 bit (http://en.wikipedia.org/wiki/Motorola_6809).
void execute_interrupt(void * isr) {
// push the registers of the program running when
// the interrupt happens onto the stack that's currently
// in use. this is something the CPU looks after usually
// (I think).
// This meant in the article that the contents of the
// registers were being pushed to the buffer holding
// the current image, overwriting whatever data was
// above the stack pointer. The fix was to swap which
// register was the source and which the dest, so that
// the ISR would be pushing to the buffer that was about
// to be written to.
pushRegs();
// do whatever's needed to get the CPU into a consistant
// state for the ISR
zeroRegs();
// call the interrupt service routine for whatever interrupt
// has occured
isr();
// just decrements the system stack pointer, leaving
// whatever was above the point corrupted in the
// original version using the two stack regs
popRegs();
}Brilliant blog post!