Also, the world has changed a lot since then. Interpreters have less penalty on a modern chip than an old stupid chip, because branch prediction, prefetching, and multiple pipelines can really help with them, so it's relatively speaking cheaper to examine data and make decisions and the CPU will spend more time "doing things" as long as the data required and the branches taken are predictable, which they often are in this sort of code. And on the flip side, modern processors really want your code to be static, precisely so that all those optimizations can work well, along with code caches, micro-op caches, etc... constantly changing the code isn't good for performance on modern chips. The 6502 doesn't care how much the code is changing, it just executes the next opcode at the same speed regardless.
Very different world.
Another case is to access a large range of memory, for instance, to fetch data from a large table. You have the code in ram and increment the high-byte address (HH) of a 'lda $HH00,x' or 'sta $HH00,x' to access larger ram area (because x index can only access 256 bytes). I've seen that in Vic-20 games, I don't know if NES games used it.
One case is loop unrolling, a.k.a. speedcode. This is very common practice in modern 6502 demos. I don't think that the old NES games used it, but some modern NES demos may use it. See http://codebase64.org/doku.php?id=base:speedcode or http://csdb.dk/forums/index.php?roomid=11&topicid=96279&show...
frac = 0;
while (len > 0) {
for (frac += scale; frac >= 1; frac--)
*dst++ = *src;
src++;
}
Instead of repeating the work of the inner loop every time, generate the code that has the scale baked in. For instance, doubling the image would output this code fragment repeated for the width of the source image: *dst++ = *src; *dst++ = *src; src++;
And scaling down by half would be: *dst++ = *src; src++; src++;
Though whether this is faster or the best technique really depends on the processor. It might be just as easy to pre-compute a lookup table pointing to the offset of the source pixel for each destination one. That's something an 8086 could do fairly easily and maybe a 6502, but not so much for a Z-80.You find yourself on an uncharted desert island with two sailors, a movie star, some other lady, a millionaire and his wife... and a crate of 4k roms. For reasons that would take far too long to explain here - your only salvation is to recreate the Atari game catalog on your coconut game console.
[edit] It's much faster than the obvious solution, which is to do LDA $D801 / STA $D800 / LDA $D802 / STA $D801 / etc. Or, even worse, a loop incrementing the X register with LDA $D801,X / STA $D800,X / etc. Or worse still: indirection via zero-page (though no-one should really ever consider that for colour RAM updates, even though at first glance it seems clever for moving the characters between screens, it eats too much time)
(This technique is pretty much only applicable if your scroll is fixed-direction and fixed-speed)
[edit #2] Credit for that goes to Jon Williams (Shadow Dancer, C64 - and others) for adding that optimisation to my scroll routines that he used in SD - and for then telling me the trick :)
Thinking about it I wonder why I didn't copy the core of that down to the zero page and write the value into the instructions there. STA $12 is a cycle faster than STA $1234 maybe I couldn't find the space (or I was being kind to the OS + basic)