As a somewhat representative example - suppose you have a bunch of objects that you want to work on. On the Z80 you'd probably adopt a modern-style struct system, whereby each object is represented by a block of memory. Offset 0 is X coordinate, offset 1 is Y coordinate, offset 2 is flags, blah de blah. Suppose each object is 8 bytes, and you want to set a flag for each object that's on screen.
LD DE,(MAX_X<<8)|MIN_X ; 10 10
LD IY,8 ; 14 24
LD IX,OBJECT0 ; 14 38
LD B,NUMOBJECTS ; 7 45
.LOOP
LD A,(IX+0) ; 19 19
CP D ; 4 23
JP C,NOTVISIBLE ; 10 33
CP E ; 4 37
JP NC,NOTVISIBLE ; 10 47
.VISIBLE
LD A,(IX+2) ; 19 19
OR VISFLAG ; 7 26
LD (IX+2),A ; 19 45
ADD IX,IY ; 15 60
DJNZ LOOP ; 13 73
JP DONE ; 10
.NOTVISIBLE
LD A,(IX+2) ; 19 19
AND ~VISFLAG ; 7 26
LD (IX+2),A ; 19 45
ADD IX,IY ; 15 60
DJNZ LOOP ; 13 73
.DONE
(For readers born after 1985 :) - numbers in the comments are cycle counts: instruction's count, then cumulative total for this block)So, for N objects: (45 + N * 120). Plus 10 if the last one was visible.
For the 6502 you'd probably adopt a striped layout, so rather than having each object as a block of data, you'd have an X table, a Y table, a flags table, and so on. So the same again:
LDX NUMOBJECTS ; 3 3
.LOOP
LDA XS-1,X ; 4 4
CMP MINX ; 3 7
BCC NOTVISIBLE ; 3 10
CMP MAXX ; 3 13
BCS NOTVISIBLE ; 3 16
.VISIBLE
LDA FLAGS-1,X ; 4 4
ORA #VISFLAG ; 2 6
STA FLAGS-1,X ; 5 11
DEX ; 2 13
BNE LOOP ; 3 16
BEQ DONE ; 3
.NOTVISIBLE
LDA FLAGS-1,X ; 4 4
AND #~VISFLAG ; 2 6
STA FLAGS-1,X ; 5 11
DEX ; 2 13
BNE LOOP ; 3 16
.DONE
So, for N objects: (3 + N * 32). Plus 3 if the last one was visible. (I've been a bit scrappy with the cycle counts here, by counting each branch as taken, even though that makes the totals invalid. This is done to favour the Z80, which doesn't appear to execute untaken branches any quicker.)That's pretty much the 4:1 improvement that's commonly claimed. 1MHz 6502 will beat 3.5MHz Z80; 2MHz 6502 will beat 4MHz Z80.
This might seem like a synthetic benchmark, designed to make the 6502 look better, but this sort of thing crops up fairly often. (I noticed it initially after noting how crappy-looking the code was for a couple of Z80 games I was disassembling; I only realised why after trying to rewrite the snippets in question myself!)
One thing to note in particular here is that in the Z80 case you're two registers down - because IY has been used for the array stride (no immediate 16-bit addition...), and DE has been used for the min/max constants (immediate instructions are more expensive).
And this is all very well as the code is presented here, but in practice, in many cases you'll probably need to use DE in your code. You can work around this by using EXX and having your constants in the shadow register bank, but you're losing 8 T-states per iteration (EXX = 4 T-states). For many purposes you'd be better off using self-modifying code - but now you lose 3 T-states per addition (immediate instructions are more expensive).
So a bit of a disappointment for me. Initial excitement at seeing just how much crazy stuff the Z80 can do rapidly turned into disillusionment as I realised just how long it takes to do anything. It's like the 68000 in this respect.
Good luck beating LDIR with a 6502 at quarter the clock rate, though...
Z80 code is also generally much easier to follow...
Z80 has a 16-bit stack pointer too...