Multiplying and Dividing on the 6502 (2021)
llx.com
llx.com
Here's one possible source from an excellent website that includes many more 6502 algorithms!
https://codebase64.org/doku.php?id=base:8bit_multiplication_...
> Numerous advantages are gained improving the speed of the multiply. Routines that rely on it can be copied from the ROM and pointed to the new routine to improve performance. The following example relies on the Steve Judd's fast multiplication and reduces the number of cycles to multiply from around 2300 to little over 1400.
That's a nice speed increase from 435 multiplies/sec to 714 ops/sec. It's downright spritely!
There's also a faster log function that
> executes in around 13000 cycles, where the built-in routine takes 19000.
...jumping from 53 to 77 ops/sec. Smokin'!
The other tricks are all worth learning, cordic, shifts, look up tables and the other tricks we use when add shifts and rotates and bit opps are the only thing the CPU can do.
Since y * 40 is (y * 8) + (y * 32), you'd take x, left-shift three times, save the result, shift left two more times, and add the two together. Pretty fast. You'd do stuff like this everywhere you knew the coefficients in advance.
Also, a lookup table for the Y address, plus a fixed X address can get bitmaps blitted to the screen quickly.
For special case graphics, just code in the addresses, unroll the loop and use index X or Y to handle screen positioning.
LDX XPOS
LDA IMAGE0
STA SCREEN0, X
LDA IMAGE1
STA SCREEN1, X
LDA IMAGE2
STA SCREEN2, X
.
.
.No multiply needed. Y addresses either coded at fixed position, or generated at runtime based on some event and or when there is time to compute it all off screen.
Trade RAM for speed!
Today, I mostly do it for fun on my old Apple machine. It is enjoyable to see what can be done. 1Mhz is actually quite a lot when one looks at the resolution and overall task to complete. And where to put it?
Things like having the game run at 30Hz or 15Hz so more can be computed in game, in real time. Then, when it comes to drawing everything, often a simple screen clear, or page flip then draw does not make sense. Two screens takes too much RAM, etc...
Single buffer, draw fast, only draw deltas, using XOR, all add up!
Getting paid?
I have done some assembly for money on Propeller chips, PIC, AVR, and the like. There are times in embedded land where doing these fun things pays off and makes sense.
For what it is worth, a little fun can be had for income, but not much.
https://atjs.mbnet.fi/mc6809/Information/6809.htm
8x8 unsigned. I started out programming assembler on the KIM, then moved to the Dragon 32 which had a 6809 and then back to the BBC micro again, which in almost every way felt like a huge step forward and in that one way felt like a step back, the lack of 8x8 multiply in hardware. Here is a very neat (and pretty quick) multiplication routine:
However, the 6502 was a challenge to implement high-level languages on, IF you also wanted speed and small code size. But since the 6502s were slow compared to modern chips, and could only address 64KiB directly, you typically also wanted speed and small size.
I posted some links to some approaches here: https://dwheeler.com/6502/
Atalan and Plasma look nice.
* mul/div take a lot of transistors and may not be a single or fixed cycle instruction.
* A lot of embedded systems don't really need those instructions anyway - they just encode discrete state machines or do something like networking that doesn't need * or / for operation.
* less transistors means the chips are cheaper, and single (or fixed) cycles per instruction means hard real-time is simpler to reason about.
Is that correct?
I think we should look at the culture of programming around architectures like SPARC and MIPS and what kind of assumptions the designers had about how people would program them. Even though you could find compilers for 6502, and you could write assembly for MIPS and SPARC, assembly programming was de rigueur for 6502 and it was not for MIPS and SPARC. 6502 systems are typically cheap and have small amounts of memory, and the architecture is a bit weird if your targeting a C compiler (inconvenient 256 byte stack, for example). MIPS and SPARC systems didn't appear until the late 1980s, they tended to have much larger amounts of memory, and it was assumed that you would use C (or something else).
SPARC and MIPS may have also omitted the multiply instruction in an attempt to make it so the processor could execute one operation per cycle for nearly every instruction.
As for gcc, I'm fairly sure gcc has at least one out-of-tree but yeah perhaps the code wasn't as good as hand rolled.
(They used ARM2 which added MUL and MLA - multiply with accumulate)
But you could write:
MOV R1, #3
MUL R0, R0, R1
In spite of the limitations, this gave the instruction set a certain 68000 quality to it (except much faster for a given clock speed).To multiply the number in R0 by 3, which was pretty convenient. The ARM had a thing called the barrel shifter though, which let you add an arbitrary shift to the last operand of any arithmetic operation. All arithmetic ops take 1 processor cycle, so you could write this instead to multiply by 3 in a single cycle:
ADD R0, R0, R0, LSL #1
Ie, add R0 to itself multiplied by 2. Constant divisions could be constructed with the SUB instruction too. Some constants required multiple instructions (but I think the maximum was something like 4 or 5 instructions for any constant? I wrote an assembler that could figure this out for you automatically in the mid-90s so I used to know for sure).This is basically a single-instruction version of the 6502 trick (handy, because the first OS for the ARM was a hurried port of a 6502 operating system), which sort of fits with the ARM's original inspiration as being a 32-bit version of the 6502. As each instruction completes in one CPU cycle, the ARM could have fairly monstrous integer performance for the mid to late 80s if you knew how to program it.
https://archive.org/details/dr_dobbs_journal_vol_01/page/n20...
Wozniak floating point code.
I took an interest due to the C64.
I love that there was a time when "just write the binary equivalent" was the easy solution.