Turns out, you do always have a calculator with you.
I think we could gain from emphasizing math with orders of magnitude, and knowing 1/n's would be super handy too.
2x = double it. 3x = triple it. 4x = double it twice. 5x = ... 6x = triple it then double it. 7x = ... 8x = double it thrice. 9x = from 10x subtract it once. 10x = put "0" on the end.
That gets you everything but 5x5, 5x7, and 7x7. That's just 3 things to memorize instead of 55. I usually get the 5s by just halving 10x, which then handles every case but 7x7. The overhead of doing 2 (sometimes 3) mental additions isn't much worse than memorizing everything as a singular operation and this way of doing things makes it a bit clearer what multiplication represents rather than being an arbitrary thing to memorize because teacher said so.
To hell with primary school. Memorizing is for suckers.
BTW 5x is "double it twice (4x) and add it (1x)"
Personally, for up to 10x10 - I prefer a lookup table (both mentally and in code)
P.S. It's been 30 years? since I had to implement it, but as I recall multiplication could be done by any number with just doubling and adding the initial value.. 2x = double 3x = double +x 4x = double the double, etc. any even multiple be done just by repeating the doubling, and odd multiples are gotten by repeated doubling + adding the starting value.
This worked great for cpu's as 'doubling' was just a bit shift left 'x' number of times - very fast.
Except that you'd get a summation with the number of terms equal to the number of bits in the source operands. Multiplication in a CPU only seems very fast because the theory is simple. Building a fast multiplier in silicon requires a lot of area and a lot of interconnect, because you don't require "just" a massive addition table, you also need carry lookahead logic to perform fast carry propagation otherwise you get no speed increase.
A naive multiplier in a CPU requires 1 cycle per result bit (multiplying two 32-bit numbers requires 32 additions, then you need to wait another 32 cycles for full carry propagation). That's very slow, in CPU terms.
But again, it was 30 years ago so my memory may be fuzzy.
Edit: Hmm - if I recall, it was something like: and the value with x01 to check LSB,
asl value, (multiplicand masked with LSB=0),
if LSB check was true, add value to result,
check overflow flag ?
Again, this was simple binary data usually of a fixed 1 byte or 2 byte length. we used BCD for 'real work' and I can't even remember how that was done...
I always remember it being much simpler then it appears it was. Guess I'm getting rusty from using fancy languages with built in multiply instructions :-)
If you add division by two into the list of primitives, 5 and 6 become accessible within 2 steps and so the system can reach the whole 10s table within a cost of 2 (times 10 is considered a zero cost step). That makes the cost one step worse than traditional memorization but you end up being faster overall because you only need to focus on being fast at the three primitives.
Recognizing the general principle that you can change A into BC or (B+C) where B and C are both easier to multiply with than A is far more useful than memorizing the table. The best part is except for the (optional) division by two, all the primitive operations are things you already learned how to do at that stage. You can short circuit a tedious, long and wasteful part of the early math curriculum AND walk away with genuine understanding of why the rules are and how to figure things out for yourself. While everyone else struggles to multiply by 17, you can be the smarty who doubled 4 times and added once. Or put zero at the end, doubled, and subtracted 3x. Or put zero at the end, subtracted 1x, doubled it, subtracted 1x. With a bit of thought, you can extend the entire multiplication table to 20x20 such that very few rarely does anything cost more than 3 steps.