Not sure if that's smart or stupid.
Not sure if that's smart or stupid.
(1 * 5 * 9 * 13 * 17 * ...) * (2 * 6 * 10 * 14 * 18 * ...) * (3 * 7 * 11 * 15 * 19 * ...) * (4 * 8 * 12 * 16 * 20 * ...)
(it's not precisely that, but close enough)
It isn't optimal however, optimal code would be pre-computed results (signed integer overflow is undefined, so n <= 12 is defined)
(if signed integer overflow would be defined to be 2's complement overflow, you can still use a table, as n > 33 gives 0)
1. Turned from recursive form into iterative form 2. Unrolled heavily 3. Autovectorized (!)
The throughput will be substantially more than the simple version.
I'd like to know how the recursion to iteration step was done.
vs. what it does with 64-bit ints:
Can anybody explain the 32 bit version?