the paper is not painful at all
unless i fucked it up, it looks like you can insert a separate renormalization step before the sorting where you shift each number to the left by a variable amount, like a floating-point unit always does with the mantissa (except subnormals), and that seems to solve the exponential distribution problem; it always seems to get down to a single item from 10000 34-bit items in about 15 steps
no wait, it doesn't really solve it, because a vector of the first 1000 fibonacci numbers still takes 485 iterations. but the last number in that vector is a 694-bit number. it does seem to improve it enormously
i thought this might make it work much worse (because in a sense it's adding bits to the numbers: what used to be a 1-bit number might now have n-bit-wide differences with the numbers before and after it) but at least in random tests it seems to make a huge improvement
just to clarify, what i'm doing (with unsigned integers) is
def normalize(v):
for vi in v:
while vi < 2**34:
vi *= 2
yield vi
def nreductions(v):
while True:
v = list(sortu(normalize(v)))
yield v
v = list(diffs(v))
with 256-bit numbers and a 2**256 normalization target it seems to typically be about 30 or 40 reduction steps, not sure if those qualify as bignums to you
the shifts of course have to be undone in the other direction, just like the permutations, but i don't think that's a problem?
(oh, now i see that in §3.1 'alignment' you are already doing something like this, except that you're shifting right to reduce the number of duplicates and eliminate one extra bit of differencing per iteration, not left to reduce the dynamic range of the data. for smallish numbers that seems to be roughly as effective, but left-shift normalizing works a lot better than right-shift aligning for 256-bit numbers)
i haven't tried doing any actual vector multiplies with this algorithm yet so if i did fuck it up i wouldn't have noticed
this is a pretty exciting algorithm, thanks for sharing