Main reasons for not benching against pdqsort:
1. stability
2. worse performance on long doubles, and I don't know why
3. A variety of hard to explain performance differences.
4. pdqsort does better on generic data, which can be very important.
So it's tricky to present a fair benchmark when two sorts behave very differently.
As to performance advantages of fluxsort:
1. It has a faster insertion sort.
2. Branchless pseudomedian of 15 gives an advantage.
3. Partial loop unrolling with: while (ptx + 8 < pte)
4. Data movement should be nearly identical, if not better, with the recursive calls through the ptx pointer. In the optimal case the memcpy only triggers when the partition shrinks below 24 elements, and in half of those cases the partition will already be in main memory.
So on random you could expect n / 2 extra data movements on top of ~ n log n moves.