Worse, a stable version of BlockQuicksort is pointless, which is to say that it is actually impossible for there to be a practical use case for Fluxsort. The branchless partitioning technique is relevant only if comparison doesn't depend on branching, which means the input array needs to consist machine-native types, and in fact Fluxsort limits itself to integers and long double. On these inputs stability isn't an observable property! The result of Fluxsort will be indistinguishable from that of any other quicksort (okay, I guess it will keep negative and positive zeros in the original order for floats, but an integer-based comparison could also accomplish this). So all you get is a slower sort that uses more memory. There are use cases for stable partitioning, such as sort-by or arg-sort, but as written Fluxsort can't do these.
The statistical method to decide whether to use quicksort or Quadsort—checking the ratio of comparisons between adjacent pairs of elements that are smaller versus greater—is really interesting. It applies to the mergesort versus quicksort problem in general.
The use of the ptx pointer is definitely not novel and I think it's a stretch to even describe it as a brain-twister. It's the obvious way to do branchless stable partitioning.
This is good work! It's well-written and well-explained, and clearly points in the direction of practical applications. However, it's presented in a way that suggests that it's a good drop-in sorting algorithm, when instead it's more of a research effort. Beware!