It sorts in time O(ku-ulog(u)+n), where k is the number of bits in each key, u is the number of unique elements, and n is the total number of elements.
It sorts in time O(ku-ulog(u)+n), where k is the number of bits in each key, u is the number of unique elements, and n is the total number of elements.
I have not made SIMD optimizations, which djbsort includes, and skylinesort does not require pre-computing the merging network, or padding the array to that size. This works great for crypto where array sizes are known ahead of time and generally always the same, but not all use cases are like this.
Second, consider this line.
> Otherwise we keep traveling to the left until we are able to make another leap.
Wouldn't sorting an array with the two elements 2^n and 2^(n+1)-1 require ~2^n repeats of this step? That's a second 2^k cost.
Overall this is just counting sort with skips, but for that I'd imagine it's cheaper to just use a bit array to track set elements, which also makes zeroing much faster.
When there's 2^n and 2^(n+1)-1, it takes n-2 steps to reach the former from the latter, not 2^n.
The other thing about it that's better is that Skylinesort is the only sorting algorithm that is downward-concave, i.e. it takes less time to sort a list of size a (keeping clustering constant) than it takes to sort two lists of size a/2. In other words, it shows economies of scale. It has been a long-remarked paradox that sorting algorithms take longer per element when there's more elements.
It sounds like this is just a variation of counting sort or bucket sort optimized for a small number of unique values.
When I say that it can sort an array of size 2n faster than two arrays of size n it is because if you plot the asymptotic time, you find it's second derivative is negative (concave-up). This is because of the subtraction in the asymptotic time, O(k u - u log u + n). A Stanford professor was reluctant to accept an asymptotic time with a subtraction in it, but I explained it was really the logarithm of a division, ku-ulogu was actually u log(K/u) where K is 2^k, the total number of elements in the empty auxiliary array. log(K / u) is the same as log K - log u and log(K)=k, so you end up with k u - u log u, which has a negative second derivative when k is fixed.
When log u = k, skylinesort does indeed run in linear time, but it runs slower that linear time before that, and log u is never greater than k, by definition.
<rant> That kind of phenomenon (reading too far into an infinite process) is ubiquitous once you start noticing it. Infinitely many universes don't imply the existence of every possible universe (if you subscribe to such theories), pie having infinitely many digits doesn't imply that it encodes every possible message somewhere in those digits (most real numbers have that property, but iirc it's still an open question for pi), and so on. </rant>