Code / scripts here: https://github.com/funrollloops/parallel-sort-bench
I had to use a cloud instance for testing since I don't have an avx512-capable personal machine.
Code / scripts here: https://github.com/funrollloops/parallel-sort-bench
I had to use a cloud instance for testing since I don't have an avx512-capable personal machine.
We also have a CMake build, not sure about well-engineered but patches welcome from anyone with more CMake expertise :)
I have heard that FetchContent can make something unreliable, as in dependent on having a connection, but I think it opens up the door for a lot of people to be willing to try it :) (That, and you can turn off mandatory updates, making it work offline too).
I think the basic rule for making FetchContent work is to just put the CMakeLists.txt in the root directory and then make sure that you can use the project as a CMake sub-project.
avx512_qsort benefits from sorted runs while vqsort does not, but vqsort is still 20% faster in this case as well.
The full results are in the README, but a short version:
--------------------------------------------------
Benchmark Time CPU
--------------------------------------------------
TbbSort/random 3.17 s 16.7 s
HwySort/random 2.27 s 2.27 s
StdPartitionHwySort/random 2.06 s 4.01 s
PdqSort/random 5.67 s 5.67 s
IntelX86SIMDSort/random 3.73 s 3.73 s
TbbSort/sorted_runs 1.58 s 8.02 s
HwySort/sorted_runs 2.38 s 2.38 s
StdPartitionHwySort/sorted_runs 1.11 s 3.21 s
PdqSort/sorted_runs 5.30 s 5.30 s
IntelX86SIMDSort/sorted_runs 2.90 s 2.90 sNote that there appears to be some bug because our benchmark's verification fails.