Here are my results:
sort_asm_recurse.asm 69 ms/loop
clang++ 3.8.0/sort_cpp_recurse.cpp 65 ms/loop
g++ 5.4.0/sort_cpp_recurse.cpp 70 ms/loop
Compiler flags: -O3 --std=c++11 -fomit-frame-pointer -march=native -mtune=nativeSo on my computer, the assembly code (barely) beat g++ but not clang++. From a cursory glance of the assembler code clang++ generates, the difference seem to be that it adds alignment to critical loops.
It is also smarter at using 32bit registers when it can get away with it. F.e the handwritten assembler code contains "xor r9, r9". An equivalent but faster variant that the compiler generates is "xor r9d, r9d".
There is also a slight error in the assembly code. rsp should be aligned to a 16 byte boundary when a call instruction is executed and the code doesn't ensure that. Likely it loses a whole bunch of performance by calling from unaligned addresses.