The code is not portable to other compilers/OSes so I can’t even run it here on Windows. The code uses inline assembly a lot, i.e. optimizer likely failed to achieve much, and for code like this optimizations are very important.
The reason for inline assembly is also worse mentioning, it’s for rdtsc instruction. Mainstream CPUs have dynamic frequency for the last decade. Recent CPUs return rdtsc values unrelated to the count of instruction it executed, it’s a counter running at some constant frequency. The code uses AVX, the CPU is therefore at least sandy bridge or AMD equivalent. Now in 2018, __asm rdtsc is no longer a good method for performance measures.