Is there a path forward for compilers to eek out these optimization gains eventually? Is there even a path?
550x gains with some C ++ / mixed gnarly low level assembly vs standard C++ is pretty shocking to me.
550x gains with some C ++ / mixed gnarly low level assembly vs standard C++ is pretty shocking to me.
const u8 *file_lo;
file_lo = (const u8*)mmap(0,250000000ull,PROT_READ,MAP_PRIVATE|MAP_POPULATE,0,0);
const u8 *file_hi = file_lo + 250000000ull;
u64 count = 0;
while (file_lo < file_hi) {
if (*file_lo == 127) {
++count;
}
file_lo++;
}
I got a bit under 54ms. The solution in the article runs in a bit under 16ms.Used clang with -Ofast -march=native -static. Funnily, gcc gets only 54000 with the same options, 1.6 times slower.
So you can get 25k with following code, clang -Ofast -std=c++17 -march=native -static
#include <iostream>
#include <cstdint>
#include <sys/mman.h>
#include <unistd.h>
int main() {
auto file_lo = (const uint8_t*)mmap(0,250000000ull,PROT_READ,MAP_PRIVATE|MAP_POPULATE,STDIN_FILENO,0);
int count = 0;
for (uint32_t i = 0; i < 250000000; ++i) {
if (file_lo[i] == 127) {
++count;
}
}
std::cout << count << std::endl;
_exit(0);
return 0;
} if (*file_lo == 127) {
++count;
with count += (*file_lo == 127);
That might save you the occasional branch mis-prediction, and might possibly open up some hardware-level loop optimisations. Any difference?