But the 1ms number is also measured with strace.
Plus, the article is in 2015, and the author did not mention the CPU and other configuration. So that also make things muddy.
Plus, the article is in 2015, and the author did not mention the CPU and other configuration. So that also make things muddy.
Profiling my trivial program with callgrind shows (unsurprisingly) that the majority of the time is in dynamic library relocation and whatever __GI__tunables_init does: https://i.imgur.com/Yligh7S.png