The speed of the Hello world application is the metric and they used hyperfine - a benchmarking tool.
The speed of the Hello world application is the metric and they used hyperfine - a benchmarking tool.
The post pretty clearly blames "C++" if you read more than the title - note the "due to C++" here:
>> Yet if these numbers are to be believed, there is a significant penalty due to C++ for tiny program executions, under Linux.
> The speed of the Hello world application is the metric
And it's a poor metric, mostly testing process setup/teardown.
> and they used hyperfine - a benchmarking tool.
And benchmarking tools are only as good as their usage and application.
A more apples-to-apples comparison would be to use `\n` instead of `std::endl`, and `std::ios_sync_with_stdio(false)` to avoid excessive syncronization with C I/O. Admittedly, neither particularly helps here: I observe most of the overhead the article does merely by switching out clang/gcc for clang++/g++ on the same C source code, with ~6ms-9ms total runtime under wsl, or ~9ms (no measurable difference between the C++ or C code) when compiled with cl 19.15.26732.1 and run on windows. This is mostly measuring process setup/teardown and OS I/O overhead: redirecting stdout to null, I don't see a significant perf impact until I put the printing in some rather large loops:
#include <stdio.h>
#include <stdlib.h>
int main() {
for (int i=0; i<1000000; ++i) printf("hello world\n");
return EXIT_SUCCESS;
}
#include <iostream>
#include <stdlib.h>
int main() {
std::ios::sync_with_stdio(false);
for (int i=0; i<1000000; ++i) std::cout << "hello world\n";
return EXIT_SUCCESS;
}
Guess which is faster? The C++ version, oddly enough, both on wsl ubuntu linux: /mnt/c/local/ben$ clang++ -Os hello_world.cpp && cargo run --bin bench --quiet --release
average runtime: 36.78434ms
/mnt/c/local/ben$ clang -Os hello_world.c && cargo run --bin bench --quiet --release
average runtime: 86.33667ms
/mnt/c/local/ben$ clang++ -O3 hello_world.cpp && cargo run --bin bench --quiet --release
average runtime: 36.34467ms
/mnt/c/local/ben$ clang -O3 hello_world.c && cargo run --bin bench --quiet --release
average runtime: 86.75787ms
And on windows: C:\local\ben>cl /nologo /O2 hello_world.cpp && cargo run --bin bench --quiet --release
hello_world.cpp
C:\Program Files (x86)\Microsoft Visual Studio\2017\Community\VC\Tools\MSVC\14.15.26726\include\xlocale(319): warning C4530: C++ exception handler used, but unwind semantics are not enabled. Specify /EHsc
average runtime: 1.44849203s
C:\local\ben>cl /nologo /O2 hello_world.c && cargo run --bin bench --quiet --release
hello_world.c
average runtime: 1.75088561s
Take these numbers with a massive grain of salt: Launching the first subprocess is extra expensive (on the order of 90-100ms) and included in the above averages, distoring them, presumably from delayed loading or initialization of a library in the benchmarking process as subsequent runs of the benchmarking process show the same overhead for the first subprocess[1]. Meanwhile, the author of the original article is calling overhead that has to be rounded up to 1ms a "huge penalty".(EDIT[1]: to rule out first-access overhead being from windows defender or similar, I confirmed that launching a sacrificial, unrelated, first subprocess such as "cmd /C ver" eliminates the overhead for the first "./a.out" execution)
However, you're testing the speed to print something, not the application. I mean, I'm assuming they wanted to include the process start and teardown.
But sure, the test isn't fair.
I'm testing what impact modifications to the application have, to better understand what is and isn't contributing to the original application's performance and bottlenecks. I've shown how far I have to modify the original to actually start to hit I/O bottlenecks, which by extension, also shows how little they contribute to the bottlenecks of the original application, despite people looking to blame it for the "C" vs "C++" differences elsewhere in the discussion tree.
> I'm assuming they wanted to include the process start and teardown.
And there are interesting discussions to have about that in terms of breaking down OS overhead, API overhead, first launch vs subsequent launches, etc. - none of which the original post bothered with. We could peek at a more realistic workflow involving, say, an xargs-spammed executable processing a piped stream, if we really wanted to simulate a startup/teardown heavy workload. We could fire up `perf` to prepare to profile - I at least installed the frontend before realizing actually using it on wsl requires compiling stuff.
The post didn't bother with any of that. And to be fair, I didn't bother too much either. Low-effort on my part as well ;)
> I do not believe that printing ‘hello world’ itself should be slower or faster in C++, at least not significantly. What we are testing by running these programs is the overhead due to the choice of programming language.