Edit: FWIW, building with "gcc foo.c -o foo -O2 -s -Wall --static" was slightly faster than without the --static. And "musl-gcc foo.c -o foo -O2 -s -Wall --static" was slightly better again. Maybe 0.5ms improvement in total. I'm on Ubuntu 20 x64, gcc 9, Intel Xeon W-2133.
$ git clone https://github.com/jart/cosmopolitan && cd cosmopolitan
$ make -j8 o//examples/rusage.com o//examples/hello.com
$ o//examples/rusage.com o//examples/hello.com
hello world 123
RL: took 90µs wall time
RL: ballooned to 108kb in size
RL: needed 106µs cpu (0% kernel)
RL: caused 13 page faults (100% memcpy)
This one has a tiny bit more overhead than the practical minimum you saw earlier in the benchmark. Note that APE binaries start off as shell scripts so the first run is going to be slower. It's something I'm working towards improving in a variety of ways. Here's the executive summary of the upcoming changes: int ws, pid;
CHECK_NE(-1, (pid = fork()));
if (!pid) {
execve("ape.com", (char *const[]){"ape.com", 0}, environ);
perror("execve");
_Exit(127);
}
CHECK_EQ(pid, wait(&ws));
CHECK_TRUE(WIFEXITED(ws));
CHECK_EQ(0, WEXITSTATUS(ws));
Here's the latency of the various execution strategies: FORK+EXEC+EXIT+WAIT │ APE │ APE-LOADER │ BINFMT_MISC
───────────────────────────────────────────────────────────────
fork() + execve() │ 55µs │ 66µs │ 56µs
vfork() + execve() │ 25µs │ 446µs │ 35µs
/bin/sh -c ./ape (1st) │ 485µs │ 457µs │ 170µs
/bin/sh -c ./ape (avg) │ 148µs │ 466µs │ 159µsThen for my musl-gcc hello world, I did, "for i in {1..10}; do o//examples/rusage.com ../hello" and the lowest wall time I got was 120us. But most of the runs were closer to 350us.
So, my attempt to answer your, "Why does hello world take 1.8ms to run?" question is: a) they timed it with "time" which measures lots of overhead, b) there's something wrong with their setup that means they are mostly timing the time taken for a CPU core to wake up.
Edit, and also no, cycles are not "more about instruction set". They're pulses of electricity that make their way through a chip, they are about as far from instruction-set-specific as it gets.
static inline unsigned long ClocksToNanos(unsigned long x, unsigned long y) {
// approximation of round(x*.323018) which is usually
// the ratio between inva rdtsc ticks and nanoseconds
unsigned long difference = y >= x ? y - x : ~x + y + 1;
return (difference * 338709) >> 20;
}
That's just one quick and dirty way your approximation could be approximated. Should work great for any benchmark games. Obviously don't use it for like, an X-ray or something. You can get the RDTSC value as follows: static inline unsigned long Rdtsc(void) {
unsigned long Rax, Rdx;
asm volatile("rdtsc" : "=a"(Rax), "=d"(Rdx) : /* no inputs */ : "memory");
return Rdx << 32 | Rax;
}
Intel also has recommendations for things like using CPUID and memory fences if you need to reign in speculative execution, but they can be costly. If you want to go even deeper then RDTSCP has a nice feature that lets you know if the operating system switched you to a different core during your measurement interval because it gives you TSC_AUX.Here's something interesting. If I run my test in a RHEL5 (Linux c. 2010) VirtualBox VM then it actually goes faster than a modern Linux kernel running on bare metal, taking only 14µs. If I run it in a RHEL7 VM then it needs 131µs.
So >1ms hello world execution should be a red flag. If that's actually the latency of GitHub's kernels and can't be explained by something like a shell script doing the time interval measurements, then it'd be interesting to learn what it's doing.
Here's the measurement code, it appears to be significantly more complicated than a simple fork/exec/wait loop but that could just be all the C# getting in the way: https://github.com/hanabi1224/Programming-Language-Benchmark... Note that we are definitely measuring the C# async runtime to some degree. Nevertheless you are probably right that the bulk of this 1.8ms is in the executable under test, and it truly is just bloat. Running `hyperfine ./empty-main-function` from rustc on my Mac gives 0.8ms.
Because instruction set is the interface and the architecture is the implementation? Or something else?
One important component of a program's performance, especially in numeric code, is how many instructions it manages to execute per cycle. If you have written your program's instructions in a 'good' order, then the CPU will e.g. not queue up many divisions in a row that depend on each other. It can do two different things at the same time. You can't do anything in less than one cycle, but you can do two of the same thing in the same cycle if you have two available execution units for that task.
The most direct answer to your question is that different instruction sets have different ideas about how much work a single instruction should do. The basic divide is between CISC (complex instructions, x86 qualifies) and RISC (basic instructions, ARM). It is a dramatic oversimplification but one instruction in a CISC architecture might involve 10 instructions in a RISC. So you can't compare instruction counts between architectures. But you can compare cycles and still make a bit of sense.
If this is interesting to you, you should try this game, where you can build a CPU purely out of the most basic logic gates: https://nandgame.com/