Hello World
drewdevault.com
drewdevault.com
I agree that software bloat is a big problem, but trivializing that problem to printing a "hello world" to the screen, punishing all languages with runtimes by measuring the syscalls involved in their startup routines, disregarding the fact that many users are going to have a single system-wide runtime for e.g. C or Python or Julia and therefore the total-kB number does not scale linearly with the number of programs written in C or Python or Julia, ignoring the massively increased software development and debugging time for writing in low-level and memory-unsafe languages like assembly, static Zig, or C, and directly implying[0] that most problems with software complexity can be solved by writing in assembly, static Zig, or C rather than in Julia/Ruby/Java/all other languages from the bottom 90% of the list (and that's the vibe that this post gives me) is, for me, more about making a venting shitpost than creating something that provides even a part of an actual solution to software bloat in general.
The "more time your users are sitting there waiting for your program" statement is especially amusing to me. Your users are not going to wait shorter for your program because you and your team are taking another year to write, test, and debug it in assembly.
[0] "These numbers are real. This is more complexity that someone has to debug, more time your users are sitting there waiting for your program, less disk space available for files which actually matter to the user."
It proves that the compilers for many programming languages emit a lot of instructions and a lot of system calls they could have optimized away theoretically.
This is particularly true when there are lots of "startup routines". When the program to be run is known, almost all of that startup can be found out to not be necessary, and stripped from the final executable.
> Your users are not going to wait shorter for your program because you and your team are taking another year to write, test, and debug it in assembly.
They might, if your program is interactive and is used often enough. Now, writing assembly is not very realistic, but considering a different language with less overhead - is.
It’s far easier to work in and translate thoughts to more feature rich languages. The vast majority of the time it’s not the language itself that will be an optimization issue down the line but rather the written code.
Sometimes language matters! Not always, but frequently!
Or you commit atrocities like rip the garbage collector out of Python. I’m not joking; Instagram did this: https://instagram-engineering.com/dismissing-python-garbage-...
A lot of games in gc-ed languages turn the gc off, and then manually run the gc whenever there's some spare time in a frame. The same thing should more or less work for a server handling web requests too. Maybe run it after every request or perhaps every couple of requests depending on memory usage.
For a hello-world program, this is true, you do not need any of the features that usually come in the language runtime - no garbage collection, no runtime memory safety, no runtime introspection, et cetera. This stems directly from the fact that many of these languages have borrowed features from Lisp, which has the whole language (GC, compiler, debug functions, etc..) always available for the code.
And that, considering your point of view, might be a feature, not a bug - especially if your program, having a runtime of its own, actually produces its own stacktrace, allows you to interactively inspect values of variables throughout the stack, and - in case of some runtimes - is capable of opening up its own debugger for you to perform introspection and issue analyzing in.
Right, I think the point was that this isn't always necessary, in which case it becomes extra baggage that you have to carry around.
https://en.wikipedia.org/wiki/RollerCoaster_Tycoon_(video_ga...
"Sawyer wrote 99% of the code for RollerCoaster Tycoon in x86 assembly language, with the remaining one percent written in C."
I wonder how that code gets maintained.
Also - that game absolutely needs better resolution to do the coasters justice.
It's the lower-level languages that, for the sake of efficiency and minimalism, cannot handle systems larger than a hello-world and actual programs written in them without long development times and a mass of bugs that come straight from their inability to abstract when necessary and from their memory management systems that lack sanity and safety. This "simplicity"[0] is often fetishized, and this blog post looks no different to me - look how many overflow bugs in operating system kernels this "simplicity" has caused.
Or, paraphrasing your quote, if your language prevents you from creating most basic abstractions, how can you trust it's doing well on actual programs?
[0] I can agree, writing in a memory unsafe language is simple, as long as it's someone else who maintains the code you wrote, accepts or reroutes all of the bugtickets that happen, and analyzes heisenbugs that come from weird memory corruptions caused by unsafe code.
Most compilers (i assume) optimize by applying specific optimizations to patterns they find in the source. most of the time this is pretty good compared to non-optimized code. but the other day on HN i learned about polyhedral compilation, which is pretty crazy of an optimization. it got me thinking: what about all the optimizations we don't know about? is it possible to find the absolute best optimization for a program given some specifications of what we want from it?
I think an ideal compiler would brute-force optimize like this so we don't miss out on optimizations we didn't think of:
1: compile the code normally without optimizations
2: generate a binary [see note1], and test if it produces the same output as the binary from (1)[see note2]
3: sort through the binaries and choose the best ones (perhaps some binaries are more optimized for speed, size, etc so there may be tradeoffs)
note1: here are a couple ways i think you could do this, in order of practicality:
A: modify the binary from (1) a bit at a time
B: modify the binary from (1) X bits at a time, or any of the other binaries from (2) that passed the test, more likely to modify binaries that perform better in the tests.
C: iterate through every possible binary up to around the size of the binary from (1)
note2: the easiest way to implement this would be to test it with a ton of sample input, but it would be better if the binary could be analyzed somehow to see if it satisfied the parameters of the source code.
I wouldn't use a such a brute-force optimizer to actually produce code which would then be undefined, but you could use it to reverse-engineer optimizations and implement them in source code.
edit: formatting
It was more difficult to write a huge program in C. But you needed a smaller team. You are so deep in the system management is always over their heads.
Perhaps it proves that if one is writing relatively small, simple programs, some of these languages are overkill.
go build -ldflags '-s -w' -o test test.go
Also, since he's complaining about the difficulty of obtaining a static build in Go: There's an issue for that. (https://github.com/golang/go/issues/26492) Drew definitely knows, he left an incredibly unhelpful comment there yesterday.I guess I could have written: I think someone should have wondered whether this was actually a reasonable way to make a statically linked binary before.
Maybe if you’ve gotten used to invoking the magic it doesn’t seem quite so arcane anymore.
Nim:
$ cat hello.nim
stdout.write("hello, world!\n")
Static (musl): $ nim --gcc.exe:musl-gcc --gcc.linkerexe:musl-gcc --passL:-static c -d:release hello.nim
$ ldd ./hello
not a dynamic executable
Execution Time: 0m0.002s (real)
Total Syscalls: 16
Unique Syscalls: 8
Size (KiB): 95K (78K stripped)
Dynamic (glibc): $ nim c -d:release hello.nim
$ ldd ./hello
linux-vdso.so.1 => (0x00007ffc994b6000)
libdl.so.2 => /lib64/libdl.so.2 (0x00007f7c88785000)
libc.so.6 => /lib64/libc.so.6 (0x00007f7c883b8000)
/lib64/ld-linux-x86-64.so.2 (0x00007f7c88989000)
Execution Time: 0m0.002s (real)
Total Syscalls: 42
Unique Syscalls: 13
Size (KiB): 91K (79K stripped)
Which I think is actually pretty reasonable for a high-level GC'd language.Seems suspicious that lines up with the 95.9 KiB the author listed for C + GCC + musl static build even though the author says they stripped the binary after. I think they might have copied the wrong number into the table :).
The author was counting the size of dynamic as binary + dynamically linked files. Should be about the same as the c dynamic ones in the table in this case anyways but just a note to anyone else running their own tests.
-d:danger --opt:size
On lobste.rs, someone also listed results for OCaml and they were quite nice IMO too.I thought it was a breezy read with a simple thesis. I don't know why such a thing should be discouraged.
Why must all blog posts be attempts to change the world?
You're looking at the numbers for PyPy, an alternative Python implementation written in Python-ish compiled to C which provides a JIT compiler for Python. A bit more understandable why that initializes slowly (though runs faster for longer programs).
OpenJDK 11.0.5 JRE
38 bytes source code, 970 KB binary (782 KB stripped) 0m0.003s execution time, 139 syscalls (26 unique).
Seeing the list of syscalls is the most interesting part of this whole exercise; the number of milliseconds it takes to print "hello world" is not super relevant (except in the few cases where start-up time is painfully long).
That gave me a good chuckle towards the end.
It'd be useful to break this out a little further as it'd have been interesting to see how small just the output is on the dynamically linked versions instead of just comparing static to whole dynamic bundle.
It's also a bit odd that e.g. zig gets optimized for size and stripped via the compiler, c gets optimized for speed and stripped via strip, and Go/Crystal just gets built standard with no stripping at all. I don't think it'd change the big picture just a bit odd.
.
Unrelated tangent/ramble, I played with Zig and Go as part of my yearly "take December off and tinker" break. Zig was really fun to work with but unfortunately still in a huge churn and development. Go was a lot better than I expected it to be (I had put off messing with Go for a few years now) and the size of the stdlib is just astounding. In the end it wasn't as "fun" as zig but it had very low friction and I definitely see myself using it for a few personal projects over the next year... and then seeing if Zig has less churn in December ;).
What is the point of this post? Yes, I fully expect a simple Hello World in assembly would be straightforward and fast. I still want the advantage of things like automated memory management, an interpreter or JIT compiler where warranted, a standard runtime environment, etc. For anything even remotely complicated.
I get it, over the past 30-40 years we've built layers upon layers of abstraction, so it's worth it to take a look back and ask "Are there some cases where we overdid it?" Still, let's not throw the baby out with the bathwater, or forget why we added those layers in the first place.
I think it’s sort of like how a lot of people fetishize a party lifestyle in their 20s and age out of it often when they get more perspective and understand bigger picture priorities.
$ cat a.c
#define _GNU_SOURCE
#include <unistd.h>
#include <sys/syscall.h>
void _start(void) {
syscall(__NR_write, 1, "Hello world!\n", 13);
syscall(__NR_exit_group, 1);
}
$ gcc -nostartfiles a.c -o a -static
$ ls -l a
-rwxr-xr-x 1 user user 9584 jan 5 00:07 a
$ strip -s a; ls -l a
-rwxr-xr-x 1 user user 9000 jan 5 00:10 a
$ strace ./a
execve("./a", ["./a"], 0x7ffd4d4f8250 /* 67 vars */) = 0
write(1, "Hello world!\n", 13) = 13
exit_group(1) = ?
+++ exited with 1 +++
I'm not actually sure why syscall() is inlined and glibc code is not included here while compiling with -static, but well, it works. Maybe it's because syscall() is a macro, or maybe it's some kind of ld code minimization (removing unused/unreachable code) trick which seems to be common in modern static linkers (ld, gold, lld).> #define _GNU_SOURCE
If it's portable, what's the point of defining _GNU_SOURCE? Honest question.
_GNU_SOURCE is requested by man 2 syscall. Interestingly it mentions it appeared in 4BSD, so possibly it might work (maybe with some changes) under other Unixy platforms.
$ gcc -nostartfiles a.c -o a -Wl,--entry=_start -static
$ ./a
-bash: ./a: cannot execute binary file: Exec format error
$ file a
a: ELF 64-bit LSB executable, x86-64, version 1 (SYSV), statically linked, not stripped\
But the final binary is impressively small, and it seem to contain all relevant code (body of _start and syscall funcs) $ strip -s a; ls -l a
-rwxr-xr-x 1 user users 2688 Jan 5 07:53 a
Works under FreeBSD though: $ clang -nostartfiles a.c -o a -static
$ ./a
Hello world!
$ strip -s a; ls -l a
-rwxr-xr-x 1 user user 9216 Jan 5 07:59 aEdit: I'm just curious and don't know how to even start testinng this, not trying to promote/demote Node.js in any way.
This is on nodejs v13.5.0
EDIT: My previous comment was made using an old version of nodejs, the update halved the number of syscalls from being around ~1350
The size would be a few bytes larger since node is scripted and that's more characters.
And well, I am thinking you're talking about "much faster" for small amount of text (small enough to not fill the stdout buffer). Actually printing huge amount of text to stdout should be much faster than printing to stderr exactly because of this buffer.
If they're not emitting the most efficient code on "hello world", that's a thing. It's not groundbreaking, but, I'd like whatever language I choose to be efficient with the small things as well as the large things.
I know what you're going to say. "printf does a lot more". Ok. But can I statically analyze: printf("hello world"), and notice that it's not doing anything interesting with zero ambiguity?
I was responding to this statement:
> This appears to have been written by someone who thinks the point of "Hello World" is to print the string "Hello World" as efficiently as possible.
To me... that would straightforwardly seem to be the point of writing a hello world program.
The blog post isn't even about how optimized the end result is. It's to get you, the programmer, to think about the cost you incur by going further down in this table when choosing your language/environment. Sometimes that's fine; Drew himself writes a lot of Python for example for his web stuff because it's the best choice for it and writing it in C or Zig is a pointless effort. The key thing to note though is that he picked Python while being well aware of this table.
Lots of programmers today aren't aware of this table. That's the point of the post.
It probably already does. In a real program run with a real compiler you cannot reasonably know what the optimization process is going to do. You have to look at the output.
>It's to get you, the programmer, to think about the cost you incur by going further down in this table when choosing your language/environment.
>Lots of programmers today aren't aware of this table.
Please don't. You don't have to be a CS prodigy to notice the size of an executable or the fact that a program is starting slow.
Put that way, most of this post seems like a tautology: if you misuse the tools you are given, of course you're going to get bad results!
It seems reasonable to me that a language should make the assumption that the programmer's use case matches the languages strengths, so by default any runtime setup / bookkeeping / teardown code should run. Failure to remove that extra "complexity" isn't a failure of the toolchain, it's a failure of the programmer to select the right tool for the job.
If the point the author was trying to make was that the complexity being added is never useful, this is not a post that argues that position. A cost / benefit discussion of the specific behaviours being supported by that complexity would be very interesting!
I'm told that the story behind it, is that he was arguing with someone about the efficiency of an operation, and actually wrote that site to prove his point.
I feel it needs some kind of normalizing. I get that it is illustrating bloat but it doesn’t really illustrate where that’s coming from.
Maybe only the output of the JITs should counts, or the syscalls required to assemble the example should be included. Are musl and glibc really wasting cycles or are they doing something that the example is missing.
Fun to think about.
No, he runs the assembly through NASM + GCC as documented on the page.
I think it's a comparison of "when the user runs the program what runs, how long does it take and how much disk space does it need" based of the column headings. It's not a comparison of the tooling prior to the user's computer as far as I can tell.
This is...highly context dependent. For highly polymorphic code, my understanding is that JITs can outperform precompiled binaries, since they can inline virtual/polymorphic calls in tight loops.
This also isn't a "performance" test in any real sense. It's a test of startup time. Where, yes, JITs lose, but unless you're writing short lived interactive command line tools, or something that runs on lambda, that shouldn't be a concern. For "normal" serverside or desktop apps that run for more than, say, 30 seconds, the difference between 0s of startup and 0.2s of startup time is literally in the noise.
The resulting generated binary? Well no, a python binary is smaller than the c binary. The toolchain? Well gcc is pretty complex and that's unaccounted for. The build process? Again, no.
The closest thing I can think of is the language runtime. But why do I care about how complex the language runtime is? Often more complex language runtimes make my life easier anyway, and they're all sitting atop the Intel microcode magic box anyway.
There's a very specific definition of complexity you're using, and I'm still not sure what it is. In my world, you usually add complexity to eek out extra performance by breaking the less complex abstractions.
If you haven't given a clear definition of what "the system" is, I can't really use your evaluation to influence my decision making.
Figuring how where the binary size cliffs are (what causes size to grow a lot) and which cliffs it's practical to avoid might be useful.
It stands to reason that the default use case of Perl is well optimized.
#![feature(start, lang_items)]
#![no_std]
#![no_main]
#[link(name = "c")]
extern "C" {
pub fn puts(s: *const u8);
}
#[no_mangle]
pub extern "C" fn main(_argc: i32, _argv: *const *const u8) -> i32 {
unsafe {
puts(b"hello world\n" as *const u8);
};
0
}
#[panic_handler]
fn panic(_info: &core::panic::PanicInfo) -> ! {
loop {}
}
#[lang = "eh_personality"]
extern "C" fn eh_personality() {}
2x assembly isn't bad, but I'm willing to bet we can go smaller.(See also http://mainisusuallyafunction.blogspot.com/2015/01/151-byte-... which was the original implementation, this shaved a few more bytes off)
I.e. you want your compiler to emit code that isn't portable across unix-like systems (and their future versions) in any degree? AFAIK Linux is the only unix-like system that guarantees ABI stability.
I'm curious what Zig does in this case that it got so "good" results. Does it forgo portability?
And yeah, I think it's an okay intro to syntax because at least it shows you the minimum boilerplate to get some output.
If you don't want that behavior, you can control this all yourself with write! and friends.
EDIT: deeper analysis on Reddit: https://www.reddit.com/r/programming/comments/ejxwlu/hello_w...
I chuckled.