A small Rust binary indeed (2022)
darkcoding.net
darkcoding.net
unsafe { asm!( "mov edi, 42", "mov eax, 60", "syscall", options(nostack, noreturn) ) // nostack prevents `asm!` from push/pop rax // noreturn prevents it putting a 'ret' at the end // but it does put a ud2 (undefined instruction) instead }
and
> We will need to tell the C compiler that we’re providing our own entry point, telling it not to include it’s own start files.
So it's a Rust program but it's just calling inline assembly and using a C compiler?
This isn't only libc -- it also includes libgcc (or compiler-rt, depending on your system toolchain), which, despite the name, may still be called "behind your back" by the LLVM toolchain.
> So it's a Rust program but it's just calling inline assembly and using a C compiler?
Yeah, I think this article is more in the tradition of [0] (but trying hard not to drop rustc) than being completely practical advice on making the binary you ship to users smaller.
[0]: http://www.muppetlabs.com/~breadbox/software/tiny/teensy.htm...
800 lines of ASM file reduced to 600 lines of Rust, including comments and constants in both cases. He might be pushing the limits, and everything's unsafe Rust, but unsafe Rust is still safer than raw assembly.
I don't think I would go that far. Assembly doesn't have undefined behavior, and especially not with the strict constraints around references as in Rust. The safe/unsafe dichotomy in Rust is better than only using C or C++ when there are concise, robust encapsulations around broken invariants.
Certainly some assembly languages do.
They do, because the machine code sometimes has undefined behaviour. E.g. on the 6502 famously 1/4 of the instructions are undefined (all of the 0b.....11 ones), and many of them behave differently on different implementations of the processor (up to and including halting it, or placing it in a strange state).
As someone who has written a fair amount of assembler over the years... Yes, it doesn't have undefined behavior, but it also lacks practically all guard rails and safeties.
The smallest error and you might do things like completely messing up your call stack – just need to forget one "POP" or mess up with stack pointer adjustment. Or for example a computed jump in the middle of an instruction.
You can create bugs that can be almost impossible to figure out from a crash dump that even something as low level as C will effectively protect you from doing.
In a lot of conversations around Rust binary sizes some people extrapolate from the "Hello, World!" size difference as if the additional cost on top of a bare C binary was linear, when in reality it is (approximately) a constant cost. That on top of completely disregarding that the "bloat" is doing something (panic machinery, string formatting, DWARF symbol storage, DWARF symbol parsing, etc.).
The standard library has 4MB of debug info baked in, which due to its special integration with Cargo is always added, even when you explicitly configure `debug=false`. This is what usually surprises people and makes Rust executables seem huge.
The point is that it'll never be doing something, and the compiler can clearly see that, but decides to add that dead code anyway.
(\* t.ml \*)
let () = print_endline "Hello, World!"
Then just doing a standard compilation and a strip: $ ocamlopt -o t t.ml && ls -l -h t | cut -d " " -f5
1.5M
$ strip t && ls -l -h t | cut -d " " -f5
356K
$ ./t
Hello, World!
I may be overlooking something, and would be interested to learn what if so, but I was surprised we got a result smaller than the rust binary in the first instance.Source: https://kobzol.github.io/rust/cargo/2024/01/23/making-rust-b... HN discussion: https://news.ycombinator.com/item?id=39112486
edit: It doesn't look like any of the techniques listed after that work if you still use the standard library.
We went from 3.6 MiB to 400 bytes.
In an ideal world, you'd get those 400 bytes from the compiler when you set it to optimise for minimum size and give it the same "simplest possible Rust program", and without optimisation the output might be a little bit larger, but not 4 orders of magnitude larger.
Frankly, 3.6MB is nothing but insane for a program that does little more than exit. That's more than 2 floppies! We can add "Rust program that does nothing" to "List of Things That Turbo Pascal is Smaller Than" (https://news.ycombinator.com/item?id=22843140)
I wonder what it is that leads to such inefficiency. It would seem reasonable that a program that does very little, should also not contain much code. Therefore a compiler shouldn't generated much code. Yet somewhere along the way, drenched in multiple layers of abstraction, we've lost common sense?
A program that returns 42 is a 5-byte file in DOS: b8 2a 4c cd 21. (And if they'd been a little more thoughtful on the initial API, it could've been 3 bytes.)
My understanding is that there is little corporate funding available to resource compiler engineers to prioritize "do nothing and exit" workloads.
Today I got rid of libc on the Windows version of a commandline tool to flash firmware via USB, which freed 7 kB of the .exe size.
The original version was done in C++ plus Qt and was ca. 3.5 MB (.exe and dependencies).
The optimized C version is 14 kB compressed with upx.
The optimal assembly code is this: "mov al, 42; xchg edi, eax; mov al, 60; syscall". Or just "mov al, 60; syscall" if you want to exit with 0 rather 42.
If not that feels a bit like baby with bath water