Making Rust binaries smaller by default
kobzol.github.io
kobzol.github.io
I do love Rust but binary sizes have always annoyed me greatly and I always had this nagging feeling that part of all programmers don't take Rust seriously because of it. And I actually have witnessed, several times in the last 2-ish years, older-school programmers berating and ignoring Rust on that basis alone (so the author is quite right to call this out as a factor).
Looking at the https://github.com/johnthagen/min-sized-rust repo, final binary size of 51 KB when compilation / linking / stripping takes stdlib into account (and not just blindly copy-pasting the 4MB binary blob) is acceptable and much more reasonable. I wouldn't care for further micro-optimizations e.g. going to 20KB or even 5KB (further down the README file).
I also don't use nightly in my Rust work so I guess I'll have to wait several more years. :(
A feature that's landed in Rust nightly will be part of the next beta release (at most 6 weeks away) and then the following full release (exactly 6 weeks away).
For this feature in particular, rustbot added a tag of 1.77.0, which is releasing on 21st March 2024, less than 2 months away.
It's possible you've confused this with more complex features that stay on nightly for a long while they are tested. This is not one of those features.
A very relevant example is `build-std`, which builds the standard library (and does LTO) instead of copying a pre-built one. This feature has been on nightly for at least a couple of years.
Are you saying that the feature still hadn't "landed in Rust nightly" until recently? If so then what's the difference between a feature just being available in Rust nightly, vs having "landed"?
The Cargo doc page on `build-std` says that it is an "unstable feature" which will only eventually end up in stable once it is stabilized: https://doc.rust-lang.org/cargo/reference/unstable.html#buil... and Kobzol's linked post above indicates that `build-std` is "sadly still unstable."
I'm no expert on Rust compiler development, but my understanding is that all code that is merged into master is available on nightly. If they're not behind a feature flag (this one isn't), they'll be available in a full release within 12 weeks of being merged. Larger features that need a lot more testing remain behind feature flags. Once they are merged into master, they remain on nightly until they're sufficiently tested. The multi-threaded frontend (https://blog.rust-lang.org/2023/11/09/parallel-rustc.html) is an example of such a feature. It'll remain nightly only for several months.
Again, I'm not an expert. This is based on what I've observed of Rust development.
> For example, one thing that was noted is that if we strip the debug symbols by default, then backtraces of release builds will... not contain any debug info, such as line numbers. That is indeed true, but my claim is that these have not been useful anyway. If you have a binary that only has debug symbols for the standard library, but not for your own code, then even though the backtrace will contain some line numbers from stdlib, it will not really give you any useful context [...]
Optimizing binary size is a worthy and important goal because some people have uses cases where it matters. If you are not one of those people why fret about it?
But binary size is absolutely a cost. Whether or not it's a cost that matters to you is a different question, but it's still a cost.
Zero cost in terms of CPU and RAM usage during execution time.
Nothing is zero-cost if you look closely enough.
Reducing the binary size is usually more important than performance (and often more important than memory safety, if we are totally honest...).
Personally, I hate bloated software just as much as slow software. It's crazy what you can do with 64 kilobytes, let the demoscene blow your mind: https://www.youtube.com/watch?v=ZfuierUvx1A
Would be awesome if compilers could do the same :)
It's also important to remember that in embedded software, you're probably working in a no_std environment anyway, which means that the Rust standard library doesn't need to be included in the binary, which will already significantly reduce the size of your binary.
A lot of this is about tradeoffs, and I think the article does a good job of explaining what tradeoffs are relevant here. Yes, binaries should be as small as possible, but shipping multiple compiled standard libraries is also not ideal. The Rust team seem to have gone for a good default for people who are beginning, with lots of ways to tailor the build process for people who don't need debugging information, or who are willing to accept longer install/compile times in exchange for smaller, more optimised binaries.
We currently have some nasty C++/boost monstrosities, if Rust can deliver a better development experience AND reliable stack traces AND smaller binaries, it would be HUGE for embedded.
If that means that 90% of the stdlib is unused (which is likely true for many small projects) then it should not be included in the binary.
I remember this being the standard decades ago. It might not make sense in certain situations (i.e. with reflection), but that shouldn't be the default case anyway. What happened?
Every demo is really amazing when you encounter them for the first time, but AFAIK 64K demos were hottest around 2000 when Farbrausch revolutionalized the scene with .fr-08: .the .product [1]. Since then, 64K turned out to be too large because most 64K demos can be divided into multiple parts---engine, data and compressor---and each part can be individually developed. There are many 64K demos but far less engines and only a handful number of compressors in this level. It also means that there are only a handful number of people that can actually make engines and compressors. It's not a fit software, it's rather an unhealthily thin software. They are still awesome but they can't be a model.
But embedded Rust usually does not use stdlib. This thread started with:
> It's really a shame that Rust includes the stdlib piecemeal in binary form, debug symbols and all, in every final binary.
Which is not case for embedded.
That's probably an extreme case though, as nobody in their right mind would actually try to run Qt on an embedded device with any reasonable limit on resources. I'm sure rM only does it because they can afford to be wasteful - they have hundreds of megabytes of space for the operating system, and all the latency-sensitive stuff is definitely not built on Qt.
Reason I know this is because they offer SSH access to the device and an SDK.
Over in C, plenty of people work on things where file size matters. It is a big deal. System constraints (embedded) wire constraints (far away systems where internet is slow and ephemeral), legacy systems where everything is going to be slow even moving that fat file around and you have to be present to update (ATM, ticket, vending machines)... The list goes on.
So it may NOT matter for the desk top, or mobile or servers, but that's a tiny fraction of the computing out there.
What I have is a wall for functionality.
Loading 700kb react blob (compressed mind you) so I can read a web page is a hard no. You want to give me a rich GUI to do data editing in a browser, bring it on, I'll take the down load.
I know that rust and tiny go swap code here and there, binary sizes getting smaller on rust might make me give it a poke for some of the more "cute" embedded/iot things I like to play with... (another place where small matters!)
If Rust really wants to minimize any overhead in spite of the necessity of backtrace supports, there are indeed many ways to minimize the cold section of executable. Even a simple compression will work---especially given that we already have a copy of miniz-oxide there! So try that if you are motivated, and I would more than welcome that effort, but Rust has way, way more important things to do than that.
(EDIT: I am talking about the `build-std` work here, not the default strip debug info flag.)
Which is one of the reasons why I tend to leave JS disabled in my browser. Too many web devs have no care or concern for these sorts of things, which often makes JS a real resource drain.
If we talk about ultra-low-power platforms, e.g. energy-harvesting IoT devices, 1MB is still quite a lot.
If we are going to argue that Rust can compete with C/C++, it needs to have similar performance, also regarding binary size.
C++ just labels a few specific features as available in a "freestanding" C++ standard library if you have one (on an embedded platform presumably you do).
This makes it very easy to know what you're getting in Rust's stdlib in #[no_std] because it's all of core, so e.g.
https://doc.rust-lang.org/core/primitive.slice.html#method.s... vs. https://doc.rust-lang.org/std/primitive.slice.html#method.so...
At first glance those are identical, but no, std is re-exporting every feature core had, but it also gains features, for example the stable sort function family only exist in std because they're using a temporary allocation whereas the unstable (ie equivalent items may be re-ordered) sort provided even in core doesn't do that.
Figuring out what you get in your stdlib with C++ often comes down to suck it and see.
Never had the time or dedication to actually verify this but I've been bitten by programs and OS-es that trash the cache too much and I've seen humanly perceivable lags because of it. But maybe in this case I am overreacting.
Another way of saying this is that the change to strip automatically is a small win for for disk space and RAM consumption, but I don't think it's going to improve performance dramatically? I could be wrong about this though.
Thanks for the nuance. If debug symbols indeed never go into the CPU cache then my remark is completely irrelevant.
Still, I hate big Rust binaries as much as the next guy.
I think that it depends on what sort of machine you're aiming for the binary to run on. I develop for a few platforms where a 1MB executable size would be completely unacceptable.
Where is this 4mb claim coming from? I just built a hello world on macos in release mode, with no special flags and the result was 400kb. Thats absolutely larger that it should be, but its a lot smaller than 4mb.
Is it really that much worse on linux?
Indeed we could always strip the binaries and I've done so on every Rust project I worked on.
Hrm, I read the comments of another user who contributes to Rust and it seems there is various formatting / panic / abort / symbol and location translation code that contributes to the size and currently there is not much that can be done about it.
Oh well.
The ~ 4 MiB of debug symbols the article talks about are for the whole libstd or just the portions that actually ended up in the binary?
I think the new default makes sense but I'd love to have the option to build a lean but debuggable release binary with just the needed symbols.
The second problem is dynamic dispatch. The overhead in "hello world" is from printing and panicking machinery (it handles closed stdout), and these features use vtables, so it's even harder to precisely analyze what's actually used and what isn't.
I know floats are full of scary corner cases, but… Assuming tens of bytes per an if statement, hundreds of corner cases just in formatting? Is it really that bad?
EDIT: oh, ok, so I guess it's because strip is "debuginfo" here, rather than "true".
[profile.release]
strip = true
opt-level = "z"
lto = true
codegen-units = 1
panic = "abort"
cargo +nightly build -Z build-std=std,panic_abort -Z build-std-features=panic_immediate_abort --target x86_64-unknown-linux-gnu --releaseAs mentioned in another thread, I've simply followed https://github.com/johnthagen/min-sized-rust
• `panic = "abort"` means that any panic terminates the program. This is not always desirable because you may want to catch and recover from panics, particularly in long-running servers.
• `strip = true` means that anything depending on DWARF would no longer work. Backtraces won't work, but also unwinding will no longer work (so this is disastrous if you haven't set `panic` above). The actual proposal has `strip = "debuginfo"` instead, so unwinding will work while backtraces won't.
• `codegen-units = 1` is the number of concurrent compilation jobs (cgu) in the LLVM codegen phase. A single cgu will significantly increase the compilation time, while allowing a bit more optimization. Otherwise this is okay.
• `lto = true` enables Rust-specific link-time optimizations across crates. The actual benefit depends on the set of crates linked, but it is significantly slower that many large enough projects wouldn't want it. It does benefit small programs like the "Hello, world" program the most though.
• `opt-level = "z"` is same to C/C++ `-Oz` and the same pros and cons apply.
[1] https://doc.rust-lang.org/cargo/reference/profiles.html#rele...
> This is not always desirable because you may want to catch and recover from panics, particularly in long-running servers.
IMHO a panic implies that execution cannot continue under any circumstances, and even any attempts for a graceful shutdown might be futile (if recovery is possible it shouldn't be a panic but done through regular error handling).
For a server process the best reaction to a panic would mean abort and clean restart.
> A single cgu will significantly increase the compilation time, while allowing a bit more optimization.
Increased build time is acceptable for release mode IMHO.
Rust panic is just a C++ exception in its implementation, and not every C++ programmer would terminate a process when an exception is thrown. Of course Rust panic is more resillient because Rust provides a memory safety and the logic error can be reasonably bounded as a result.
> Increased build time is acceptable for release mode IMHO.
I think the slowest possible configuration is at least 3x slower than the default, and that's too slow to be acceptable for most people. But you can always tune them up if you want---please note that this issue is all about defaults.
That's often infeasible, and creates a major DoS/reliability problem.
The server may be processing hundreds of different requests at the same time, and panic=abort will kill all of them. That creates a visible failure to other unrelated clients of the server, not just the offending request.
If you try to fix that and retry aborted requests, you'll retry the panic-inducing request and cause another failure (and it'll take several restarts to bisect out the offending request).
Plus a restart may be costly, require loading data, warming up caches, etc.
It's just way cheaper to catch a panic and return 500 to the offending request. Rust guarantees to panic before anything terrible memory-corrupting happens. Even if you don't trust it and would prefer to restart anyway, you have an option to gracefully hand over the traffic before the restart.
I suppose you could have a panic in one thread whilst others are still running fine. You may just want to shut those other threads down gracefully.
You shouldn't use them as a general exception mechanism. They aren't the same, even if under the hood they both use stack unwinding.
These are still okay, but then `panic = "abort"` shakes a lot of code off by making the final executable unable to print a nice stack trace (even without symbols) and instead immediately traps/abort()-s.
Edit: and GP also used "panic-immediate-abort", which also removes the dependency to std::fmt::format!() because it now silently abort()s without even printing an error string.
I just tried this. The stack traces on panic seem more or less the same with or without panic = "abort" in Cargo.toml.
For example, this program:
fn main() {
let v = vec![1, 2, 3];
v[99];
}
Compiled with panic="abort" outputs this stack trace: $ RUST_BACKTRACE=1 cargo run --release
Compiling rust-panic v0.1.0 (/Users/seph/temp/rust-panic)
Finished release [optimized] target(s) in 1.82s
Running `target/release/rust-panic`
thread 'main' panicked at src/main.rs:4:6:
index out of bounds: the len is 3 but the index is 99
stack backtrace:
0: _rust_begin_unwind
1: core::panicking::panic_fmt
2: core::panicking::panic_bounds_check
3: <alloc::vec::Vec<T,A> as core::ops::index::Index<I>>::index
4: rust_panic::main
note: Some details are omitted, run with `RUST_BACKTRACE=full` for a verbose backtrace.
[1] 95405 abort RUST_BACKTRACE=1 cargo run --release
Weirdly, in this test if I don't strip the binary, I get a larger binary size with panic="abort" than when I leave that out. That is surely due to a bug somewhere.One can get binaries pretty damn small (low-mid tens of kilobytes for a basic cli program doing something like hashing of a file).
Problem I've found with manually compiling std (which has ancillary benefits of being able to compile to a specific uarch) is it can break the compilation process when bringing in third-party deps. The config.toml (stored in $PROJECT_ROOT/.cargo) overrides cargo's behaviour for all dependencies as well - which may break those compilations.
Tbh, it's one reason I don't particularly rate the rustc+cargo toolchain - but for most people writing regular applications: just being able to do ```cargo build -r``` and not care about binary size, uarch optimization or custom llvm/rustc optimizations (PGO etc), most won't care.
I believe it does print backtraces then terminate, since the backtrace is printed via panic hooks, which happen before the actual unwinding.
It's comparable to C++ land -fno-exceptions (not exactly, but similar).
Ah, and it should be also noted that some non-fatal signals were also delivered only via panic. The best-known example is a memory allocation failure, which is recoverable in Rust but needed unwinding for a long time. Nowadays you have an unsafe but non-unwinding alternative.
[1] https://blog.rust-lang.org/2020/12/11/lock-poisoning-survey....
lto and codegen-units=1 have a huge compile time cost. For release for general distribution you should tend to favour them, but the release profile isn’t just about that activity (especially because debug/opt-level<2 is often just too slow to use while developing). You commonly want to create another profile for production distribution.
Abort on panic changes runtime behaviour by stopping you from catching panics, which will completely break some programs, and harm the failure mode of others, so that e.g. one defective route on a web server will suddenly take the entire website down for everyone (or, if you have a supervisor that can restart the server, at least disrupt it for everyone).
IME optimize for speed vs size usually isn't as clear cut as the name says though, sometimes smaller code does indeed run faster, but in most cases I've seen there's not much of a difference between -O3, -O2, -Os and -Oz.
For example, java has to set up a constants pool, parse it from the classfile, run init and cinit, resolve references, allocatea a frame, and so on, just to get to the entrypoint.
Compared to C without stdlib which needs to basically just run a few bytes to syscall write().
There are obviously massive advantages to Java for which these steps are needed, but you do pay for it.
See issue #13871
https://github.com/rust-lang/rust/issues/13871
I exaggerate, point is that we've come a long way and are still getting better. Different people look at different metrics and the more popular the language becomes the bigger the variety of metrics the come into focus.
[1] https://web.archive.org/web/20230708052006/https://www.gamed... (search for "The Programming Antihero")
#include <stdio.h>
int main() { printf("Hello, World!"); return 0;}
Of course, If I had read more than 3 words in the man page, the answer was easy to understand: "Os Optimize for size. -Os enables all -O2 optimizations except those that often increase code size", so you can't really get lower size than the default using only the optimization system, there's also "-finline-functions" included in -Os, but it won't help you in a Hello World. #include <stdio.h>
#include <stdlib.h>
#include <execinfo.h>
#include <errno.h>
void print_backtrace(void) {
void *traces[50];
char **symbols;
int num_traces, i;
num_traces = backtrace(traces, sizeof(traces) / sizeof(*traces));
strings = backtrace_symbols(traces, num_traces);
if (!strings) return;
for (i = 0; i < num_traces; ++i) {
fprintf("%d: %s\n", i + 1, strings[i]);
}
free(strings);
}
int main(void) {
static const char FMT[] = "Hello, World!\n";
static int EXPECTED = (int) (sizeof(FMT) - 1);
int ret = printf(FMT);
if (ret < 0) {
fprintf(stderr, "printf failed: %s\n", strerror(errno));
print_backtrace();
return 1;
}
if (ret != EXPECTED) {
fprintf(stderr, "printf failed: only %d characters were written\n", ret);
print_backtrace();
return 1;
}
return 0;
}
While this is still substantially different (for example, Rust's I/O buffering is different from C), this should be enough to demonstrate that this comparison is very unfair. gcc -s -Os -fuse-ld=lld a.c && ls -al a.out
leads to a 5496 bytes ELF though. Which is not much larger than just printf("Hello World!\n"), see sibling comment.I think the point is C "cheated" by including a lot of goodness (format, backtrace etc) in the shared library so they does not have to be copied to each binary.
$ gcc -Os a.c && ls -al a.out && size
-rwxr-xr-x 1 user user 15952 Jan 24 17:49 a.out
text data bss dec hex filename
1316 584 8 1908 774 a.out
$ gcc -s -Os a.c && ls -al a.out && size
-rwxr-xr-x 1 user user 14472 Jan 24 17:50 a.out
text data bss dec hex filename
1316 584 8 1908 774 a.out
$ gcc -s -Os -fuse-ld=lld a.c && ls -al a.out && size
-rwxr-xr-x 1 user user 4552 Jan 24 17:50 a.out
text data bss dec hex filename
1199 528 1 1728 6c0 a.out zig cc -Os -target x86_64-linux-musl hello.c -o hello
...which basically calls Clang under the hood, but comes with out-of-the-box cross-compilation support for Linux and MUSL creates a 5136 bytes executable.You're perfectly welcome to just use a raw for loop with numeric indexes in Rust , in which case you wouldn't have those extra function calls for iterators, etc. Likewise, if you implemented an iterator abstraction in C, you'd end up paying a similar cost. So, that's not what the grandparent comment was talking about.
But I do find it useful now and then to have debug symbols in production, so that monitoring and telemetry gets some context. But to me, it's rather logical that I need to add this to the build.
I would, however, love it when it's trivial to get these symbols set up external.
https://github.com/llvm/llvm-project/commit/36f01909a0e29c10...
And naturally we already have a port: https://crates.io/crates/debuginfod-rs
Looks more like a systemic issue in the Rust development process to me tbh.
What's more shocking though is that even after a 90% size reduction, a vanilla hello world is still 415 KBytes. That's about 10..100x bigger than I would expect from a low level "systems programming language".
Also, Rust has no direct platform support unlike C. So everything has to be statically linked to be portable. A statically linked glibc is indeed much larger than that (~800 KB in my machine). Conversely, you can sacrifice portability and link to `libstd*.so` dynamically to get a very small binary (~17 KB in my machine, both for C and Rust).
For instance a statically linked C hello world for Linux via MUSL (cross-compiled with `zig cc -Os -target x86_64-linux-musl hello.c -o hello` because I'm currently on a Mac) is just 5 KBytes.
[1] https://www.gnu.org/software/libc/manual/html_node/Backtrace...
Because in the end almost nobody actually cares about it enough to create a fix.
Small binaries are great, but people care mainly about how fast it compiles and how fast it runs, and in the few cases where the binary size is important it was already possible to shrink it significantly (more than with this new change). In my entire life I have never heard a user complain about the size of the binary, what people really care about is efficiency/speed at runtime. That's why people regularly mention VS Code using Electron, but not that its installation package alone has >500MB.
C was created in a time when binary sizes used to be important for every program. These days, it only matters for a small subset of use cases.
If anything, the fact that no one fixed it for this long is an indicator that it doesn't matter much for the kinds of things people typically build with rust.
It's nice that there's a fix now, but would I care if Firefox or Chrome binary was 4MB bigger?
The `strip` was added to rust nightly in 2020.
1: https://github.com/johnthagen/min-sized-rust
2: https://kerkour.com/optimize-rust-binary-size
3: https://rustrepo.com/repo/johnthagen-min-sized-rust
4: https://sing.stanford.edu/site/publications/rust-lctes22.pdf
5: https://arusahni.net/blog/2020/03/optimizing-rust-binary-siz...
That happens all the time, it's called prioritizing. If you don't let people prioritize, they will burn out and leave the project. That's not what you want.
after all: without debug info (which you do not want to ship to customers necessarily), you cannot do profiling or debugging in any meaningful way...
[1] https://doc.rust-lang.org/cargo/reference/profiles.html#spli...
Is this really a thing? Do C folks make fun of Rust?
Anyone who’s been doing this for long enough is a polyglot.
Reddit gave me some interesting insights into people like this. Whenever I come across a very radical opinion, I check their whole profile. It is often someone without much if any experience, and clearly doing it just for the sake of tribalism.
Because despite its complexity, slow compile times, lack of a specification or more than one viable implementation and bloated binary size, people still tout it as a C replacement.
For application programming, Rust is fine, but for embedded and systems programming, nearly any of those on its own can be enough to eliminate Rust as an option, depending on the situation.
Except for complexity, they are all solvable and being worked on, for this very reason. But Rust can never replace C completely, because it's not simple, and there is a sizeable portion of primarily C programmers who are minimalists.
In this comment I describe what I did to remove the bloat: https://news.ycombinator.com/item?id=36394426
It's not the default and it took me some searching and reading to get there. I wonder if Cargo could have some vastly different default per target. As everybody said in this thread, defaults matters.
Hmm, maybe that's the fundamental difference in thinking; the people I'm thinking of would reach for the simplest tool that can feasibly handle the job, even if it's somewhat harder to use.
I am sure a pure assembly implementation would have saved some room in the FLASH, but the generated machine code wasn't that bad for what I can tell with my limited experience anyways. And it sure was nice to write it in Rust. Except a codegen bug in the version of Rust at the time ;) That one was painful!
won't this destroy stack trace support?
Why not instead only without?
Sure, you can rebuild the standard library, but it's a lot simpler to strip the output than set things up so the first time you want debug symbols, it has to rebuild std, cache it somewhere, not rebuild it again on the next build, but be sure to invalidate that cache the next time std gets updated.
And in general, not wanting debug symbols is the last step in development, before making your first release. Before then, you pretty much always want those debug symbols, except for when you're benchmarking binary size or something.
So if they were going to ship std without debug symbols, they'd probably just be better off not shipping a prebuilt std at all, as pretty much everyone would end up having to build it on first run of 'cargo build' anyway. (Which is maybe fine, actually?)
For the C++ standary library maybe, but for pretty much all others which provide ABI compatibility it's a concious and properly followed decision.
For C++, yes sure, but I thought C usually has a pretty well specified ABI on most platforms, no?
The compilation runs not from the SD card but from a chroot to a large external drive plugged into the RPi via USB.
Once I have GCC and rest of the toolchain built, I then use it to compile custom kernels with embedded ramdisks.
https://en.wikipedia.org/wiki/Self-hosting_(compilers)
Wish I had a computer with 32GB RAM but I don't.
I use strip -s every day.
When evaluating languages, you rarely look foor good enough results to work around.
If it was like 1GB, sure. But it aint.
Every company makes some effort to be green and eco friendly these days, yet we're perfectly happy wasting billions of hours of compute time every second because "who cares" and "we can"... Same mindset.
While I'm not a seasoned C or C++ programmer I have definitely done this to a few new languages when playing around with them. My thought was "If helloworld.rust is this large, it'll be huge once I've actually written more code".
https://git.musl-libc.org/cgit/musl/tree/src/stdio
Not sure if this can easily be replicated in Rust though and LTO must be used instead for dead code removal.
When the program grows, the fix costs stop to matter.