Minimizing Rust Binary Size
github.com
github.com
https://www.codeslow.com/2020/01/writing-4k-intro-in-rust.ht...
On the other hand, reified generics increases the binary size. The advantage is that this is more digestible for the optimizer for inlining and specific optimization, which gives you more predictable performance than the alternative of devirtualization, but of course size can also have negative consequences, like for example slower compile time. One good example is this recent PR that improved compile time by making parts of the Vec implementation in the standard library non-generic: https://github.com/rust-lang/rust/pull/72013
2. Rust doesn't want to commit to a stable ABI yet, so even if an OS wanted to ship libstd, it'd be logistically difficult due to libstd version having to exactly match compiler version.
2. is a fair point. At this point in Rust's life I see how avoiding binary conpatibility questions makes sense. Is anybody even working on it, though? Otherwise Rust could be making system-provided, dynamically linked libraries almost impossible due to earlier design decisions without consideration for BC.
There is no work on MVPs or RFCs for now, but stable ABIs are being discussed currently: https://internals.rust-lang.org/t/a-stable-modular-abi-for-r...
They're very much possible; they're just limited to using the C ABI when interacting with dynamically-linked code. More complex features can be implemented by providing thin wrappers as part of language-specific bindings.
In what scenarios is optimizing for binary size preferred over optimizing for speed?
64kb intros:
https://www.reddit.com/r/rust/comments/597hhv/logicoma_elysi...
SQLite recommends optimizing for size, because it yields a significantly smaller binary with minimal speed impact: https://www.sqlite.org/footprint.html
In that case having a big exe file can really hurt start-up performance, so optimizing for that might be preferable if your application is not very speed sensitive.
[1]: https://docs.microsoft.com/en-us/windows/win32/api/winnt/ns-... (IMAGE_FILE_REMOVABLE_RUN_FROM_SWAP or IMAGE_FILE_NET_RUN_FROM_SWAP)
I see some of them change behaviour or make things complex, but I don't see a performance impact or anything that indicates software will be more crashy.
At the assembly level, it's sometimes true that a single instruction or smaller sequence of instructions takes more cpu cycles.
Analyze the different instruction sequences here:
https://www.nxp.com/docs/en/supporting-information/MC680X0OP...
This was not an exhaustive test of optimization combinations, just a single stack, but the results are interesting!
To read the table: default is no changes in the release profile, ie:
[profile.release]
# opt-level = 'z'
# lto = true
# panic = 'abort'
# codegen-units = 1
After that each additional line gets uncommented and rebuilt, then sizes recorded before and after cargo-strip, so the final line is all four optimizations applied.Results:
as generated after cargo-strip
default 8675776 5069328
opt-level = 'z' 9023200 4676112
lto = true 5943312 3586584
panic = 'abort' 5062456 3135928
codegen-units = 1 4747000 3013048
Tests run on ubuntu 20.04 (Linux d2836c103a22 5.4.0-37-generic #41-Ubuntu SMP Wed Jun 3 18:57:02 UTC 2020 x86_64 GNU/Linux) with rustc 1.43.1 (8d69840ab 2020-05-04)
cargo 1.43.0 (2cbe9048e 2020-05-03)
cargo-strip - reduces the size of binaries using the `strip` command 0.2.2
Fascinatingly, opt-level='z' produced a LARGER binary than the default, before stripping. That was unexpected.-Brian