Life without line numbers for 6% smaller Go binaries
commaok.xyz
commaok.xyz
This is great article instructing how to shrink the Go Binaries. Author is lead of cryptography and security team of Go Team at Google.
I performed the steps in the article and reduced the output binary size from ~35 MB for raw go build to around 29 MB for go build with the ldflags all the way down to 8.3 MB for compressed with upx.
Really impressive results.
I hadn't tested the startup times of the upx binary before, so I just did by running `time $binary_name --help`. For my binary, this should print a simple help message and quit.
Big: 0.04s user 0.01s system 101% cpu 0.054 total
Med: 0.04s user 0.01s system 99% cpu 0.048 total
Sm : 0.14s user 0.02s system 99% cpu 0.163 total
So you're right - it makes execution time take about 3x as long. However, in my case, 3x slower is a fine trade off to make.
But to answer your question, you're going to see some performance improvements if your binary can fit into lower levels of the CPU cache.
Kdb+/q for example is less than 700kb, which mean that the entire thing can fit into the L1 instruction cache on high end server CPUs. And that size is even more impressive considering its only dynamically linked to libc and libpthread.
This small size contributes to the out of this world performance the system has. And remember that's the language and the time series db.
35 MB is huge! Are there any assets inside the binary?
For context, 35MB is two and a half boxes of high-density floppy disks. There were whole operating systems, complete with applications, smaller than that.
Native Android doesn't need such big sizes.
go version go1.14.4 darwin/amd64
Looks like the Linux and Darwin versions of the binaries are only a few hundred kb different.
(The program we at Google use just to build our bloated programs is itself 180MB.)
But also, since this is client side code, a smaller download means smaller utilities on client machines. Which is nice.
go build -ldflags="-s -w" main.go
upx --brute main
Just tried that on one of my binaries and went from 12 mB down to 3.6 mB. func codeToError(i int) error {
switch i {
case ...
...
}
}
func stuff() error {
if rv := callJunk(); rv < 0 {
return codeToError(rv)
}
if rv := moreJunk(); rv < 0 {
return codeToError(rv)
}
...
return nil
}
You're going to get multiple whole copies of the inner function inlined into the outer one. This can add up, if the outer function is long and/or if the inner switch statement is huge.You can save a lot of code by jumping through an adapter function instead, like:
func callWithErrorsConverted(f func() int, g func(int) error) error {
return g(f())
}
... because now the compiler doesn't get to inline g into f. These things make a big difference in whole-program size. I wrote about it in the context of tailscale and protobufs here:https://paper.dropbox.com/doc/Bloaty-Puffs-and-the-Go-Compil...
Can you elaborate on what you mean by this? I’m new to go and want to understand the performance implications.
EDIT: and all registers are caller-saved. so when calling a function, all register contents will be saved on the stack (and restored later) even if the function doesn't touch certain registers
i don't think its intended as a stable target and will probably change at some point. the top google hit was a 3rd party site, and this issue is still open [#16922 doc: document assembly calling convention](https://github.com/golang/go/issues/16922)
btw the source i linked in my original comment links to a proposal to allow passing args in registers
There's a lesson in there about success, or something...
https://dr-knz.net/go-calling-convention-x86-64.html#differe...
This has never seemed like such a clear distinction to me. A language is essentially just an instruction set for a compiler.
It being a compiler problem, you can just accept the performance/size penalty and leave your code as is, but hope that in the future it will eventually change and the eventually really smart compiler will accept your current code as is and produce output without the penalty.
Language problem would means you have to rewrite it differently.
And as other comments call out, it is important to not conflate the two. A language issue is one that can't be easily fixed without breaking backwards compatibility. Compiler optimizations can be made (and is done all the time) without breaking the language.
The C compiler would do something similar here and also increase code size for performance unless you explicitly use `-Os`
Any inlining routine needs a cost model. From the article it seems like Go compiler's cost model is not yet well tuned.
That doesn't sound like Go's compiler is primitive, a primitive compiler wouldn't attempt the inlining at all. This sounds like a performance optimization which isn't something that a primitive compiler would attempt.
I went through most of the compiler about two years ago to see ,amongst other, how much of its time was spent during IO since both compile time and correctness of dependency handling varied wildly in virtualized environments.
Which is where I learned a lot, as for example that the correct way to invoke the compiler is essentially impossible to do with the supplied tools, but if you do it, it was actually possible to not recompile system libraries each time you build, that refactoring code where module scope is used for a most things is incredibly tedious. To compile to/from memory, as a service I had to change approximately 10 000 lines of code without being entirely done, as everything was tied into the module namespace and a little bit too concrete vfs like file access mechanism. Thus both the lifetimes of data, and where IO went were impossible to change without changing almost every type declaration site, hence often declaration, and attribute reference cite. Anything module based needed to be threaded through in appropriate places. Yes, and because some package names are hardcoded to be forbidden by parts of the toolchain unless it knows it is compiling itself, you have to patch the compiler (iirc) and the build tool if you want it to compile itself to a different name, which you kind of have to when you are doing something incredibly experimental. Otherwise it'd just overwrite itself because of gopath. Because of how the compiler/build works, it becomes contagious, so you'll need different names for everything. Well, you could probably get away with a little less renaming if you compiled the compiler with make, but at the point I realized that, I was already too far in for it to be any point in backtracking.
Hardest issue was when it deadlocked on a mutex for a while, turns out one the main compiler go routines concurrency were not implemented correctly, so without the implicit serialization of the disk IO, it deadlocked almost every compile.
Apologies for the digression, go was kind of forced on me, and some of the most touted "opinionated" parts lost me countless of hours of grief.
If you manage to supply exactly the packages (transitively), and not a wildcard more, it would avoid recompiling the system libraries on each invocation. One would think a compiler where so much effort was put into making it fast would make it easy to have it run fast? The primitives were sort of there, but they didn't really fit together, and I actually had to write quite a bit of script/code to make it work. Got very fast though, thought I broke it the first time I got it to work. <return> and instant prompt. Toke me a while to convince myself it was actually compiling anything at all.
Cut the recompile time in a VM to less than a tenth. Still slowed down by some pointless file IO (don't remember exactly what) but that was less of an issue.
Vanity package URL's though was what broke this camels back, after GOPATH and some other miscellanea had strained it. When I ... probably needed to quickly fork+fix an external package, and suddenly had this additional obstacle, I was so ... unhappy. Already too much work, and then you get more work for exactly no discernible reason at all. It's even called a vanity url officially.
(I'm not for/against the philosophy of the decision making, just pointing it out).
you could have two totally separate compilers, which go does, there is the gcc variant. i don’t know if it tends to produce faster binaries though.
Another thing Ive always found strange is people always recommend (-ldflags -w), when (-ldflags -s) gives smaller result, and (-ldflags '-s -w') is identical to (-ldflags -s).
No, Go doesn't use .pdb files.
On all platforms Go stores a significant amount of meta-data in the binary, in the same format so that they can access it using the same code on all platforms.
This information is used by the runtime. Precise garbage collector needs to know the layout of structures and stack to know which fields are pointers; To generate readable callstack Go needs info about function and where they come from (source code file name, position). etc.
A sidenote: you can convert DWARF debug info to .pdb using external tool which allows debugging Go programs with Windows native debuggers like WinDBG.
Is there a way to make `go build` for a particular project always use `-ldflags=-w`? Or do I have to remember to type it in every time (or hide my `go build` command inside a Makefile or `build.sh`)?
It's serving a certificate for github.com.
In an older thread you mention an acronym TFA. The thread was a discussion on sparse files and removing bytes from the front of a file.
What is TFA?