Mid-stack inlining in the Go compiler
docs.google.com
docs.google.com
It's not some weird, exotic form of inlining; it's just a more complete version than Go used to have.
What does it include? You've conveniently managed to skip over the crucial part with ellipses. Calls in the middle of stacks perhaps?
You people have too much time on your hands to go and pick at the wording people use in their presentations.
What kind of argument is that?
There isn't a standard term, because virtually all inliners always did what Go is now doing. It's pretty basic functionality.
But instead, you come along and criticize them for being too specific with their terms (!!) and also accuse them of being deceptive. This is not a marketing exercise. There is no conspiracy here.
Perhaps Ericson2314's gripe is that it's being posted (and upvoted) on HN, rather than just a Go-specific forum (e.g. reddit.com/r/golang), and by implication it's intended to be read by a more general audience.
Why would correct use of terminology be "pedantry at its finest"?
What people are concerned about is intellectual dishonesty. The presentation does not go to much effort to show, for example, why the term "mid-stack inlining" needs to be introduced for Go. Do other languages not have this kind of inlining? Does comparing and contrasting other implementations just not matter?
Because you're debating terminology when the wrong term caused no one confusion, thereby detracting from the actual conversation.
> What people are concerned about is intellectual dishonesty.
There's no cause for this concern.
> The presentation does not go to much effort to show, for example, why the term "mid-stack inlining" needs to be introduced for Go.
Sure it does; see slides 3-6.
> Do other languages not have this kind of inlining? Does comparing and contrasting other implementations just not matter?
Perhaps the author omitted it from his deck because it's impractical to cover the whole breadth of inlining in a deck that's already 35 slides long. Perhaps he simply didn't think to include it. There are a lot of likely explanations for why this wasn't included besides nefarious motives. You sound paranoid.
Well, why do you think we are bringing it up, then? Due to some nefarious motive?
As far as programming languages go, Go is very new. Of course, there will be a number of areas that still need to be optimized or fully fleshed out. I think the Go compiler has only been written in Go for about a year. I'm sure you you find this pretty lame as well.
I for one am very glad that more work is put into inlining
Since I've learned about continuation passing style (which Go channels could probably be formally transformed into), I've been convinced that there's a better way to do codegen. Better calling convention, better stack representation, better instruction architecture; I'm not yet sure - it's a nag continuously at the back of my mind, almost as though it's at the tip of my tongue. In this specific case, it must surely be possible to inline a continuation with some foreign architecture. I'd love to see some literature on the more experimental end of this stuff, if anyone has it.
https://www.amazon.com/Compiling-Continuations-Andrew-W-Appe...
More recently though I heard someone proved mathematically that CPS can be transformed one-for-one into one of the more conventional models. That doesn't mean it might not still be easier for the humans to deal with however.
Byproduct of that could be little hacky things like this below that make code faster by restructuring code a little
https://techblug.wordpress.com/2013/08/19/java-jit-compiler-...
You should be able to get about 15-20% with about 3-5% binary increase size.
In fact, with ThinLTO, we often see that gain with binary size decrease from smart inlining choices.
(The heuristics for inlining take a very very long time to get right and tune)
The issue they will next hit is that inlining is going to make the compiler slower until they tune the heuristics well.
1. The compiler is a lot slower, which is usually from code growth and not debugging info growth. If the compiler is that much slower from debugging info growth, they have larger issues :) 2. Usually people do not include debug info sizes in binary sizes, because DWARF/et al info can be stripped and put alongside the binary (IE it doesn't even have to be part of the binary)
So no, it doesn't say that it's "mostly due", it says ~25% is due to debugging information.
https://github.com/golang/go/issues/19386
Those numbers are just for my CLs that fix stack traces but with mid-stack inlining still off. Turning it on makes builds noticeably slower:
$ time ./make.bash
real: 45.32s user: 118.67s cpu: 5.85s
$ time GO_GCFLAGS='-l=4' ./make.bash
real: 64.51s user: 167.04s cpu: 7.12s
We'll need to tweak the inlining heuristic to find a good balance between performance, build times, and binary size.dl.google.com has been golang since 2013. Imagine the traffic that application gets!
That said, it's a little disappointing when runtimes require custom algorithms or metadata to walk the stack and construct a stack trace. It makes it harder to build debuggers that grok the state of multiple runtimes (e.g., the Go code and the C code in the same program). This also affects runtime tracing tools like DTrace, which by construction can't rely on runtime support for help.
So it records the calling convention, architecture flags, alignment, and other ABI pieces etc? As well as an estimate of instruction-level inlining cost, summary info about arguments, etc, so you effectively decide whether inlining it will help or hurt, without having the IR around to try?
FWIW: Writing out the info is usually not the hard part, actually, and is unrelated to the DAG-ness of the packages.
GCC is just the perennial example here, but they refused to write it out for years for political reasons, not technical ones :)
No. At the moment it records the AST in the object file, because the inliner works at the Go AST level. In the future it may instead record the SSA representation (which would obviously give better cost estimates; the current heuristics are really extremely simple).
"FWIW: Writing out the info is usually not the hard part, actually, and is unrelated to the DAG-ness of the packages."
The DAG-ness means it's always available when compiling the call site, even if it's a cross-package call. It means you don't have to do it at link time.
This would be identical to what others do then :)
"The DAG-ness means it's always available when compiling the call site, even if it's a cross-package call. It means you don't have to do it at link time. "
This is unrelated to DAG-ness. Unless you mean something else by DAG-ness. DAG-ness means it's a directed-acyclic graph. That is, all other things being equal, it has no cycles.
This is unrelated to the problem.
For example, in other languages/etc, it could be weakly defined, or other some form of overridable, regardless of whether it has cycles, is actually multiply defined etc. That requires link-time resolution, becuase you can't optimize it making an assumption about its callees/callers, or even inline it, and then just hand that version to others because you may have screwed it up in a way one of those other callers depend on.
That is, the overridability is an attribute of the function, not a problem of how it is used.
Ditto on the ABI, alignment, etc.
None of the interesting problems they have to solve are related to packaging. They occur with DAGs or non-DAGs, are just related to these languages supporting a richer set of things you can do to functions :)
Why is it any harder to do at link time?
(I've implemented this in a production compiler, and choosing whether to do it at compile time or link time was a trivial decision.)
Every good production C++ compiler has had some form of link time optimization for many years.
IBM's, for example, has been happily cross-optimizing between C++, java, fortran, PL/IX, etc without any issues, going on at least 15, maybe 25+ years now (I know it's 15 for sure, i suspect it's closer to 25).
That doesn't make sense to me. The complicated part is storing the IR in packages in a form that can be read back into the compiler later. That's needed to do inlining at all. Once you have that done, doing LTO is trivial: you just slurp your IR for all modules linked together into the compiler and emit a single binary. (You can be fancier, like ThinLTO does, but again, the effort needed to do ThinLTO is independent of whether you have cyclic dependencies or not.)