Go compiler: Initial support for concurrent backend compilation
go-review.googlesource.com
go-review.googlesource.com
All benchmarks are from my 8 core 2.9 GHz Intel Core i7 darwin/amd64 laptop.
First, going from tip to this CL with c=1 costs about 3% CPU and has almost no memory impact.
Comparing this CL to itself, from c=1 to c=2 improves real times 20-30%, costs 5-10% more CPU time, and adds about 2% alloc.
From c=1 to c=4, real time is down ~40%, CPU usage up 10-20%, alloc up ~5%
Going beyond c=4 on my machine tends to increase CPU time and allocs without impacting real time.
Go is a great language which combines C-like speeds for low-level code with a good concurrency model, making your program comparatively easy to parallelize on a per-function basis. As the work on the Go compiler shows, it is not trivial for complex programs, but definitely possible.
I'm in the market for a new developer desktop and if it shakes out the way it's looking an 8 core/16 thread Ryzen is hitting the price/perf sweetspot.
No, Go doesn't give you simd/vector primitives for data parallelism.
What do you mean by "finally pushing 8 cores"? AMD has been selling 8 core CPUs for years with their AMD FX 8xxx line.
I am using Go in production. It does not combine C-like speeds (I've comparison after rewriting one of the services from C to Go), but there is a lot of space for additional optimizations in future. Go will never be as efficient as C because of additional indirection in certain cases and additional bookkeeping. But in many cases it's faster than Java.
- in general, understand which types are values, and which have to be allocated. E.g. the slice structure itself is a value type (typically 24 bytes), but it points to an array, which might be heap allocated. Re-slicing creates a new slice structure, but does not touch the heap - extending a slice might.
- to analyze your code, use go build -gcflags="-m", this will print out optimization information, especially which allocations "escape to heap", which means, they are heap allocated and not stack allocated.
- Interfaces are really great, but copying values of interface{} does involve heap allocation of 16 bytes. Also method invocation via interfaces has some overhead. Method invocation of non-interface types is a fast a plain function call.
- profile :)
http://benchmarksgame.alioth.debian.org/u64q/compare.php?lan...
This is not really true, at least not as such a general claim.
1. malloc()/free() can be either faster or slower than garbage collection. It depends on implementation details and is situational. Note in particular the issues of memory locality and temporary allocations.
2. This assumes an omniscient programmer who has perfect knowledge of when to allocate or free. In reality, memory management is a traditional cross-cutting concern, and in order to deal with it manually in a modular fashion, you will typically introduce additional overhead. Examples:
- in C++, m[a] = b, where m is a map, and a and b are strings requires that both a and b are being copied. This is not necessary in a GCed language.
- naive reference counting has significant overhead, significantly more than a modern tracing GC.
- even unique_ptr has overhead (though less so than reference counting).
Garbage collection allows you to have the necessary global knowledge to (mostly) eliminate memory management as a cross-cutting concern, because the GC is allowed to peek behind the scenes and ignore abstraction boundaries.
3. Things that C is fast at is usually also allocation-free code, i.e. iterating over arrays or matrices. The performance of such code is not going to be any different in a GCed language, assuming an optimizer of comparable caliber.
You don't need to make copy.
std::map<std::string, std::string> m;
m.emplace(std::make_pair(std::string("a"), std::string("a")));
2. The general problem that I was trying to illustrate is that shared ownership requires either reference counting or avoiding it through copying if you don't have automatic memory management.
Go is GCed language. Show me the code that will do what you claim (map[var1] = var2 and use var1(type string) and var2(type string) in later code) in Go.
let addr s = 2 * (Obj.magic s) + 1
let main () =
let table = Hashtbl.create 0 in
let key = "foo" and value = "bar" in
Hashtbl.add table key value;
Hashtbl.iter (fun key' value' ->
Printf.printf "%b %b %x %x %x %x\n"
(key == key') (value == value')
(addr key) (addr key') (addr value) (addr value')
) table
let () = main ()
This prints: true true 10cbd2490 10cbd2490 10cbd2378 10cbd2378
The exact addresses may vary from run to run.Note that == is reference equality in OCaml.
I'm honestly baffled why you think that there is a need to copy the variables or why you can't use them later or whatever you believe there.
> I'm honestly baffled why you think that there is a need to copy the variables or why you can't use them later or whatever you believe there.
What you wrote is SPECIAL case in SOME of the GC'd languages that have immutable strings. If you have language with mutable strings then this will not work. I know it will not work in Go also and they have immutable strings (so even not all gc'd languages with immutable strings optimize this) that's why I've asked for Go code in Go thread.
I don't really use Go and made a general comment about GC that was not limited to Go in a subthread about fairly general observations about GC, because somebody made a general statement about GC that was not limited to Go.
The language is not relevant for the observation I made. Nor does it really matter if we're storing strings or other heap-allocated objects. It's a question of lifetime management.
> What you wrote is SPECIAL case in SOME of the GC'd languages that have immutable strings.
OCaml strings are mutable, actually. But you seem to be confusing lifetime issues with aliasing issues, anyway.
The underlying problem is that without copying C++ would not know when to free the strings. It would be whenever they went out of scope in the calling function or when they were deleted from the map, whichever is later; the alternative is reference counting (also expensive, and not used by std::string). A GCed language can avoid that because the garbage collector will only free them when they are no longer reachable from either (or any other location).
https://news.ycombinator.com/item?id=14116435
You replied "No" to that what I wrote about Go which is not true. I never even once wrote about lifetimes. Don't you see it? Look closely again what I wrote in my comments. I think you are confused about what I was disagreeing because you didn't read carefully what I wrote. I agree with what you wrote about lifetimes 100% and I was before that discussion started because I write C++ and those things are basic knowledge. But it wasn't about lifetimes from the beginning...
> The language is not relevant for the observation I made. Nor does it really matter if we're storing strings or other heap-allocated objects.
No, because what you wrote is only true depending on the language and depending on where those strings are stored and in which language, because in some cases they will be copied so what you wrote at the beginning would not be entirely true. And I disagreed with that only, I didn't mention lifetimes even once.
If instead of what you wrote you would write (big letters to show diff):
- in C++, m[a] = b, IN SPECIAL CASE where m is a map, and a and b are strings AND A AND B OUTLIVE MAP ASSIGNMENT requires that both a and b are being copied UNLESS YOU USE SHARED OWNERSHIP. This is not necessary in SOME OF GCed languages in SPECIAL CASES.
Then I would have no reason to disagree with you in the first place.
Then you wasted a lot of time for both of us by getting sidetracked by a detail that I hadn't even mentioned in my original comment. This is about the issues with having one object being referenced from multiple locations. I gave a concrete example to illustrate this issue, and you spent an entire subthread arguing the specifics of the example without seeming to understand the issue it was meant to illustrate. This is not specific to maps or strings. It's a general issue of manual memory management whenever you're dealing with shared ownership of the same object.
> You replied "No" to that what I wrote about Go which is not true.
I think you are reading to much into an expression of disagreement. You kept fixating on the specifics of go, I was trying to get back to the semantics of GCed languages vs. manual memory management. That's what my "no" was about.
Edit: I also did a quick test for Go, just to end that part of the argument, too. Contrary to your statement, Go doesn't seem to copy strings when adding them to maps, either. While there's no way to take the address of a string literal in Go, you can add lots of instances of a very large string and check the memory usage (compared to an empty string). It turns out to be nearly the same.
> I think you are reading to much into an expression of disagreement. You kept fixating on the specifics of go, I was trying to get back to the semantics of GCed languages vs. manual memory management. That's what my "no" was about.
There was only one sentence in that comment to which you could write No, next time if you want to get back to lifetimes in discussion with someone, write "I want to discuss lifetimes" instead of writing "No" to something that someone wrote, you see a difference? Words matter.
> I also did a quick test for Go, just to end that part of the argument, too. Contrary to your statement, Go doesn't seem to copy strings when adding them to maps, either. While there's no way to take the address of a string literal in Go, you can add lots of instances of a very large string and check the memory usage (compared to an empty string). It turns out to be nearly the same.
No, you are wrong. Look at runtime hashmap implementation and generated assembly (go tool objdump binary_name > asm.S) for function that do what I was writing in my comments. You will see that there is a copy.
Because you were derailing the discussion and I was trying to get it back on track. My whole original comment was about lifetimes and GC vs. manual memory management. You were getting sidetracked by details that weren't relevant for that point, which was specifically about GCed languages vs. manual memory management in general and didn't even mention Go [1].
> No, you are wrong. Look at runtime hashmap implementation and generated assembly (go tool objdump binary_name > asm.S) for function that do what I was writing in my comments. You will see that there is a copy.
If this were true, the following program would take >65 GB of memory. In reality, it requires some 110 MB.
package main
import ("strings"; "fmt")
func main() {
m := make(map[int]string)
value := strings.Repeat(".", 65536)
for i := 0; i < 1000000; i++ {
m[i] = value
}
fmt.Println(len(m))
}
Note that even if it were as you said, nothing would prevent Go from changing to an implementation that doesn't copy strings.Frequencies have not increased much, but performance certainly has. Better single-core performance at the same frequency is the main advantage Intel CPUs currently have over AMD ones.
Also, why LLVM has spent time recently creating a fast(er) linker.
Why do we use GPUs, for example? Data parallelism. What have all the advances in video/audio codec implementation been the result of? Data parallelism again. What makes hashing fast? You guessed it...
Actually Go's shared-memory multithreading concurrency model encourages program architectures that are not easily parallelizable at all. And it's harmful to praise it. We have much better share-nothing models that encourage parallelizable programs from the start.
How was this achieved? Was it by a dev who was intimately familiar with the code base? Was it just applying the race detector repeatedly until it stopped squawking? Something else?
Both.
[0] https://github.com/golang/go/commit/b902a63ade47cf69218c9b38...
Similar words are commit, diff, patch, etc. (though sometimes those other words have other implications)
I feel like the real gains would be with a threaded linker instead, but I can't say I have any data to back that up. But I mean, you can't link until everything is compiled, so presumably there are free CPU cores that could be used.
You can also try https://github.com/derekparker/delve
But I haven't used a debugger in at least ten years. Production Go for the past few.
I debug with good log messages.
Of course, not everyone's problems are exactly like mine, so I will believe you if you say it's necessary for your situation.
I think that this point about team conversations is incredibly important. Apart from Ruby, Go is the only language that I know where the technical conversations seem to consistently focus on the problem at hand, rather than coding concerns like the finer points of language features. The company that I work for does both Ruby and C#, and the difference in the conversations around the two is very noticeable.
https://github.com/Microsoft/vscode-go/wiki/Debugging-Go-cod...
I've used it for debugging unit tests. Very handy.
Question - is it possible to turn on PolyGerrit with ability to switch back to Old UI in open source gerrit? I found the option to turn on PolyGerrit here[0]
[0] https://groups.google.com/forum/#!msg/repo-discuss/3JjvqB-mK...
[1] https://chrome.google.com/webstore/detail/gerrit-ui-switcher...
Go would already work on more than one _package_ at once, using multiple processes, which helped big builds covering many packages (think building Juju, Kubernetes, or Docker from scratch). The benefit here is in the common situation where you make a change in one package and recompile: more of your cores can be put to use getting that package rebuilt.
This is a changelist for review, after preliminary work to move state out of globals, etc. If you click "Show more" it will show much more, including discussion of the internals and wall- and CPU-time benchmark results.
It's not really significant at all. Compile-times of bigger modules will get a bit faster, such as e.g. with zapcc, which got C++ compile times 50% faster by caching.