Go: Execution Tracer Overhaul
go.googlesource.com
go.googlesource.com
We started implementing distributed tracing using the precursor to OpenTelemetry with the vendor Lightstep. They have a neat solution that allows streaming OTel spans: use large ring buffers to hold and analyze as many traces as you can afford. It finds relations between spans and surfaces interesting spans. This you run in your own cluster. Then that ships a subset of total spans and traces to their saas for a good UI.
Unfortunately, at our scale, that was really expensive and we still had to resort to sampling.
Edit: It seems like
- P: Logical CPU cores
- M: OS-level threads
- G: Goroutines
A somewhat out of date explanation of Ps, Ms, and Gs by yours truly:
The problem is you won't see them connected in a single stack trace which makes looking at the flow of data in the system difficult. I mean there are profiling tools but they don't go that far.
Another way is to put print on log statements all over function on their entry and exit and try to find a coherent flow.
How do people wrriting golang everyday deal with this problem?
For your development setup, have everything run in a docker container and use breakpoints in your IDE to catch program flow. Similar to what you do with a standard/monolith program.
Is this the kind of thing you're asking about, or something different?
How do I know where the call chain goes after I have produced on a channel? There can be many subscribers on the same channel.
I guess I amasking to debug a multithreaded program to get some kind of unifrom stack trace.
Instrumentation will be richer, more accurate, low cost, and able to be of value throughout the application for the purposes of tracing and profiling.
What it will mean is that thing like https://github.com/open-telemetry/opentelemetry-go will be able to use this to better instrument Go applications, and when OTel includes profiling (as well as traces) then you'll be able to use Pyroscope, Polar Signals, etc for that in addition to Tempo (or whatever you use for tracing), as well as using something like Grafana and Datadog to view all of the above.
(I work for Grafana, but the above is relatively vendor neutral as yes we recently acquired Pyroscope and launched Grafana Cloud Profiles, and yes some of the people cited work for Grafana, but others are doing great work here too and it all benefits people who write and run applications using Go).
This design document is about improving the implementation of the Go execution tracer, but IIUC the user-level tooling and behavior will remain the same (except that the wire format will be documented).
I was surprised a bit by the discussion about using the clock instead of rdtsc equivalents. I think because I’m not familiar with the synchronisation issues between cores. I was also surprised by their timings (rdtsc should take ~20 cycles so I would have expected more like 6ns than 10, and I wasn’t expecting the clock call to be so fast). I know some processors round rdtsc values giving pretty poor precision so maybe the clock call will do better in those cases too. It seems like a reasonable approach.
Supporting tracing in the compiler (rather than as a library) seems right to me. Being able to mess with the compiler is useful because it is important to keep these traces low overhead and that generally means you would much rather write down a small constant integer than copy in some strings for each event. The compiler can ensure these integers are unique and you have some table of them somewhere. (I think it’s a little possible to do this sort of thing in C with just macros as you can drop into assembly and use more advanced assembler features like pushing/popping different sections and suchlike). And the go runtime does enough scheduling that you’ll want to trace so you’ll want to be able to get that information too.
Stepping back a bit to look at tracing more generally, the landscape of tools just feels kinda bad to me. There are some powerful things at low levels but they are very hard to use.
- Intel ipt and Arm coresight offer hardware-assisted tracing of things like jumps/function calls (from which you can mostly figure out control flow though exceptions/things like go routines can mess that up) and there is great support for both in perf but you only get a textual ‘script’ output which is obviously going to be a lot and hard to deal with. I think these tracing features can be virtualised (in newer chips) but I think hypervisors often don’t do it so eg you can’t access these features on a typical cloud vm. You also don’t have it on amd cpus.
- Linux has ebpf and such programs can be attached to various events in the kernel to trace them (I’m not sure how to efficiently get data out yet. There’s a thing called a ‘map’ which can be a ring buffer but I only know about an interface where you can read from it at one element per syscall. Someone else told me that it’s possible to just mmap it and then read it if you know the format). There is a great tool in bpftrace but it doesn’t easily give you fine-grained output, I think.
- Solaris / MacOS have dtrace, which I confess I don’t know much about.
- strace exists. It gives a lot of data but the output format isn’t exactly machine readable (I admit I haven’t tried very hard). For a busy application built in something like go or nodejs where many threads are doing many different things, the output is pretty hard to deal with. I think the ptrace(2) api it uses isn’t great for having a low overhead and I don’t know how it will trace io_uring syscalls.
- there is an Xcode tool called instruments which uses dtrace to offer a visualisation of syscalls and various call stacks but it is mostly a statistical profiler and I find the ui a bit unpleasant.
- Chrome and Firefox have their own trace recorders/profilers and viewers. You can also (somewhat) view chrome traces in perfetto and some programs will generate the chrome trace format. I haven’t used the Firefox profiler but it seems more geared towards statistical profiles than traces. There’s a cli tool called samply which outputs a format the Firefox profiler can read.
- There seem to be a bunch of different trace formats and viewers and producers and they aren’t generally very compatible with each other. I’m not sure how that will get better over time.
- There are various tools for windows which I don’t know much about (many targeted at game developers), of which many are proprietary and I think some are well loved.
I hope other language implementations will be able to copy some of the things go does as the advantage of this kind of tracing becomes more clear. It might be harder to have a good way of tracing JavaScript (which has promises instead of goroutines (in some sense these are different features but I think they are often used for similar kinds of concurrency)) but the async/await syntax and all the effort that went into improving debugger support for it suggests to me that it might be possible.