Jaeger – A Distributed Tracing System
github.com
github.com
Disclosure: I’m the executive director of CNCF, which just adopted Jaeger 2 weeks ago, and I’m an author of the landscape.
The CNCF storage WG is also looking at creating a "zoomed in" version of the storage section with higher fidelity information. That's one model of providing more detail.
We also have an interactive version of the landscape coming that will provide filtering, zooming, etc.
An analogy is SLF4J for Java logging. All libraries, etc use the same interface and the final user determines the backend: java.util, Logback. This makes sense if you have many authors of libraries with a cross-cutting concern.
This really makes OpenTracing half a dozen different standards, one per language, with common semantics.
Should it be about a wire protocol instead? Discussion at https://github.com/opentracing/specification/issues/34
There is an open issue about Envoy/linkerd/Istio support here: https://github.com/opentracing/specification/issues/86 (as well as in a number of other locations)
As an OpenTracing contributor, the core value prop still seems quite strong in that instrumentation of OSS dependencies is a massive pain point and should not be tracing-system-specific since it doesn't need to be. There is also value in common protocols and formats, and in that spirit there is interest in broadening scope to include those... though from seeing many companies adopt tracing tech, I haven't observed protocol compatibility as the main pain point or blocker.
I have no doubt that a Go tracer would start orders of magnitudes faster than a Java one (especially if it pulls in spring or other web-related dependencies for the zipkin UI) but I think it is irrelevant.
There are a few things I wish it did, but they are all on the roadmap: http://jaeger.readthedocs.io/en/latest/roadmap/
A friend of mine, Felix Barnsteiner, wrote a profiler for Java based applications, called stagemonitor [1]. He started working on it in 2013 for his masters thesis. Since then, he steadily worked on it in the company as well as in his spare time. Some months ago, he implemented support for distributed tracing. Stagemonitor implements Open Tracing. It also collects frontend performance data, called end user monitoring. They get correlated automatically. And the best thing about it: stagemonitor is free and open source (APL). Get in touch if you have any questions.
A bit of background about how Dapper-style distributed tracing works. Things typically start with an RPC call of some kind (typically from an external source like a public load balancer). At that point, you must decide whether to trace this request or not, which is typically done as a random sample (say, 1% of requests). At that point, the request gets assigned a _trace id_, a random identifier for that request.
The trace id is stored in some request context and propagated to each subsequent service. Each service, meanwhile, divides up its request processing flow into a series of "spans" which represent some piece of computation. For example, a span cover an RPC call or a DB query. Spans are identified by a random _span id_. Once a request has been sampled, all spans for that request are sent to a central span collector where they're stored for later querying.
This model is simple but very limited. It's often hard to know whether a trace is interesting at the outset, hence the reliance on random sampling. For example, you might want to understand why your 99p latency is high, but if you're just sampling 1% the 99p requests will only be 0.01% of your sample.
More generally, interesting events (like errors or slow requests) tend to be rare, and sampling a random, small percent of requests is unlikely to turn up the interesting cases.
A better model, as implemented by lightstep [1] (and an in-progress distributed tracer I've been working on) is to collect all spans. Even with very high request volume it's reasonable to store all spans for at least a few minutes. Doing so opens up all sorts of interesting possibilities, because you can start tracing a request at any point during that window. For example, you might want to trace all requests that have errors in them. Or all requests that take longer than a certain time. Or get a google sample of requests across different latency buckets. Or requests that violate some application invariants you've defined.
Ultimately, though, distributed tracing is so helpful for understanding complex distributed systems and webs of microservices, and it's exciting to see more open-source competition for zipkin.
[0] Dapper is Google's distributed tracing system. The paper (https://research.google.com/pubs/pub36356.html) kicked off a lot of interest in distributed tracing in the broader community. [1] http://lightstep.com/
That being the case, it's not hard to see why people are going with the (existing) OSS solutions. :/
UPDATE - Found this on GitHub, is this the whole thing?
https://github.com/lightstep/lightstep-tracer-go
If so, pointing people towards it from the .com website might help get people trying it out, as the .com website makes it seem non-OSS. :)
I hope to be releasing an open-source version of that approach in the next several months.
LightStep specifically is meant for large-scale enterprise deployments and their specific needs and has focused on that for now.
Having said that, Dapper-style tracing systems with head-based sampling still provide enormous benefits that are often underestimated. In fact, even if we support tail-based sampling in the future, we're still going to run a certain portion of the traffic through probabilistic, head-based sampling, because it allows to reason about statistical patterns observed in the systems at large scale.
https://github.com/jaegertracing/jaeger/issues/425
The Dapper model is interesting, but Jaeger is not 100% based on that.
From the documentation it doesn't look like much, except probably the biggest downside is that you have to add your instrumentation points manually? i.e. there is code change required.
https://github.com/cncf/toc/blob/master/proposals/jaeger.ado...
Instead of just method calls in a process, you take a top-level (HTTP) request and see all the various upstream service calls and their internal logic involved in completing that request. Useful for micro/multi-service based architectures but you can trace anything because it's a standard output format.
For those who don't know german, Jaeger is "Jäger" is hunter/ranger. A somewhat neutral term in itself but there is also a slight, somewhat remote connection towards some part of the history ("Jagdstaffel" and what not). I have absolutely no idea if this has anything to do with it, mind you - but since the main authors appear to be in the USA, I find that very awkward. Why not just stick to some english name? That would seem a much better it. Or perhaps they think german names are awesome ... it's also weird when you see all the people write Jaeger rather Jäger...
There are plenty of software projects named from a variety of languages not native to the creator of the package. And Jäger has so many neutral meanings and even for a German wouldn't first bring thoughts of the Jagdstaffel I don't think...
Also writing the ä as ae is very common for those who have keyboards without umlaut characters... I've seen it lots of times and doesn't look too weird.
Edit: typo
We may not know our history, but we do like drinking.
FWIW, it seems to be an appropriated term.
You might be reading a bit more into it than is warranted.