Grafana Tempo, a scalable distributed tracing system
grafana.com
grafana.com
It feels like Grafana Labs is stepping into the same "if it doesn't help sales tick checkboxes it's not getting development time" trap so many B2B companies hit as they grow and it's a shame. I wish there was something else I could switch to.
To illustrate, here's some of the paper cuts that instantly come to mind: mod-s not working in panel editor, gradient fill incompatible with series overrides, various missing tooltips, missing query error reporting and syntax highlighting in variable editor, variable editor needing many clicks for most common actions, hidden variables can only be changed by repeatedly opening and closing the settings, some dashboard settings pages being easy to accidentally close without saving, g+[key] keyboard shortcuts randomly not working, random scroll jumps with multiple queries in a panel, switching between cursor modes with mod+o is super slow, Prometheus autocomplete not working when editing existing queries, clicking the dashboard name filters by folder while clicking the folder name shows all dashboards, adhoc filters silently don't work, etc.
The "not responding" one specifically happens a lot while editing queries. It'll randomly decide to evaluate the query while I'm typing, get thousands of results and lock up.
Open since 2014, hitting it on every other panel...
I can't help but feel they've learned the wrong lessons from their challenges with tracing.
I'm telling you, from first hand experience, this does not end well.
There's no reason that your tracing system should not be indexing your tags in an engine that provides advanced search features through a powerful query language.
Tempo's implementation seems pragmatic as a short to medium term solution though. Log engines still have a lot more investment and maturity than trace engines. In my work, even though the trace tags contain the best data quality, the tracing system is currently worse at answering a good deal of my questions. It's simply that Splunk has many tools that work well, and the tracing system is behind.
What features are you looking for and how do you rank them?
It's really nice to have some metrics when for instance a service goes down. It's super easy to spot a OOM situation or other vertical scaling issues.
I'm now wondering where the line is between necessary monitoring and bikeshedding? I could look at introducing distributed tracing, but will it actually add any meaningful value?
Prometheus is by far the cheapest, both for infrastructure spend and for human cost. The only time we spend really working on Prometheus is configuring our service discovery, which is the shipping component you’d have with any monitoring tool. I estimate we spend about 1 developer day a month un Prometheus upkeep.
Logging is much more fiddly, and required a large up front investment. Now it’s running though, we get a lot of value from our setup. It costs much more though, think about $10k/month in infrastructure spend and about 3 developer days a month to maintain.
Grafana is effortless to run, really dead simple. You can consider this to be ‘free’, both for infra and maintenance.
Hope that gives you a sense of this? I think we come out equivalent to managed services for cost when you account for infra and human time, but have far more flexibility in how we use the tools and develop the skills to properly leverage each product along the way.
My initial draft of this stated that we are using Prometheus/Grafana + ELK across the company, although its still being standardised and turned from individuals creating their own deployments to a proper managed, strategic system with SLAs/documentation.
We're looking at the feature lists of SaaS offerings like Datadog, Logz.io, but I think we're too big to get a sensible price. That said, I don't think we're big enough to justify something like Thanos that turns Prometheus from a really simple application, into a big time investment.
Out of interest, are you storing 2TB of Prometheus data in a single instance? Or spread across multiple?
The 2TB of data is what we have stored across all those Prometheus instances. We use thanos as the entry point that Grafana speaks to, so you get aggregated results.
Thanos as a querier is very simple to setup, and is very low maintenance. We have intended to configure long term storage using the GCS backend for years now, but sadly this project always ended up losing to other (genuinely!) more impactful work.
We hope to do this within the next 6 months though, and reckon the project will take about 2 weeks of our teams time.
For the monitoring angle then, I can recommend Prometheus and thanos as very easy systems to configure. Even for a small team with no prior experience, you'll probably have a good time.
The one to watch out for is Elasticsearch, as that is a fundamentally more complex system in which you plan to store much more data. Loki looks much easier to setup and benefits from the Grafana ecosystem integration, if you're looking for a shorter/cheaper/less featureful option.
You comment on standardising the logging schema is a great idea as well. I'll circulate that around.
Thanks for your help!
I think ELK is the most challenging bit as you outlined yourself but is still very doable.
My company uses ELK, and I personally really dislike it. It's okay when it works — filtering logs with queries is the primary use case, and it's decent at it when it's not returning 503s. But it's also based on Elasticsearch, inheriting all its warts. I really wish someone would build a better competitor. Loki/Grafana isn't anywhere close yet.
One of my pet peeves with ELK is how the indexer assumes all log entries have the same schema. So if one app has the "error" field as a string, and other logs it as an object, then ELK will reject the second one. It boggles my mind that someone thought this was a good technical design. It could easily suffix internal field names with its schema (type, analyzer, etc.) and then unify the fields at the UI level. But no, instead it discards log data.
Sluggish performance (ELK requires enormous machine researches for no particular reason, though the JVM is a major driver) and lack of support for tailing are two other pet peeves. (I have many more.)
That can be fixed by:
- logging to different indexes
- preprocessing your logs so the keys have their own schema prefix, as you mention
The way we've tackled this is to have an official company-wide logging schema. It's just a GitHub repo at gocardless/logging that has an exhaustive list of logging keys, with an explanation of what they should contain.
This has the benefit of encouraging consistent logging practices over many teams, as well as improving the chance your logs will get indexed correctly. If the field type doesn't match, we won't index that field but it will appear in the _log field, where you can do a full text search as a fallback if you really needed to find your log.
It's not perfect though, and I still hate the dynamic type assignment.
- Service discovery can get tricky. In our case we manage the physical machines by ourselves so at the beginning we just had to write down static configs, although by now we have automatic discovery (took a few days of development).
- Depending on what you want to monitor, you might need to write your own exporters to Prometheus. However, mtail [1] has been really useful to create metrics from logs without too much work. In any case, you'll have to put time in deploying and configuring those exporters.
- Dashboards and alerts. There are dashboards for a lot of exporters, and there are collections of alerts too [2], but you will need to put time and effort in modifying/creating dashboards and writing down alerts. However, it's a productive effort because it helps in having a better understanding of which metrics are important and how do they relate to the workings of the software you use. Also, PromQL is a pretty nice query language for the purposes of Prometheus.
- Notification integrations. In my case we had to put some time to properly configure a Microsoft Teams integration and a deadman switch channel, but in most cases it will be pretty straightforward.
All in all, you'll need to invest some time in the integrations, but those are things that you need to do in any case. Prometheus itself is pretty easy to set up and maintain, and doesn't stand in your way. No tweaks, no undocumented settings, no bugs. I'm pretty happy in that regard, once you get it running you don't have to worry about it. Storage usage is pretty low even with a high amount of exporters and metrics per node, maybe around 10GB for 60 days of data of a single node? (I'm not sure because Prometheus does some compression and it's not exactly linear with the time or the number of nodes).
And that relatively low investment pays off quickly. The machines we manage use various tools and programs to deal with quite a lot of data at high bandwidths, so performance problems and bugs can be difficult to debug. The Prometheus + Grafana setup has made several times easier the debugging of issues and performance problems, the alerting system helps us prevent outages and we have even discovered issues that were unknown to us. For me, the moment you manage machines with even just a little bit of complexity in the software or setup, it's already worth it to look into monitoring.
1: https://github.com/google/mtail 2: https://awesome-prometheus-alerts.grep.to/
We spit out some 30k+ spans per second, FWIW. :-)
Edit: Disclaimer, we're not using Tempo.
I think I must be missing something, but it seems like the big difference between Tempo and a traditional tracing system is the storage indexing & database (ES/C* vs object store and index all fields vs key/value lookup by ID). I vaguely remember reading something that latency even from EC2 -> S3 can be around 200-300ms. Wouldn't this cause the overhead to rise?
Feel free to point me to any documentation that clear this up!
Disclaimer: I am using Tempo :) (and from Grafana)
I was testing Google Pub/Sub's Go client for publishing internal API event data for later ingest to BigQuery, and it turns out Pub/Sub publishing is not that much faster than writing directly to BigQuery. The buffer sizes we'd need to avoid adding latency to our APIs would have to be ridiculously high; the Pub/Sub client buffers and submits batches in the background (its default buffer size is 100MB!). I don't like the idea of having huge buffers that increase with the request rate.
Conversely, pushing the data to NATS in recent time without any buffering or batching turned out to be fast enough to not add any latency. You have to be able to receive messages very fast on the consumer side (as NATS will start dropping messages if consumers can't keep up), but you can simply run a few big horizontally autoscaled ingest processes that can sit there ingesting as fast as they can, which never impacts API latency at all.
if you need to see the full request as it passed through your system then distributed tracing/Tempo is a great fit!
https://www.honeycomb.io/blog/secondary-storage-to-just-stor...
Backing Jaeger with Elasticsearch or Cassandra was a nightmare. :-|
It will be easier and cheaper to manage compared to Jaeger but doesn't yet have any ability to search inbuilt (which is one of the reasons Jaeger is expensive).
Not to bash Jaeger though, its more powerful than tempo in that it allows you to search for traces. Tempo is about integrating with Grafana, Loki and Prometheus for finding traces.
Unless you're already running one of these, in which case deploying Jaeger is very easy (at least, that was my experience with our Elasticsearch backend).
> much more performant.
You expect an object store to be more performant than Elasticsearch?
Jaeger supports native search, but requires Elastic or Cassandra.
Tempo relies on discovery from logs/exemplars, but puts everything in object storage (s3/gcs).
Tempo is cheaper and easier to operate but lacks native search.