The current state of OpenTelemetry
signoz.io
signoz.io
- Stability in the specification
- Stability in semantic conventions
- Stability in the protocol representation
- Stability in SDKs that can generate data
- Stability in the Collector that can receive, process, and export that data
Unfortunately, for many people, they may interpret "stable" in one of those categories as "stable for everything", and then get really annoyed when they find their language doesn't actually have stable support (or any support!) for that concept.
What I'm most proud of in 2023 is all of the little things we made progress on with components that engineers have to materially deal with. On the website, we documented what feels like a million little things and clarified tons of concepts that people told us were confusing. Across all the SDKs, we fixed tons of little bugs, added more and more instrumentations, and completed the unsexy work to make metrics generation stable across most of our 11+ languages. The Collector added oodles and oodles of support for different data sources, and OTTL went from a neat component to a rock-solid general-purpose data transformation tool.
There's so much more work to do, but I'm really happy about the progress.
Beyond that, it gives off an "over-engineered" vibe. It's probably not, and the complexity of being a unified standard that can work across so many different variations is inherently going to need a lot of abstractions, but it feels so much more difficult to go through OpenTelemetry compared to an opinionated observability SaaS.
There are too many layers of indirections across the documentation, the specification, instrumentation libraries, collectors, protocols, backends, github issues, enhancement proposals, distributions, api vs sdk. Try to file an issue and you will be bounced betweeen at least three of them.
I get it, the idea is to be vendor agnostic. But the central hub needs to be a lot more refined. Especially the language instrumentation api don’t need to pretend they are independent of the central project.
Presumably this means it's costly, which would be a reason for them to recommend it.
https://github.com/open-telemetry/opentelemetry-lambda/issue...
https://github.com/aws-observability/aws-otel-lambda/issues/...
I dislike: 1. Context/Scope: super complicated and too much abstraction/generalization here. 2. Span->activate returns a scope that needs to be deactivated manually. Which is different from ending the Span.
It gets very complicated and I assume supports folks swapping parts of sdk out within scope down the call stack. I'm curious how much that is used and if it justifies it's conceptual weight.
Again some of this stuff is lang specific problems. In java, mostly fine. In PHP (and I assume other environment per request dynamic langs) deep overkill.
It was fun having freedom to work on OTel within x-ray to provide better instrumentation users but was always frustrating how whenever pushing for more native support within the team such as ingesting otlp directly, the answer was always no since it meant losing control. Note that I don't think the actual reason is to overcharge users (otherwise why invest in things like graviton?), though the result does end up being that.
But one of our coworkers was gung ho about tracing, so I obliged not knowing what I had gotten myself into. And the moment I turned tracing on is got quenched because in a mature app the amount of tracing you want to do is about ten times the limit for messages per second per sender. It’s a toy and that functionality is now turned off, though the call graph changes to support it are still there and make our stack traces ugly.
We don’t historically do a lot of things well, but everything does correlationID propagation properly and some have pretty good telemetry, and have done since before I got here. Most other things in the “well” column were heavy lifting by myself and a handful of other instigators, some of whom gave up and left.
That’s probably what upsets me so much about OTEL. It made one of our strengths into another thing to complain about.
We moved a project from statsd to otel last year, and I really wish we had spent that time on something else. I really wish I had gotten to spend my time in something else. Statsd makes it the aggregator’s problem to deal with stats, so application lifetime (you mention lambda) is not a problem. They can fire and forget, and at a much higher data rate than we achieved with OTEL.
The main feature of OTEL is the tagging, but Amazon charges you for using it. So much so that we had to impoverish our tags until they provided only a small factor of improvement in observability that was ultimately not worth the cost of migration. If I had straight ported, we would have generated about 40x as much traffic as we got to in the end by cutting corners.
OTEL is actively hostile to programming languages with a global interpreter lock. Languages where you run processes proportional to the number of cores on the machine need to tag each process separately. If you don’t, then the stats interfere with each other. And that brings us back to pricing by tag, because now you have 16-128x times as many combinations of tags because each process has a unique tag per machine.
We ended up having to put our own aggregator sidecar in place that could merge the counters and histograms from the same machine. If you cross your eyes it just looks like one of the statsd forks that adds tags. Which would have been so much easier for us to do.
And each process remembers every stat it has ever seen, otherwise it will self clobber. So again, 16-128x as much memory for bookkeeping unless you send them to a bookkeeping process.
They had weird memory leaks that only got sorted out in the summer. There’s a workaround, but it nearly caused a production outage for us and we violated our SLAs, which costs us money.
We also had total stats loss because the JavaScript implementation is Typescript and their typescript implementation did not assert that numeric inputs were numbers instead of numeric strings. That lead to number + string arithmetic bugs, which lead to giant numbers that OTRL choked on and dropped.
It also took a lot of work to get our sidecar to consume input from 30+ processes without dropping any. Most of our boxes run about cpucount + 1-2. We don’t have an obscene amount of telemetry, but it’s a mature app with years of “we need to track X” conversations. It wasn’t until September that I was confident we could ingest from 64 cores at once, and I have no idea how we’d handle 128. Because again, OTEL does not like more than one process that thinks it’s the “same” app, so you have to tag or centralize to disambiguate.
And the thing I hate the most about OpenTelemetry: it has One Bad Apple Syndrome. Because the stats are accumulated and sent in bulk, if it does not like one value in the update, it drops the entire message. One poison pill stat causes 100% loss of telemetry from that machine. That is a stupid fucking design decision and I want the author of that particular level of hell to feel my anger about this jackass decision with every fiber of my being. Fuck you, sir. You have no business working in standards track software. Get out.
This bit is really unfortunate. Hopefully not too unhelpful, but there are several other vendors out there that don't limit you like this. Other tradeoffs to be sure (no such thing as a perfect o11y tool), but I think most dedicated observability backends have moved away from this kind of pricing structure.
I feel like you're always left to choose between obscene pricing models and AbstractSingletonProxyFactoryBeanProvider level "enterprise" configuration.
If an attribute is important for debugging something later you shouldn't have to pay 100x the cost to be able to use it. Unfortunately, when you're using a "1.0" type tool such as CloudWatch or DD Metrics you end up needing to guess the economic cost of data you include and measure it against perceived economic value later down the line, which is a terrible experience.
(Switching observability tools is no joke though, so I won't say "just switch tools!" -- if the current pains are high it may be worth it, but there's no simple way to switch that I've seen)
If one connection said it had seen 200 events, and another 190, it ping ponged between them instead of deciding 200+190 = 390. What keeps them from clobbering is distinct tags per connection. If you're running Ruby or Python or Node, that's one connection per thread, and that's ridiculously expensive.
I am glad that the observability sector has standardized on a common protocol but my god are the reference implementations lacking.
https://learn.microsoft.com/en-us/dotnet/core/diagnostics/ob...
Can you share some of your experience, what do you mean by that? Are there edge cases causing problems, or major missing features? Easy or difficult to use?
[1] https://opentelemetry.io/docs/specs/otel/metrics/data-model/...
[2] https://github.com/open-telemetry/opentelemetry-python
[3] https://github.com/open-telemetry/opentelemetry-python/issue...
I've mentally decided to just go Prometheus and ignore OpenTelemetry for the foreseeable future.
It's one of those things big players are hyping to preemptively lock you in their solution, but it's actually just alpha-quality new tech and "boring" "old" tech like Prometheus or statsd are simply more functional and better supported in the wild.
Btw, metric generation is not enabled in Tempo by default.
# tempo.yml
overrides:
defaults:
metrics_generator:
processors: [service-graphs, span-metrics]
# Prometheus
--web.enable-remote-write-receiver
# Grafana.yml
[feature_toggles]
enable = tempoSearch tempoBackendSearch traceToMetrics
[1] https://grafana.com/docs/tempo/latest/metrics-generator/
[2] https://hexdocs.pm/opentelemetry_process_propagator/Opentele...Just the general problem you get with big, slow moving OSS projects like this. Mostly just docs not current and a massive delta between certain languages; a feature is `stable` for some languages but not others which makes it hard to push for consistent otel roll out in a mixed-language environment.
Some other "misc" points:
- Google how to do $thing and you might find the proposed spec which gives example code ... that isn't what actually got implemented. That's a different link further down on your google results.
- Python auto-instrumentation is ... fragile at best. It's not super clear if instrumentation is supported only with well known frameworks or just ... in general. I'd sure love some docs that explain how it works, too.
- certain things require the collector use GRPC, others work with grpc or http... and I only found this out after googling an obscure error and reading through a _very_ long GH issue thread.
> What we ended up implementing was a little tee inside the o11y library. As well as sending events to Honeycomb, we also converted them to JSON, and wrote them to stdout. That way, after sending to stdout, we then pumped off to our standard log aggregation system. This way, we've got a fallback. If Honeycomb is not working, we can just see our logs normally. We could also send these off to S3 or some other long term storage system if we wanted.
I'd like to go a step further, and say that in addition to being worried about honeycomb being down, sometimes you just want to check with kubectl to get an idea what is going on.
Our current projects are very log light because of the heavy tracing instrumentation, but it'd be nice to integrate this with the otel paradigms as they were originally intended
So. Much.
Flames. Flames! On the sides of my face.
Breath… Heaving breaths.
One other key area is resources which can help get engineers/implementors to get organizational buy-in
I think it's a mistake for Otel to do its own thing instead of just building on top of Prometheus.
https://grafana.com/docs/grafana-cloud/send-data/otlp/send-d...
Yeah, Java is what I'm most familiar with. The "Getting Started" shows how to do some basic manual instrumentation and collect the output with curl. Then the "Next Steps" are just random things with no guidance about why I would or wouldn't choose any of them for my next step.
But, ok, I choose "Automatic Instrumentation", that sounds promising. And it actually is really easy to set up auto instrumentation. But then at the end it says
> After you have automatic instrumentation configured for your app or service, you might want to annotate selected methods or add manual instrumentation to collect custom telemetry data.
Uh... no... after I have automatic instrumentation enabled I want to do something with the output
The two major flaws in the docs seem to be
1. The common failure of docs to explain to users why they might choose one thing or another. "If you want to do x.. If you want to do y.." what if I don't know?
2. Because otel is agnostic to the consumer of the output, there's very little in the way of explaining how to get value out of what otel produces. To connect the dots, you really need to use the docs of your observability tool. Which I understand, but then most of them have their own setup directions because they want some extra fields included in the data, or they have their own fork, so not everything in the otel docs is actually usable.
I'm not sure what the answer is. It's not like I expect otel to document how to build a dashboard in Grafana. And a lot of frustration I've experienced has been with the observability tools themselves. But at the same time, I always feel like the otel docs just don't get you anywhere close to getting value out of the library. Which is a shame, because turning on auto-instrumentation and seeing all your traces with literally no extra work is a magical moment.
Observability docs in general struggle with this. So many data sources can emit so many types of metrics in so many formats, and every tool makes this impossible promise of consolidating it all into one space seamlessly. But tools like Grafana pride themselves so much on visualizing _anything_ that they paint themselves into a corner where they can't be prescriptive about common uses or methods without excluding or confusing others.
So a lot of the prescriptive answers to "what if I don't know?" gets chucked onto account and support teams of commercial vendors, because the docs can't anticipate every possible context in which an observability tool will get deployed. Each solution ends up being custom tailored and poorly portable to anyone else's, often not even to other customers with the same data sources and goals at the same scale due to wacky labelling differences or legacy requirements or some internal stakeholder demand.
More narrowly focused tools don't have as many of these problems, but not many organizations want narrowly focused observability tools. (Lots of _people_ do, but orgs don't want to pay out deals to multiple vendors for what looks like different flavors of the same result. And hey look it's Grafana Cloud or Datadog or whatever, it can do _anything_, so you devs and also bizops and SRE and IT and hey sales wants a dashboard too and so does the company cafeteria, why not, you all can just use this one tool and we just deal with one bill with a volume discount, right? Right??)
Smarter tools don't have as many of these problems by papering over the docs limitations by being better able to anticipate or surface connections between data sources, metrics, logs, traces, events, etc., and does so with better interfaces. But especially for high-cardinality data the usability of those tools either seems to fall apart or their companies charge Datadog-sized invoices.
I was shopping for one after being outside of this field for a while, and they all do the 101 features and the kitchen sink model, which adds onto the complexity. DataDog, Grafana, but also the open source ones like SigNoz itself.
Ages ago it was all about metrics, today it's metrics traces logs APM alerting exceptions and a dozen other acronyms, on top of the protocols (statsd, Prometheus, OpenTelemetry), paired with crazy complicated yet unwieldy graph building UIs. Let's not even talk about pricing models. The entire business model is based around having one more checkmark in the feature list than the competition. The wire format (OpenTelemetry) has never been the pain point in this space.
For a moment, I seriously considered just going back to the 2000s and using RRDtool.
Put differently, when you have sufficient observability of your entire system, you now have a complete abstraction of that system represented in some other UI and data streams. There's just no way out of the fact that for larger systems, this will be complicated, and the tools that can represent this reality must also be complex.
Relooking at the docs from the eyes of a newcomers if you don't already have a destination in mind they don't really help you. It's a little tricky because my setup with Grafana will be somewhat different (but similar) from someone using honeycomb or signoz or what have you, but even just having a "want to visualize your data? Check out the list of compatible vendors", with a link that direction would probably go a long way.
All I wanted to do was instrument an application and write its telemetry data to a file in a standard way, and have some story regarding combining metrics, traces, and logs as necessary. Ideally this would use minimal system resources when idle. That's it.
Here's how I run it locally for my little shovel project - https://github.com/bbkane/shovel#run-the-webapp-locally-with... .
Also linked from that README is an Ansible playbook to start OpenObserve as a systems service on a Linux VM.
Alternatively, see the shovel codebase I linked above for a "stdout" TracerProvider. You could do something like that to save to a file, and then use a tool to prettify the JSON. I have a small script to format json logs at https://github.com/bbkane/dotfiles/blob/2df9af5a9bbb40f2e101...
Amusingly I can run my application, if I generate custom formatted .json and write it to a file, I can bulk ingest it... which is pretty much what I do now without the fancy visualization app. I think this speaks to my point that the OpenTelemetry part of the pipeline wouldn't be doing much of anything in this case. (The reason I care about files is that applications run in places where internet connectivity is intermittent, so generating and exporting telemetry from an application/process needs to be independent from the task of transferring the collected data to another host.)
EDIT: it's what I've used when bridging between "this is a CLI app for maybe 3 people" and "this will need to be monitored"
It’s all moving too fast and yet not fast enough.
I get it, Google made trade offs that work for them, and I agree with their position - but for someone at a smaller company working in a non-Go/Java/C programming language it was just a ton of friction for no benefit.
In fact, the tooling in Go isn't even an example of the easiest way to do it and requires you to do more steps than for example .NET where getting server boilerplate, fully working client or just generated POCO contracts from .proto boils down to
dotnet add package Grpc.Tools
<Protobuf Include="MyApiContracts.proto" /> (in .csproj)At best it works tolerably in a monorepo with tightly controlled deployment scenarios and great tooling.
But if you don't have a Google-like operations environment, it's a lot of extra overhead for a mostly meaningless benefit.
Also depending on the environment you run in, can code size bloat vs alternatives can matter
You mean like an IETF standard? That is true, although the specification is quite simple to implement. It is certainly a de-facto standard, even if it hasn’t been standardized by the IETF or IEEE or ANSI or ECMA.
> inherently limits anything built on top of them to not be a standard either
I’m not sure that strictly follows. https://datatracker.ietf.org/doc/html/rfc9232 for example directly references the protobuf spec at https://protobuf.dev/ and includes protobufs as a valid encoding.
> depending on the environment
I’ve had several projects that ran on wimpy Cortex M0 processors and printf() has generally taken more code space in flash than NanoPB. This is generally with the same device doing both encoding and decoding.
If you’re only encoding, the amount of code required to encode a given structure into a PB is very close to trivial. If I recall it can also be done in a streaming fashion so you don’t even need a RAM buffer necessarily to handle the encoded output.
Do I love protobufs? Not really. There’s often some issue with protoc when running it in a new environment. The APIs sometimes bother me, especially the callback structure in NanoPB. But it’s been a workhorse for probably 15 years now and as a straightforward TLV encoding it works pretty darned well.
My largest complaint: observability. With almost literally any other protocol, if you can mitm on the wire, your human brain can parse it. You can just take a glance at it and see any issues. With grpc/pbuf ... nope. not happening.
Also, I really don't like how it tries to shim data into bitmasks. Going back to debugging two systems talking to each other, I'm not a computer. Needing special tooling just to figure out what two systems are saying to each other to shave a quarter of a packet is barely worth it, if at all, IMHO.
Sure, but on the other hand, the number of times I’ve needed to do this, compared to JSON/string/untyped/etc systems is precisely zero. There’s just a whole lot of failure that are just non-issues with typed setups like protobufs. Protobuf still has plenty of flaws and annoying-google-isms, but not being human readable isn’t one of them IMO.
Be careful about “need”. When people are avoiding doing something painful they invent all sorts of rationalizations to try to avoid cognitive dissonance. You don’t reach for to tool that hurts to pick up. You reach for something else, and most do it subconsciously.
Nobody is going to try to read protobuf data. Doesn’t mean they don’t need to understand why the wire protocol fucked up.
If you haven't needed to do this, perhaps you aren't working on big enough systems? I've primarily needed to do this when dealing with hundreds of millions of disparate clients, not so much on smaller systems.
I guess it depends on where you come down on Postel’s law. If you’re an adherent, and are prepared to be flexible in what you accept, then yeah, you will have extra work on your hands.
Personally, I’m not a fan of Postels law, and I’m camp “send correct data, that deserializes and upholds invariants, or get your request rejected”. I’ve played enough games with systems that weren’t strict enough about their boundaries and it’s just not worth the pain.
This requires looking at what is going over the wire in a lot of cases. It has nothing to do with Postel’s Law, but more to telling a customer what is wrong and making your own software more resilient in the face of the world.
Implement a human readable protocol, then use standardized streaming compression on the wire to get your message size down. Something LZ family because there are tools everywhere that speak them. And consider turning off transport encoding for local development.
Being able to scan data saves so much time on triage. And using zgrep and friends on production data is almost as easy. You will spend tons of effort trying to make something 10% more efficient than zlib or for certain zstd, and the cost is externalized onto your team.
Self-describing comes with rather large costs over a compact format in essentially all cases, there are lots of good reasons to prefer it. Particularly in internal infrastructure, like telemetry tends to involve.
I’ve heard the ‘binary makes debugging difficult’ excuse before and it’s just nonsense really.
Protobufs are a really great idea that's hampered by heinously subpar tooling for everything but Go.
Is observability/telemetry only about engineering-related issues (performance, downtimes, bottlenecks etc.) or does it include the "phone-home" type of telemetry (user usage statistics, user journeys)? Looking through the websites of most of the observability SaaSes it seems to only talk about the first. Then how do people solve the second? Is it with manual logging to the server from the client?
It might work now, it definitely did not as recently as about six months ago. At some point I had to get other things done this year besides OTEL.