OpenTelemetry and vendor neutrality: how to build an observability strategy
grafana.com
grafana.com
I like how OpenTelemetry decouples the signal sink from the context, compared to other structured logging libs where you wrap your sink in layers. The main thing that I dislike is the auto-instrumentation of third-party libraries: it works great most of the time, but when it doesn't it's hard to debug. Maintainers of the different OpenTelemetry repos are fairly active and respond quickly.
It's still relatively recent, but I would recommend OpenTelemetry if you're looking for an observability framework.
For example at SigNoz [0], we support OpenTelemetry format natively from Day 1 and use ClickHouse as the datastore layer which makes it very performant for aggregation and analytics queries.
There are alternative approaches like what Loki and Tempo does with blob storage based framework.
If your data is instrumented with Otel, you can easily switch between open source projects like SigNoz or Loki/Tempo/Grafana which IMO is very powerful.
We have seen users switch within a matter of hours from another backend to SigNoz, weh they are instrumented with Otel. This makes testing and evaluating new products super efficient.
Otherwise just the effort required to switch instrumentation to another vendor would have been enough to not ever think about evaluating another product
(Full Disclosure : I am one of the maintainers at SigNoz)
https://thenewstack.io/reimagining-observability-the-case-fo...
do you mean one can write query on StarTree using PromQL?
One other nice thing about OTEL is the standardization. I can plugin 3rd party components and they'll seamlessly integrate with 1st party. For instance, you might add some load balancers and a data store. If these support OTEL, you can view latency they added in your normal traces.
It was easy enough that I, as a single SRE (at the time) could write and implement across dozens of services in a few months of part-time work while handling all my other normal duties. OpenTelemetry has proved to be worth the investment, and we have stayed within the Grafana ecosystem, now paying for Grafana Cloud (to save our time on maintaining the stack in our Kubernetes clusters).
I would absolutely recommend it. I would recommend it and hopefully use it at any new future positions.
There appeared to be wide array of library and framework support across various stacks, but I can only attest personally to the quality of the above setup (Java, Boot, etc).
> It's one of those nice things that Just Works™
did it Just Work™ or did you have to do work to make it Just Work™?
Would say .NET isn't too far behind. Especially since there are built-in observability primitives and Microsoft is big on OTel. ASP.NET Core and other first party libraries already emit OTel compliant metrics and traces out of the box. Instrumenting an application is pretty straightforward.
I have less experience with the other languages. Can say there is plenty of opportunity to contribute upstream in a meaningful way. The OpenTelemetry SIGs are very welcoming and all the meetings are open.
Full disclosure: work at Grafana Labs on OpenTelemetry instrumentation
A lot of those java built-in libraries personally, I think they make easy things easy. Will it Just Work™ Out Of The Box™?
Hey we're engineers right? We know that the right answer to every question, bar none is "it depends on what you're trying to do" right? ;)
If you have the ability to use or upgrade to the newest version of Spring Boot, you will not have to go through what I did finding the "correct way" because iirc there was a lot of shifting happening in those deps between v3.1 and 3.2.
Corporate IT were not interested in reducing vendor lock-in; in fact, they asked us to ditch the OTel Collector in favour of Dynatrace OneAgent even though we could have integrated with Dynatrace without it.
It's funny because Dynatrace fully support OpenTelemetry, even having a distribution of the OpenTelemetry Collector.
Also, to say OTel threatens these proprietary agents would be an understatement. The OTel Java agent comes with 100+ OOTB integrations right now. If I were a Dynatrace sales leader and I know that we sunk a ton of cost into creating our own stuff, I'd be casting FUD into the world over OTel too.
>We realize, however, that [vendor] neutrality can have its limits when it comes to real-world use cases.
In the short term (at least the next 18 months), data collection decisions remain important, especially for vendors like Dynatrace that provide added value beyond standard trace views, golden signals, alerting, and dashboards. Organizations need to think about these instrumentation choices on a workload-by-workload, and team-by-team basis.
Choosing between OneAgent and OpenTelemetry instrumentation is pretty straightforward. Teams use OpenTelemetry to send observability signals to more than one backend. Teams use OneAgent for its added value, such as deep code-level insights in the context of traces or performance anomalies.
Over time, these decisions will become less critical as data collection becomes more of a commodity.
Full disclosure: I am a Dynatrace Product Manager.
PS: And devs (Lightspeed?) seem to really like "Open" prefix: OpenTracing + OpenCensus = OpenTelemetry.
For logs it's going to necessarily be a bit messier IMO. Logs in OTel are designed to just be your existing application logs, but an SDK or agent can wrap those with a trace and span ID to correlate them for you. And so the benefit is you can bring your existing logs, but the downside is there's no clean framework to use for logging effectively, meaning there's a great deal of variability in the quality of those logs. It's also still rolling out across the languages, so while you might have excellent support in something like Java, the support in Node isn't as clean right now.
Metrics is pretty well-baked, but it's a different model and wire format than Prometheus or other systems.
For metrics, we're shipping a bunch of numbers over the wire, with some string tags. So why not something like:
message Measurements {
uint32 metric_id = 1;
uint64 t0_seconds = 2;
uint32 t0_nanoseconds = 3;
repeated uint64 delta_nanoseconds [packed = true] = 4;
repeated int64 values [packed = true] = 5;
}
Where delta_nanoseconds represents a series of deltas from timestamp t0 and values has the same length as delta_nanoseconds. Tags could be sent separately: message Tags {
uint32 metric_id = 1;
repeated string tags = 2;
}
That way you only have to send the tags if they change and the values are encoded efficiently. I bet you could have really nice granular monitoring e.g. sub ms precision quite cheaply this way.Obviously there are further optimizations we can make if e.g. we know the values will respond nicely to delta encoding.
Yet I applaud your desire to make things more wire-efficient.
[1] https://github.com/open-telemetry/opentelemetry-proto/tree/v...
I suspect something like clp is the way to go for logs-like data, that is, low entropy text with a lot of numerical content.
[1] https://www.uber.com/blog/reducing-logging-cost-by-two-order... [2] https://www.uber.com/blog/modernizing-logging-with-clp-ii/ [3] https://github.com/apache/datafusion
Let's say I'm measuring some quantity, maybe the execution time of an http request handler, and that handler is firing roughly every millisecond and taking a certain amount of time to complete. I'd have about a thousand measurements per second, each with their own timestamp--which to be clear can be aliased if something happens in the same nanosecond! It's totally fine to have a delta of zero. But the point is this value is scalar--it's represented by a single point.
But it seems like you're suggesting vector-valued measurements are a common thing as well--e.g. I should expect to send multiple points per measurement? I'm struggling to think of an application where I'd want this.. I guess it would be easy enough to add more columns.. e.g. values0, values1, ...
EDIT: oh, I see, I think you're saying I should locally aggregate the measurements with some aggregation function and publish the aggregated values.. Yeah that's something I'd really prefer to avoid if possible. By aggressively aggregating to a coarse timestamp we throw away lots of interesting frequency information. But either way, I don't think that really affects this format much. You could totally use it for an aggregated measurement as well. And yeah each of these Measurements objects represents a timeseries of measurements--we'd simultaneously append to a bunch of them, one for each timeseries. I probably should have called it "Timeseries" instead.
EDIT2: It might be worth spelling out a bit why this format is efficient. It has to do with the details of how Google Protocol Buffers are encoded[1]. In particular, the timestamps are very cheap so long as the delta is small, and the values can also be cheap if they're small numbers--e.g. also deltas--which for ~continuously and sufficiently slowly varying phenomena is usually the case. Moreover, packed repeated fields[2] (which is unnecessary to specify in proto3 but I included here to be explicit about what I mean) are further efficient because they omit tags and are just a length followed by a bunch of variable-width encoded values. So this is leaning on packing, varints, and delta-encoding to be as compact as possible.
[1] https://protobuf.dev/programming-guides/encoding/ [2] https://protobuf.dev/programming-guides/encoding/#packed
repeated int64 values
I’m saying that in most cases there will only be one value, hence ‘repeated’ is unnecessary.I didn’t say anything about aggregation, but yes one counts things going at a thousand per second rather than sending all the detail. The Otel signal if you want detail is traces, not metrics.
Excuse me? Modify your tone, read what I wrote again, and this time make an effort to understand it. I'd be happy to answer any questions you might have.
I'm sorry if this sounds harsh but I truly cannot tell if you're trolling or what.. I think I made a serious effort to understand what you were talking about, and it seems like you haven't done the same.
- otel collector?
- kafka (or other mq)?
- cribl?
- vector?
- other?Otel collector is very useful for gathering multiple different sources, eg I am at a big corporation and we both have department level Grafana stack (Prometheus Loki etc) and we need to also send the data to Dynatrace. With otel collector these things are a minor configuration away.
For Kafka if you mean tracing through Kafka messages previously we did it by propagating it in message headers. Done at a shared team level library the effort was minimal.
OTEL is great, I just wish the CNCF had better alternatives to Grafana labs.
Main issue is cost management
Less mature than Grafana but recently accepted by the CNCF as a sandbox project, hopefully a positive leading indicator of success.
The team around Perses did a solid job coming up with the data model and making it look very Kubernetes manifest-like. This makes for good consistency, especially when configuring via Kubernetes CRs.
Going back to the basics - 12 factor app principles must also be adhered to in scenarios where opentelemtry might not be an option for observability. e.g. Logging is not very mature in Opentlemetry for all the languages as of now. Sending logs to stdout provides a good way to allow the infrastructure to capture logs in a vendor neutral way using standard log forwarders of your choice like fluentbit and otel-collector. Refer - https://12factor.net/logs
OTLP is a great leveler in terms of choice that allows people to switch backends seamlessly and will force vendors to be nice to customers and ensure that enough value is provided for the price.
For those who are using kubernetes you should check the opentelemtry operator, which allows you to autoinstrument your applications written in Java, NodeJS, Python, PHP and Go by adding a single annotation to your manifest file. Check an example here of sutoinstrumentation -
/-> review (python)
/
frontend (go) -> shop (nodejs) -> product (java)
\
\-> price (dotnet)Check for complete code - https://github.com/openobserve/hotcommerce
p.s. An OpenObserve maintainer here.
Given that, it makes sense that if you have an OLAP data, like Apache Pinot, you can do the same and better than OLTP. You can do faster aggregations. Quicker large range or full table scans.
(Disclosure: I work at StarTree, which is powered by Apache Pinot)
Pinot is cool.
OpenObserve is similar to Pinot but built in rust with Apache arrow Datafusion as the underlying technology and for a different, targeted and tailored use case of Observability.
And that's cool! I'll look into OpenObserve.
For context: I am the Head of Product of Dash0, a recently-launched Observability product 100% based on OpenTelemetry. (And Dash0 is not even the first observability based on OpenTelemetry I work on.)
OTLP as a wire protocol goes a long way in ensuring that your telemetry can be ingested by a variety of vendors, and software like the OpenTelemetry Collector enables you to forward the same data to multiple backends at the same time.
Semantic conventions, when implemented correctly by the instrumentations, put the burden of "making telemetry look right" on the side of the vendor, and that is a fantastic development for the practice of observability.
However, of course, there is more to vendor lock-in than "can it ingest the same data". The two other biggest sources of lock in are:
1) Query languages: Vendors that use proprietary query languages lock your alerting rules and dashboards (and institutional knowledge!) behind them. There is no "official" OpenTelemetry query language, but at Dash0 we found that PromQL suffices to do all types of alerting and dashboards. (Yes, even for logs and traces!)
2) Integrations with your company processes, e.g., reporting or on-call.