OpenTelemetry in 2023
bit.kevinslin.com
bit.kevinslin.com
1. It doesn't know what the hell it is. Is it a semantic standard? Is a protocol? It is a facade? It is a library? What layer of abstraction does it provide? Answer: All of the above! All the things! All the layers!
2. No one from OpenTelemetry has actually tried instrumenting a library. And if they have, they haven't the first suggestion on how instrumenters should actually use metrics, traces, and logs. Do you write to all three? To one? I asked this question two years ago, zero answers :( [1]
[1] https://github.com/open-telemetry/opentelemetry-specificatio...
You can go read the leading observability companies web pages and they'll have a 4 page writeup on custom instrumentation. That's not much, just covers very elementary basics! It's not like OTel is behind. The answer just tends heavily towards "it depends."
Once you have experience - OTel or other - you can work through these things that might confound a neophyte.
For example, you cannot define a histogram's buckets near where you define the histogram. You have to give the global exporter (or w/e the type is) a list of "overrides" that map each histogram name => their buckets. This makes it extremely ugly when you have libraries that emit metrics.
https://github.com/open-telemetry/opentelemetry-go/issues/38...
Actually it's silent pretty much all of the time. Very reminiscent of J2EE coding. Stare at the configs and hope for enlightenment.
2. I had a similar experience to you. I wanted to implement a simple heartbeat in our desktop app to get an idea of usage numbers. This is surprisingly not possible, which greatly confuses me given the name of the project. The low engagement on my question put me off and I abandoned my OpenTelemetry planning completely [1][2].
[1] https://github.com/open-telemetry/community/discussions/1598
[2] https://github.com/open-telemetry/semantic-conventions/issue...
https://clickhouse.com/blog/how-we-used-clickhouse-to-store-...
At a former place, we were doing 5% of non-error traces.
Imagine you're sampling successful traces at, say, 1%, but sending all error traces. If your error rate is low, maybe also 1%, your trace volume will be about 2% of your overall request volume.
Then you push an update that introduces a bug and now all requests fail with an error, and all those traces get sampled. Your trace volume just increased 50x, and your infrastructure may not be prepared for that.
In general you don't want your system to do more work in a bad state. In fact as the AWS well architected guide say when you're overloaded or you're in a heavy air State you should be doing as little work as possible. So that you can recover
With head sampling, the first service in the request chain can make the decision about whether to trace, which can reduce tracing overhead on services further down.
With tail-based sampling, the tracing backend can make a determination about whether to persist the trace after the trace has been collected. This has tracing overheads, but allows you to make decisions like “always keep errors”.
I don’t mean for this to sound insulting but I honestly do not think this is an acceptable take to have as a developer.
Not knowing SQL is like refusing to learn any language that has classes in it, simply because you don’t like it.
I’ve heard stories of huge corporations failing product launches because some code was written to SELECT * from a database and filtering it in-app instead of doing the queries correctly, and what’s so fun with these types of issues is that they usually don’t appear until weeks later when the table has grown to a size where it becomes a problem.
When you’re saying that you’d rather find the data in-app than in-database, you’re putting the work on an inferior party in the transaction simply because you can’t be bothered.
The code will never* find the correct data faster than the database.
* there may be exceptions, but they’re far enough between to still say “never”.
Of course not all engineers can operate with SQL as efficiently as code -- that's the whole point. Otherwise why would we be writing code? Learning SQL intimately doesn't change that fact.
I'll admit I'm a little curious about what exactly you mean here.
If you've got a successful data hungry web service with a reasonably normalized schema and moderately complex access patterns though, you're not going to be looping over the whole thing on every page load.
We’re not talking about Assembly here, “dropping down” to SQL is something that anyone should be expected to do as soon as you’re grabbing or modifying any data from a database in any scenario where performance or integrity matters. The errors you can see in situations like this are extremely complex and databases literally exist to solve them for us.
Also, if we just completely disregard the performance for a second and focus on data security instead, how do you ensure sensitive data isn’t passed to the wrong party if you don’t care about what queries are being sent?
I mean, it doesn’t matter if it’s not “in the end” displayed to an end user in the application you’re writing, or if its not stored in the intermediary node where your code is running, that data is now unnecessarily on the wire in a situation where it never should have been in the first place. If you end up mixing one customers data with another’s and sending all of it in such a way that it could even theoretically be accessed by a third party, that’s a lawsuit waiting to happen regardless of whether it was “displayed” or “forwarded” or not.
Imagine if you sniffed the packets going to some logistics app you use on your phone and you saw meta-data for all packages in your zip code in the response, or if some widget showing you your carbon footprint actually was based on a response containing the carbon footprint of every customer in the database. Even if it’s just [user_id,co2] it’s still completely unacceptable.
Never mind scenarios where you’re modifying, adding or deleting data, those are even worse and no explanation should be necessary for why.
This opens some interesting options if you want to join the result with data from your database.
To be fair, we wanted to get experience on ClickHouse and it's a special database need special attention to details on both ops and schema design.
A decent sized server will host a hugely capable instance that you may not have to think about for years. The scoffing down at DIY has made sense to some degree, but it just works brilliantly keeps getting to be a stronger & stronger case & most just assume reality can't actually work that well, that it'll be bad, and those folks won't always be right.
SaaS, especially in this space, can be *extremely* costly and its cost will scale up quickly as you send more traffic (either willingly or by mistake). Yes, Datadog, NewRelic etc will give you many pre-built and well-thought dashboards and some fancy AI-powered auto-detection thing but they will charge many $$$ for it. Consider that now cost management/analysis tools that were historically focused only on cloud, are now adding the same tooling for costly SaaS solutions!
I understand that many HN readers are skewed towards SaaS solutions, usually because they work at a SaaS shop, but depending on the size of the company, the overhead for managing it internally can totally be worth. There is overhead with SaaS as well...
w/ 1 month retention for traces:
┌─parts.table─────────────────┬──────rows─┬─disk_size──┬─engine────┬─compressed_size─┬─uncompressed_size─┬────ratio─┐
│ signoz_index_v2 │ 26902115 │ 17.06 GiB │ MergeTree │ 6.21 GiB │ 66.74 GiB │ 0.0930 │
│ durationSort │ 26901998 │ 5.44 GiB │ MergeTree │ 5.40 GiB │ 53.02 GiB │ 0.10190 │
│ trace_log │ 123185362 │ 2.64 GiB │ MergeTree │ 2.64 GiB │ 37.96 GiB │ 0.0695 │
│ trace_log_0 │ 120052084 │ 2.46 GiB │ MergeTree │ 2.45 GiB │ 37.60 GiB │ 0.06528 │
│ signoz_spans │ 26902115 │ 2.21 GiB │ MergeTree │ 2.21 GiB │ 76.73 GiB │ 0.028784 │
│ query_log │ 16384865 │ 1.91 GiB │ MergeTree │ 1.90 GiB │ 18.31 GiB │ 0.10398 │
│ part_log │ 17906105 │ 846.73 MiB │ MergeTree │ 845.39 MiB │ 3.84 GiB │ 0.21521 │
│ metric_log │ 4713151 │ 820.92 MiB │ MergeTree │ 806.13 MiB │ 14.56 GiB │ 0.05405 │
│ part_log_0 │ 15632289 │ 702.82 MiB │ MergeTree │ 701.70 MiB │ 3.34 GiB │ 0.20490 │
│ asynchronous_metric_log │ 795170674 │ 576.24 MiB │ MergeTree │ 562.50 MiB │ 11.11 GiB │ 0.049429 │
│ query_views_log │ 6597156 │ 461.35 MiB │ MergeTree │ 459.75 MiB │ 6.36 GiB │ 0.07060 │
│ logs │ 6448259 │ 408.59 MiB │ MergeTree │ 406.65 MiB │ 5.99 GiB │ 0.06627 │
│ samples_v2 │ 949110122 │ 345.01 MiB │ MergeTree │ 325.31 MiB │ 22.09 GiB │ 0.014382 │
If I was less stupid I'd get a machine with the recommended Clickhouse specs and save myself a few hours of tuning, but this works great.Downsides:
- clickhouse takes about 5 minute to start up because my tiny sc1 drive has like 4 IOPS allowed
- signoz's UI isn't amazing. It's totally functional, and they've been improving very quickly, but don't expect datadog-level polish
If anyone wants to check our project, here’s our GitHub repo - https://github.com/SigNoz/signoz
- I really like your new Logs & Traces Explorers. I spend a lot of time coming up with queries, and having a focused place for that is great. Especially since there's now a way to quickly turn my query into an alert or a dashboard item.
- You've also recently (6mo?) improved the autocomplete dramatically! This is awesome, and one of my annoyances with Datadog
Other feedback, and honestly this is all very minor. I'd be perfectly happy if nothing ever changed.
- where do I go see the metrics? There's no "Metrics" tab the way there's a "Logs" and "Traces" tab. A "Metrics Explorer" would be great.
- when I add a new plot, having to start out with a blank slate is not great. Datadog defaults to a generic system.cpu query just to fill something in, I find this helpful.
- when I have a plot in a dashboard and I see it is trending in the wrong direction, it would be nice to be able to create an alert directly from the chart rather than have to copy the query over.
- the exceptions tab is very helpful, but I've only recently discovered the LOW_CARDINAL_EXCEPTION_GROUPING flag. It'd be super nice if the variable part of exceptions was automatically detected and they were grouped
- once nice thing in DD is being able to preview a span from a log or logs from a span without opening a new page. Or previewing a span from the global page. Temporary popping this stuff up in a sidebar would be great.
- I'm not sure if there's a way to view only root spans in the trace viewer.
- This might be a problem with the spring boot instrumentation, but I can't see how to figure out what kind of span it is. Is it a `http.request`, `db.query`, etc?
> - where do I go see the metrics? There's no "Metrics" tab the way there's a "Logs" and "Traces" tab. A "Metrics Explorer" would be great.
Great, idea. This is some thing which few users have asked for and we will be shipping this in few releases
> - when I have a plot in a dashboard and I see it is trending in the wrong direction, it would be nice to be able to create an alert directly from the chart rather than have to copy the query over.
Fair point, this is something which is also in the pipeline.
> - I'm not sure if there's a way to view only root spans in the trace viewer.
We launched a tab in the new traces explorer for this, does it not serve your use case?
> - when I add a new plot, having to start out with a blank slate is not great. Datadog defaults to a generic system.cpu query just to fill something in, I find this helpful.
We can do something like this, but we don't necessarily know the name of metrics users are sending us, wrt. DataDog which has some default. metrics which their agents generate.
Will also look into other feedbacks you have given
https://github.com/open-telemetry/opentelemetry-go/blob/trac...
It's a list of the behaviors you need to implement if you're rolling your own OTEL Tracer Span implementation, and not using one of the multiple available.
In contrast, OpenTracing's interfaces had hardly any required methods, so you had to do a runtime type-cast to the whichever implementation you were using in order to access anything useful on the Span like the OperationName.
That said, if we disregard the leaky SDK APIs and half-implemented everything, it does somewhat deliver on the pluggability promise. Before OTel, you had bespoke stacks for everything. Now there is some commonality - you can plug in different logging backends to one standard SDK and expect it to more or less work. Yes, it works less well than a vertically integrated stack but this is still something. It enables competition and evolution piece by piece, without having to replace an observability stack outright (never going to be a convincing proposition).
So while the developer experience is pretty unpleasant and I am also disappointed with the actual daily usage, from an architectural perspective it opens up new opportunities that did not exist before. It is at least a partial win.
The example shows you how to use the swagger tool, parse the OpenAPI spec [2], auto-generate GoLang glue code, call __one__ of those auto-generated functions and log a trace.
However, there is zero documentation, zero other examples, and I'm left scratching my head whether there's even one person in the world using this approach. I eventually ended up just directly using the service APIs [3] via REST calls.
OTEL is painful, but the alternatives are no better :( I really wish there's some interest in this space, since SLO's and SLI measurements are becoming increasingly important.
[1] https://github.com/openzipkin/zipkin-api-example
[2] https://github.com/openzipkin/zipkin-api/blob/master/zipkin2...
[1] https://github.com/prometheus/docs/blob/main/content/docs/in...
The web browser collector published by the OTEL project uses Zone.js to hijack just about everything in the browser into contexts. If you used modern Angular before, you may recognize zone.js, its a real pain sometimes, and messes with globals, which isn't great, as it can create situations where behavior isn't predictable.
I don't know that OTEL has any standard around things like session replays either. Lots of telemetry platforms support this (Sentry, Rollbar, DataDog etc)
I think alot of backend teams have really like it. I do like the cross boundary nature of spans where you can follow them by a unique tag across your entire system.
I personally find its extremely verbose at times, in terms of the payload it generates, some logging platforms are more compact in this regard, but in practice, I haven't noticed it to be an issue
Though I agree that async context is better fit for this generally, the ERM should be good for telemetry around objects that have defined lifetime semantics, which is a step in the right direction you can use today
[0]: https://github.com/tc39/proposal-explicit-resource-managemen...
[1]: https://www.totaltypescript.com/typescript-5-2-new-keyword-u...
It is not completely solving the issue, but a starting point
Does anyone have an even part way start at doing front end tracing?
I've had better runs with Bugsnag for pure error reporting and more recently Sentry, which can do RUM / Session Replay collections.
None are what you expect though. If you want really good user behavior analytics FullStory is still top notch
Being able to connect frontend sessions -> backend trace/logs has been a huge DX change though imo.
[1] https://www.hyperdx.io/blog/browser-based-distributed-tracin...
Edit: Actually, those colleagues are doing a talk about that topic. So, if you are in Germany and Hannover area, have a look at [2] and search for "Nie wieder Log-Files!".
[0]: https://opentelemetry.io/docs/instrumentation/ruby/manual/#a...
[1]:
tracing.ts:38 Usecase: Handle Auth
tracing.ts:47 http://localhost:16686/trace/ec7ffb1e23ddbb8dd770a3f08028666b
tracing.ts:38 Adapter: Find Personal Board
tracing.ts:47 http://localhost:16686/trace/e22d342316ab0d7d23230864008e27bc
tracing.ts:38 Adapter: Find Starred Board List
tracing.ts:47 http://localhost:16686/trace/129f89cee26d54cfdc38abea368d9b4e
tracing.ts:38 Adapter: Find Personal Board List
tracing.ts:47 http://localhost:16686/trace/97948127d77501ff0c65a5db21b21b5a
[2]: https://javaforumnord.de/2023/programm/This approach isn't possible for a lot of systems that have existing logs they need to bring along, but if you're greenfield enough, I'd recommend it.
I did a talk about this at QCon London last year
https://www.infoq.com/presentations/event-tracing-monitoring...
Same is true for metrics derived from spans. (Though for metrics, you don't need to sample, and for spans you might. So keep in mind.)
It made it debuggable via output if needed, but the primary consumption became span oriented.
It's not something that anyone with a choice in the matter should be using.
I guess this is something you don't notice as merely a "user", but IMHO it's horribly overengineered for what it does and I'm absolutely not a fan.
I also disliked the Go tooling for it, which is "badly written Java in Go syntax", or something along these lines.
This was 2 years ago. Maybe it's better now, but I doubt it.
In our case it was 100% a "tick off some boxes on their strategy roadmap documents" project too and we had much much better solutions.
OTel is one of those "yeah, it works ... I guess" but also "ewwww".
I'm not an expert in every language, but I am an expert in a few, and this just isn't something that you can assume. Languages like Go deliberately do not provide the sorts of features needed to support "automatic instrumentation" in this sense. You have to fold those concerns into the design of the program itself, via modules or packages which authors explicitly opt-in to at the source level.
I completely understand the enormous value of a single, cross-language, cross-backend set of abstractions and patterns for automatic instrumentation. But (IMO and IME) current technology makes that goal mutually exclusive with performance requirements at any non-trivial scale. You have to specialize -- by language, by access pattern (metrics, logs, etc.), by concrete system (backend), and so on -- to get any kind of reasonable user experience.
That is, until some open standard is defined by said Java astronauts.
With basic parameters in place so it doesn’t eat your billing, it’s been working great with me for years. Initially with New Relic, then Datadog, now a setup with OpenTelemetry is good enough.
You should "use" metrics, logs, and traces thru dependencies that are specific to your organization. The interface between your business logic and operational telemetry should be abstract, essentially the same as a database or a remote HTTP endpoint or etc. The concrete system(s) collecting and serving telemetry data are the responsibility of your dev/ops or whatever team.
Main point: instrumentation is part of the development process, not something that's automatic or that can be bolted-on.
There is no one true interface! The interface is a function of the sub-class of telemetry data it serves, the specific properties of the service(s) it supports, the teams it's used by, the organization that maintains it, etc. etc.
OTel tries to assert a general-purpose interface. But this is exactly the issue with the project. That interface doesn't exist.
I recall Apache Skywalking being pretty good, especially for smaller/medium scale projects: https://skywalking.apache.org/
The architecture is simple, the performance is adequate, it doesn't make you spend days configuring it and it even supports various different data stores: https://skywalking.apache.org/docs/main/v9.5.0/en/setup/back...
The problems with it are that it isn't super popular (although has agents for most popular stacks), the docs could be slightly better and I recall them also working on a new UI so there is a little bit of churn: https://skywalking.apache.org/downloads/
Still better versus some of the other options when you need something that just works instead of spending a lot of time configuring something (even when that something might be superior in regards to the features): https://github.com/apache/skywalking/blob/master/docker/dock...
Sentry comes to mind (OpenTelemetry also isn't simpler due to how much it tries to do, given all the separate parts), compare its complexity to Skywalking: https://github.com/getsentry/self-hosted/blob/master/docker-...
I wish there was more self-hosted software like that out there, enough to address certain concerns in a simple way on day 1 and leave branching out to more complex options like OpenTelemetry once you have a separate team for that and the cash is rolling in.
Internally, OTEL has to keep track of every combination of labels it's seen since process start, which can easily come to dominate the processing time in an existing project. It's another in a long line of tools that dovetail with my overall software development philosophy which is that you can make pretty much any process work for 18 months before the wheels fall off.
By the time you notice OpenTelemetry is a problem, you've got 18 months of work to start trying to roll back.
I'm sure there are valid qualms with OTEL in general, but this ain't one of them. Any and all metrics telemetry systems can fall into the same design constraint you pointed out.
Push implementations do not have this problem at the client end.
Well, every unique combination of labels represents a discrete time series of telemetry data, and the total set of all time series in your entire organization always has to be of finite and reasonable cardinality. This means that label values always have to be finite e.g. enumerations, and never e.g. arbitrary values from user input.
> my overall software development philosophy which is that you can make pretty much any process work for 18 months before the wheels fall off.
The size of the set of labels in your process after (say) 1d of regular traffic should be basically the same size as after (say) 18m of regular traffic. If this isn't the case, it usually signals that you're stuffing invalid data into label values.
> OpenTelemetry is a collection of APIs, SDKs, and tools. Use it to instrument, generate, collect, and export telemetry data (metrics, logs, and traces) to help you analyze your software’s performance and behavior.
You can absolutely categorize telemetry into these high-level pillars, true. But the specifics on how that data is captured, exported, collected, queried, etc. is necessarily unique to each pillar, programming language, backend system, organization, etc.
That's because telemetry data is always larger than the original data it represents: a production request will be of some well-defined size, but the metadata about that request is potentially infinite. Consequently, the main design constraint for telemetry systems is always efficiency.
Efficiency requires specialization, which is in direct tension with features that generalize over backends and tools, e.g.
> Traces, Metrics, Logs -- Create and collect telemetry data from your services and software, then forward them to a variety of analysis tools.
and features that generalize over languages, e.g.
> Drop-In Instrumentation -- OpenTelemetry integrates with popular libraries and frameworks such as Spring, ASP.NET Core, Express, Quarkus, and more! Installation and integration can be as simple as a few lines of code.
I think OTel treats these goals -- which are very valuable to end users!! -- as inviolable core requirements, and then does whatever is necessary to implement them. But these goals are not actually valid, and so the resulting code is often inefficient, incoherent, or even unsound.
Their processors are quite capable and the entire receiver and exporter contrib collection is pretty good.
I'm not saying it's the best solution out there because that clearly depends on each use case but I don't think such harsh criticism makes sense.
Disclaimer: I'm part of the fluent-bit maintainer team.
I'm the founder of highlight.io. On the consumer side as a company, we've seen a lot of value of from OTEL; we've used it to build out language support for quite a few customers at this point, and the community is very receptive.
Here's an example of us putting up a change: https://github.com/open-telemetry/opentelemetry-js/pull/4049
Do you mind sharing why you think no-one should be using it? Some reasoning would be nice.
OTel lets the open source projects use an abstraction layer so that you can buy instead of self-host.
None of this has ever made me feel super great, but in the end I would probably consider OTel today for services that people other than my company operate. That way if some user wants to use Datadog, we're not in their way.
I used OTel in the very very early days and was rather disappointed; the Go APIs were extremely inefficient (a context.Context is needed to increment a counter? no IO in my request path please), and abstracted leakily (no way to set histogram buckets when exporting to Prometheus). I assume they probably fixed that stuff at some point, though.
More and more solutions are getting built in OTEL support, which means you can relatively seamlessly switch between backends without changing anything in your application code.
In the 2 jobs where I've set up the production environment, I just picked Prometheus/Jaeger/cloud provider log storage/Grafana on day 1 and have never been disappointed. You explode the helm chart into your cluster over the course of 30 minutes, and then move on to making something great (or spending a week debugging Webpack; can't help you with that one).
https://thenewstack.io/datadogs-65m-bill-and-why-developers-...
From an ops point-of-view devs can add whatever observability to their code and I can enforce certain filtering in a central place as well as only needing one central ingress path that applications talk to.
Also because everything emits OTLP if we ever want to move to new backends it's just a matter of changing a yaml file and not rewriting applications to support a new logging backend.
Given the choice of going back to the old way of using vendor-specific logging libraries, I will continue using OTEL 10/10 times because even given its warts, it's still a lot nicer than the alternatives.
You can't build a sound product if that's one of the design requirements.
Switching to a new backend is as simple as deploying the new backend, changing 1 line in the OTEL Collector yaml, then having your front-end pull from the new backend. 0 changes to application code necessary.
These domain concepts are descriptive, not proscriptive. They don't, and can't possibly, have specific wire-level definitions. So another way to phrase my point might be to say that OTel is asserting definitions which doesn't actually exist.
Telemetry necessarily requires specialization between producer (application/service) and consumer (observability backend) in order to deliver acceptable performance. It's core to the program as it is written: more like error handling than e.g. containerization.
On the topic of open telemetry. I have been long wanting to play with it and see if it offers all of the capabilities when we send the data to datadog. But I have been reluctant to add in another thing to manage and train on if it means that the datadog agent is still necessary for anything outside of the basics.
Has anyone else actually tried hooking this up to datadog?
Edit: just to be clear, my driving goal of this is not necessarily to keep it with datadog. But that is currently where much of our alerting and logs are now. So the idea would be to switch to open telemetry which would then allow us to (theoretically) move to something else down the line.
Nice.
FWIW I now use Tempo because I have everything else in Grafana (Prometheus, Loki), but I do miss using Jaeger.
Tempo can be spun up with docker compose using a local disk for ephemeral storage/querying: https://github.com/grafana/tempo/blob/main/example/docker-co...
Maybe this meets your needs?
> Jaeger is easier to setup/manage and has a better interface than Grafana/Tempo
What do you enjoy about the Jaeger interface? Perhaps it's a gap in Tempo we can improve.
I'm fairly sure there was an official Grafana-provided Jaeger gRPC plugin for Tempo, but can't easily find it, only this one: https://github.com/flitnetics/jaeger-tempo
Could Prometheus be augmented to store metrics, logs, and traces somehow? I don't really mind if it doesn't scale well on a single instance or is highly available, I'll just add more instances and aggregate them.
This TF project does most of the heavy lift. https://github.com/telia-oss/terraform-aws-jaeger
It's in Rust to add some HN catnip.
Cost is one thing but you would be surprised how heavy observability can be on a service, it's uses a lot of %cpu.
The former use case is often solved by a specific and narrow kind of observability data, which is metrics. A common tool for that purpose is Prometheus. You certainly can't query Prometheus for individual requests, which is fine, and accepting that invariant allows Prometheus to treat input data as statistical, in the sense that you mean in your comment.
But if we're talking about general-purpose telemetry, we're talking about more than just high-level summaries of system behavior, we also need to be able to inspect individual log events, trace spans, etc. If a user made a request an hour ago with reqid 123, I expect to be able to query my telemetry system for reqid 123 and see all of the metadata related to that request.
A telemetry system that samples prior to ingest certainly delivers value, but it can only ever solve the first use case, and never the second.
Traces came on the order of hundreds/second, but we didn't have downsampling turned on, just collected all. Traces were saved 7 days (configurable). Very little actual optimization at the point where I left.
I think it cost on the order of dozens to hundreds of dollars a month.
There's an environment variable which you can set on your containers which defines how the tracing sampler behaves. It's in the docs. See OTEL_TRACES_SAMPLER
https://opentelemetry.io/docs/specs/otel/configuration/sdk-e...
Typically seeing a 0.1 compression ratio on data before other optimizations.
I have it connected to Fly.io here: https://scratchdb.com/blog/fly-logs-to-clickhouse/
I'd be really grateful to learn more about what you're looking for (how are you even managing logs today?) Even if you end up not using scratchdb it'll help me figure out the next thing to build!
For metrics and logs, sampling isn't so useful, so I don't have a good answer. Datadog has 80% gross margin, so at most 20% of what you pay them is the infra, so you stand to save a lot of money running your own open source stacks if your labor costs would be less than that 80%. With datadog, we have a project every 3 months to reduce usage, so it's not like we aren't constantly babysitting it anyway.
FWIW we sample in datadog APM and it works fine to control costs, I'm not sure what issues you hit.
In my experience, Datadog's ingestion sampling works pretty well.
And there's retention filters you can use to override.
The high-level view and being able to draw a box around something that looks weird on my graph, and honeycomb tells me what's different inside and outside the box is amazing (called bubble up, if you're searching).
It's faster than any other tool I've used; usually data is available by the time I've switched to their ui from the curl command to our API. It's mind blowing actually.
Other saas and self hosted options I've tried have all been awful in some way or other; honeycomb is a breath of fresh air, and going back to other tools after using honeycomb is painful.
I'm not sponsored or working for them btw, they just make one of the few products that I genuinely love using.
The hypothesis being that you can save money by turning things down, but easily turn them back up when you're actively investigating. Or turn the volume up for a targeted sub-segment of your traffic.
We've done some exploration into providing the same for APM and the rest of OTEL and I think it's pretty doable. hmu if you want to talk.
For traces, as mentioned elsewhere sampling is key. Some systems use % of requests, and other systems cap traces to traces/sec.
Logs are just expensive but easy to use, so the best strategy is to lean on filtering to avoid a bunch of pointless repetition which don't add value.
TLDR, metrics require more planning but can end up being much cheaper than logs.
Thanks to this amazing group!
[1]: https://github.com/aws-observability/aws-otel-lambda/issues/...
[2]: https://github.com/open-telemetry/opentelemetry-lambda/issue...
[1]: https://docs.powertools.aws.dev/lambda/typescript/latest/
I see a lot of comments about how overly complex OTEL is. I don't disagree with this. in some sense, OTEL is very much the k8 for observability (good and bad)
The good is that it is a standard that can support every conceivable use case and has wide industry adoption The bad is that there is inherent complexity in needing to support the wide array of use cases
That said the main trouble I’ve had in the past is instability in the SDKs. Java one’s decent. Go, not so much, for example. Traditionally Otel has been awesome for tracing but not so much logs, events, and even metrics (obviously depending on your language).
Other than that, auto-instrumentation has been nice. We had this exact functionality when I worked on observability at Netflix and made starting up and maintaining microservices easy and really helped with adoption.
The lack of lock-in is fantastic, don't know if I've ever just switched technologies that easily, even SQL databases.
OpenTelemetry has been INCREDIBLY valuable to us. Not only has it made it super fast to build out SDKs for our customers, but the fact that its maintained actively gives us confidence that we're rolling out stable logic to customers' environments.
I agree with the author that OpenTelemetry has succeeded, and its pretty obvious from the fact that most major observability vendors support.
In short, to a developer it may not seem like its particularly valuable because a metrics/logs/traces API is quite simple whether or not you use OTEL. But the fact that this is an industry wide spec is where it becomes powerful.
A few
Many comments complain about the complexity of using OpenTelemetry, I recommend checking out Odigos, an open-source project which makes working with OpenTelemetry much easier: https://github.com/keyval-dev/odigos
We combine OpenTelemetry and eBPF to instantly generate distributed traces without any code changes.
I would say the hardest part was getting other devs to use it, a lot of them are stuck in their own way and did not want to go through the relatively small learning curve...
What I really like about this stack is that you can use our end to end solution, but you're not locked into it. We provide a full service, but you can also just use your own OTEL backend if you want to eject.
But if you ever decide to go this path - VictoriaMetrics supports OpenTelemetry protocol for metrics [1]
[0] https://github.com/VictoriaMetrics/VictoriaMetrics/pull/2570
[1] https://docs.victoriametrics.com/Single-server-VictoriaMetri...
But, for example, if you write a library and you want your downstream users to be able to see the telemetry. OpenTelemetry provides a standardized interface, so you don't need to make assumptions.
Consider this scenario: There is a collection of services that talk to one another, and not all use HTTP. Say agent A0 makes a connection to agent A1, this is observed by service S0 which triggers service S1 to make calls to S2 and S3, which propagate elsewhere and return answers.
If we limit the scope of this problem to services explicitly making HTTP calls to other services, we can easily use the Propagators API [1] and use X-B3 headers [2] to propagate the trace context (trace ID, span ID, parent span ID) across this graph, from the origin through to the destination and back. This allows me to query the metrics collector (Jaeger or Zipkin) using this trace ID, look at the timestamps originating at the various services and do a T_end - T_start to determine the overall time taken by one call for a round trip across all the related services.
However, this breaks when a subset of these functions cannot propagate the B3 trace IDs for various reasons (e.g., a service is watching a specific state and acts when the state changes). I've been looking into OTEL and other related non-OTEL ways to capture metrics, but it appears there's not much research into this area though it does not seem like a unique or new problem.
Has anyone here looked at this scenario, and have you had any luck with OTEL or other mechanisms to get results?
[1] https://opentelemetry.io/docs/specs/otel/context/api-propaga...
You can generate RED metrics then tail based sample and send a subset of full traces.
As with many new things, there's varying maturity in client libraries and still things missing. The Redis .NET client (maybe others?) wasn't very good (maybe alpha/beta quality) but the other .NET stuff seemed reasonable. At least with .NET, it integrates with the existing System.Diagnostics API.
I think libraries in some other languages (go?) are a bit clumsier depending on what type of diagnostic/debug/event/performance/reflection APIs that language runtime exposes for interspection.
That's not actually something you can do much about considering the sheer size of opentelemetry (both in term of implementation and vendors working on it) and i expect for people implementing nowadays, proto should be pretty stable and my experience should theorically not be the case anymore.
Let's hear some great debugging stories that have been powered by OTEL. I'd love to hear from the horses mouth without marketing speak how it was worth $$$$ collecting and storing all of this info.
From a quick glance it seems to be simple, free and open-source deployed as a single Go binary. They use ClickHouse to store data. Almost too good to be true.
I'm contemplating them for a new project.
Does anybody have any experience with the intersection of these technologies?
[0] - https://www.honeycomb.io/resources/why-we-built-our-own-dist...
[0]: https://github.com/open-telemetry/oteps/pull/171
[1]: https://github.com/open-telemetry/opentelemetry-proto/pull/3...
* For univariate time series, OTel Arrow is 2 to 2.5 better in terms of bandwidth reduction ... and the end-to-end speed is 3.1 to 11.2 times faster
* For multivariate time series, OTel Arrow is 3 to 7 times better in terms of bandwidth reduction ... Phase 2 has [not yet] been .. estimated but similar results are expected.
* For logs, OTel Arrow is 1.6 to 2 times better in terms of bandwidth reduction ... and the end-to-end speed is 2.3 to 4.86 times faster
* For traces, OTel Arrow is 1.7 to 2.8 times better in terms of bandwidth reduction ... and the end-to-end speed is 3.37 to 6.16 times faster
Pretty exciting results! The OTEL-Arrow adapter has subsequently been donated to the otel community; here's a comment that does a good job of summarizing the results and the recommendations that came out of the test [3].
[0]: https://github.com/open-telemetry/opentelemetry-proto/pull/3...
[1]: https://github.com/open-telemetry/oteps/blob/main/text/0156-...
[2]: https://github.com/open-telemetry/oteps/blob/main/text/0156-...
[3]: https://github.com/open-telemetry/community/issues/1332#issu...
The relevant code for parquet storage backend can be found here: https://github.com/grafana/tempo/tree/main/tempodb/encoding
disclosure: I work for Grafana!
I don't have the exact CPU/bandwidth numbers on me right now but CPU usage has went up by about ~50% on our "Ingester" and "Compactor" components (you can read up about the architecture here - https://grafana.com/docs/tempo/latest/operations/architectur...). But this is optimising for read performance which improved significantly.
I'm currently looking into monkey patching, but that seems dirty.
We're using middleware to start spans for http and message handlers, and then adding `startSpan` where we need.
I don't see a problem with `startSpan` everywhere as it's not much noiser than the `log.info` that would be there instead if we didn't have otel.
I'd love to accomplish that.
What you mentioned is all nice and well (optimal route), but right now, i'm working with some applications that needs it, but has i don't even know how many methods / classes that i'd need to go through and implement it on.
I wish observability actually observed more than it contributed. Once https://peps.python.org/pep-0669/ is available I'm gonna try my damndest to get otel working through it. Just give me a config file that says what functions you're interested in and I'll do the rest.
I agree 100% in javascript/typescript its annoying, and I would love to get rid of them, In go however, there isn't an extra stack frame. Nor in C# thinking about it.
The config file of functions to trace is a really interesting idea. How would you handle wanting to add data to the spans from inside those functions though? e.g. I want to add all kinds of props to the spans that can't be known except inside the function executing.
FWIW the philosophy here is that observability is a part of the application rather than something separate. That's distinctly different from the APM philosophy, which is that a separate process "does the observability" and your app is "clean" from that. I think there's quite a benefit to manually instrumenting in your codebase intentionally rather than having an automated process do it for you. But I can understand not wanting to go through and do that.
I wonder if there is any reason OTEL can't have both.
[1] https://opentelemetry.io/docs/instrumentation/js/automatic/ [2] https://github.com/open-telemetry/opentelemetry-js-contrib/t...
[1] https://opentelemetry.io/ecosystem/registry/?s=prometheus&co... [2] https://opentelemetry.io/docs/specs/otel/metrics/sdk_exporte...
Many metrics/timeseries databases make an effort to be "Prometheus-compatible" as it was kind of the unofficial standard.
OTEL: new open standards which are supposed to provide vendor-independent formats which are compatible and composable between metrics, logs, and traces, and easily enable things like deriving metrics from traces (latency would be an obvious one here).
Grafana Agent: a metrics? maybe also logs and traces? collector that supports Prometheus, OTEL and other metrics formats, which can forward, sample, and transform the data. Made by Grafana, open source, etc.
Grafana's metrics DB Mimir and maybe some others are essentially "more scalable prometheus", and use prometheus metrics format on disc, so one large concern of the Grafana Agent would be converting OTEL metrics to prometheus metrics format for ingestion into the Grafana databases - but the agent has a whole bunch of other functions supported as well.
OTEL Collector - non-vendor-specific collector, like Grafana Agent but largely just concerned with allowing collection/ingestion of OTEL and coverting other formats into OTEL. Allows extensions and plugins to be added for other purposes.
Grafana Agent is a bundling of a bunch of things for collecting instrumentation. The idea is to simplify deployment and allow opinionated setup. The Agent can upload to Grafana Cloud, or to any compatible backend (all the basic components are open source).
In particular it bundles Prometheus Agent, which does metrics collection from anything Prometheus-compatible but not queries, and OTel Collector.
It also bundles Promtail for logs.
(I work for Grafana Labs)
[1] https://github.com/VictoriaMetrics/VictoriaMetrics/pull/2570...