Vendor lock-in in the observability space
keephq.dev
keephq.dev
There was definitely a period where Datadog seemed to be talking up OTel but not actually doing anything of note to support that ecosystem.
I'd say in the last year or two, they've done a bit of a 180 and been embracing it quite a lot.
One major change is that they not only added support for the W3C Context header format but actually set it as the default over their own proprietary format.
The reason that's a pretty big deal is that W3C Context is set as a MUST for OTel clients to implement so it goes a long way to making interoperability (and migration) pretty painless.
Prior to that, you could use OTel but the actual distributed aspect (spans between services linking up) probably wouldn't work as OTel services wouldn't recognise the Datadog header format and vice versa.
There are, of course, still some features that you would miss out on by using OTel over the Datadog SDKs like I don't believe the profiler would work necessarily but that's a tradeoff to be made.
I know that Open Telemetry has it's own issues, but between low ergonomics and amenities and a lock-in; nowadays I will choose the first.
(We've been working on some Otel-based SDKs to try to make it easier to onboard [1] as an example, though curious where the general sentiment is)
I can't speak generally, but in the Rust ecosystem the various crates don't play well together. Here's one example: <https://github.com/tokio-rs/tracing/issues/2648> There are four crates involved (tracing-attributes, tracing-opentelemetry, opentelemetry, and either of opentelemetry-{datadog,otlp}) and none of them fit properly into any of the others.
At least I don't have to use all these things. `tracing` itself is essentially mandatory if I want to pick up all the spans and events created by the crates I depend on. But I can (and maybe will...) write my own `tracing::Subscriber`, sidestepping all the various bugs and incompatibilities of `tracing-subscriber` and `opentelemetry{-,otlp,-datadog}`.
(Here's another fun one in `tracing-subscriber`: <https://github.com/tokio-rs/tracing/issues/2519#issuecomment...> These interfaces just aren't right.)
Admittedly, I think part of it is just where Rust is on the language adoption curve and where Otel is in its project lifecycle as well. Writing your own subscriber might not be a terrible idea, we internally end up needing to do similar things as we find limitations/bugs that we can't wait for upstream to fix (but we're a vendor and that makes more sense for us than end users!)
I'm considering writing a tracing subscriber that dumps events, span starts/stops, and span field updates to a terse local log file format. This is a superset of what OpenTelemetry offers. (OpenTelemetry only has the concept of a completed span, which I find really unfortunate.) So I'd write a tool that takes that and pushes it to OpenTelemetry (otelcol-contrib plugin maybe) and more local-focused tools like `logcat`.
A tangent on logcat - local observability to me is a really intriguing area, I think there's a story of Otel for local as well if someone can build a good enough local DX for consuming them (we've been told a number of times about this https://github.com/hyperdxio/hyperdx/issues/7 as an example)
* When browsing locally, I'd be able to see spans that never closed (because they were super long-running and/or because the process crashed mid-span). I suppose for the latter case, the otel collector could upload them (marking them as incomplete somehow) when it knows the process shut down.
* When looking at events(/logs), I'd be able to see the current state of all the enclosing spans, including their fields. Easiest to do locally, but ideally also in otel. Maybe some mechanism for automatically copying select attributes from the span to the event for use in otel (details tbd, whether it's selected in code at span creation time, tweaked in the otel collector config, or what).
Step 3 doesn't seem logical here... Who doesn't try to control spend on something they've already invested in?
Step 3 should be to audit what you are spending the most on, and how to manage that in a way its still useful.
I've seen so many people not understand things like custom metrics billing, or log/trace retention and get burned for it.
If you're using a tool lime datadog, you should understand its billing structure a bit. I couldn't imagine setting up a redshift instance without understanding and tracking redshift spending. And then one it is high, just switching to RDS or something without even taking a look.
DataDog have an extractive pricing model. It’s undergone a number of changes over the years that separated out features previously under the host price into their own products with separate monthly costs. They put limits on the numbers of unique containers you can have per host before charging container fees, so better hope you don’t have too many crash-restart loops or. Theres the per unique host per billing period pricing model that makes it impossible to scale your own infrastructure up and down to save on hardware costs, one personal case I can attest to was case the changes would halve the AWS bill but increase the DataDog bill by more than 10 times, a thousand percent… and you would think surely you can talk to the sales team if your use case fits extremely poorly with the pricing model…. Hahaha no… unless your a whale they give zero fucks about giving you any flexibility on price.
Even the flexible pricing they offer ends up being a sham. It just seems better on paper but you end up paying nearly the same because they have a really complicated billing model where they give you free stuff with Infra hosts, once you switch away from this model you stop getting those freebies, so your Infra hosts might cost less now but everything else is more expensive now. The house always wins!
And it's a surprisingly high bar!
It's a bullshit meaningless metric, of course, and wouldn't pass the Tufte test, but the company gets a profit center.
We (the bacalhau.org[0] project) are interested in helping with this - one of our philosophies has been that part of the problem is with that first step. By first moving to a lake of some kind, you end up giving up lots of optionality. Even basic things like aggregation, filtering, windowing, etc now need to be in the "locked-in" tool, which is exactly the wrong first step to take.
SHOW HN: We have a solution that uses DuckDB to do some of this initial work[1] which can save you 70%+ or more on total data throughput. Further, it allows you to do interesting things like eventing, multi-homing observability data, etc.
I'd be very interested to hear any/all thoughts!
[0] https://github.com/bacalhau-project/bacalhau
[1] https://blog.bacalhau.org/p/bacalhau-x-duckdb-deploying-appl...
Disclosure: I co-founded the Bacalhau project.
One of the reasons many people use SigNoz is to avoid the vendor lock-in which comes with adding proprietary SDKs of closed source products like DataDog and New Relic in their code.
If anyone is starting their observability journey today, I think OpenTelemetry is a very good place to start. You can instrument with Otel SDKs and chose a visualization layer/backend which suits your needs best
(Disclaimer: I am maintainer at SigNoz)
that's the reason terraform is not really a solution for what I speak about in the blog post.