Concretely: if user U1 does an HTTP GET to /spotify/track/123, that's perhaps 10KB of production traffic, but easily 10MB of telemetry traffic. In effect that telemetry can be modeled as a huge key/val map of metadata, but you can't do it that way and remain efficient, you have to optimize for observability use cases. You have to increment a cardinality-bound set of metric counters for the request outcomes, and maybe emit some best-effort trace data for the request ID, and etc. etc. — _as separate things_!
The engineering costs dominate the design. But OpenTelemetry says this isn't the case. OpenTelemetry says that whatever requirements are on FX00 company CTO feature checklists are valid a priori, and commits to doing whatever is necessary to satisfy them. That's because OpenTelemetry is evaluated not on any technical merits, but on the adoption rate of the CNCF stack among those FX00 companies.
OpenTelemetry is explicitly and exclusively a thing meant to tick off an "observability" checklist item on the checklist of a FX00 CTO's due diligence form. That's it. That's the only goal. Nobody with a choice should be using it. Read the source code, it's abysmal.
Alternatives? Write code that leverages each pillar of observability directly. There's no short-cut. That's the whole point.
If your world is only time series databases that struggle with cardinality, then sure. Fortunately, there's a lot more tools out there, some of which do just fine with high cardinality data.
I don't really agree much with your entire comment. I think you're looking at observability through the lens of 2010s-era tools, and falling deeply into the trap of thinking that this is the only way to do things.
Assuming yes, everything else I'm claiming is noncontroversial.
most of your takes here sound like they're from somewhere around 2016-2018
That is what it would be nice to have you justify some more. After all, the same could be said of logs, or internal traffic when using microservices/DBMS (request cardinality/traffic will be multiplied).
I think I am not able to follow you correctly. Is your entire point that auto-instrumentation is too much and one should default to manual instrumentation instead?
Observability is not something that a vendor can provide without meaningful and deep integration in your infrastructure. You have to do some amount of work, and IMO few vendors deliver value beyond what a single engineer can produce with with a basic internal Prometheus infrastructure + short-term log aggregation.
The whole ball game for observability systems is optimization for specific consumption use cases. That _must_ occur at the point of origin, it _cannot_ be deferred. OTel says that it's possible to define a general-purpose exporter for arbitrary telemetry data, and that specialization and optimization of that generalized data can be done later, downstream. This is simply not true.
There are tons of startups and other not-fortune-x00 orgs benefitting from what OTel provides. Your claim that OTel is irrelevant outside fortune-x00 cos is very clearly not true.
That said, OTel is far from being easy to adopt still, and despite a lot of us trying hard to change that, it's got a long way to go. If you're using one of the "major" languages (Java, .NET, JS, Python) then it's pretty easy to set up automatic instrumentation and a Collector that you can tee off to your preferred backend analysis tool. But if you need more context from your apps, manual instrumentation outside of tracing is pretty hit-or-miss, and you need to build up a vocabulary around concepts (Resources, attributes, baggage, context, spans, etc.) that isn't easy unless you've got the time to sink into it. It's extremely powerful and has the building blocks to let you capture just about any data you need, pluggable processors and exporters, etc. -- but these building blocks are very unevenly composed into easy-to-use, turnkey-ish components that a lot of people ultimately want.