The Modern Observability Problem
failingfast.io
failingfast.io
IMO this is doubling effort. Given all the logging,metrics and traces, can we not tell how good the product is performing? Given we know how a feature is used by users, all the clicks and APIs, can we not guaranteed future changes not breaking a user journey? I'm at the moment prototyping a tool to capture my thinking in product and software development. Having OtEL is making my idea so much easier to realise
So while there is indeed overlap, the two types of metrics involve different people, different teams, different concerns, different data, different privacy concerns, and different reliability/uptime concerns.
That's also how compiler speed up a program sometime, by inline expansion or by loop unrolling. It doesn't work all the time. But as an option.
Right now there seem to be an orthodox in product development
Engineers need to be serving the orgs goals, and one of the best ways to do that is to give them incentive and tooling to be thinking about the same things the product owner is. Have them build metrics for business metrics. Those are the things you should be alerting on, and to the extent you alert on technical pieces it should be in the service of those business metrics; don't alert on latency because it's an arbitrary goal, alert on latency because you can see how it affects user experience.
I have noticed the same thing with performance and UX. Most SaaS vendors only care about client / frontend UX and neglect the backend aspect of it. After a few years of intense development, they own good looking software that is also mindbogglingly slow and in some extreme cases unusable.
As pointed out by sister comment, developer need to know how their code is impacting their business. I would even argue developers are inherently a product person. You are always building for someone, be it another developer, internal users of you company, or general public. You have to know how people are using your tool. Hence as a developer you should want to be able to zoom out from whatever observability platform is currently offering.
Guess what, developer thinks product managers are stupid because all they know is either looking at data and measuring the wrong thing, or chasing for status update. Product managers thinks developer are nerds that only know punching on keyboard and not talking to the users. They are both right and wrong. But we create the chasm ourselves. There are alternative ways of building things.
All these tools have their limitations (and we have all of them, we use Prometheus, we have tracing, we have logs - your entire stack of everything ;) ). There is a limit to your ability to tell what's going on inside a black box based on those, sometimes they'll answer the question you're interested in, and sometimes they will not. As pointed somewhere else in the comments tracing every single interaction in your system doesn't work/scale and often the one failure you care about is not going to leave a trace. Similarly with metrics at some point just measuring everything with the right labels becomes too expensive. More than once I'm looking for a specific metric to help troubleshoot something and we don't have it (despite having a ton of metrics for everything). Alerting on metrics can be very tricky because you may not have good context, some requests might be slow because they're big, some might be fast, finding rules that tell you when the system isn't behaving is extremely difficult. Usually it's the users/customers that are going to tell you that.
Adding metrics, tracing, alerts, dashboards etc. etc. takes time/effort. This needs to be weighed against time spent on other things that can improve the quality of the product. Like design, testing, etc. Really understanding what the requirements are and how the system behaves. Just because Google or Meta set that balance somewhere doesn't mean you need to. Likely your system is significantly smaller and less complex.
This is not a new problem, logging and other methods of observability have been with us since the beginning of time, and it's always been something that needs to be approached with balance. There's some logging that adds value and there's some point where it is counter-productive. When things break, more than often the logging just gives you a starting point for debugging- not the answer.
My personal philosophy is invest in quality early on and you will reduce your operational costs. Simpler and more reliable software needs less monitoring and conversely no amount of monitoring is going to turn poor quality software into reliable software. There are many domains where the software just has to be right (lessay in your car or airplane) and you can't rely on someone monitoring the software to go and fix things if they go wrong... That said, it's always about the balance. You shouldn't care about the fashion of the day or what Google does. You need to decide where the balance is for your product that optimizes things over its lifetime with the given constraints. Every project is going to have a different balance.
Micro/3rd party services exacerbate this problem. You may see latency increase for a particular call, but what tells you why latency is increasing? What's measuring all of these tech choices, how do you know your 3rd party api is serving you traffic reliably?
Concretely: if user U1 does an HTTP GET to /spotify/track/123, that's perhaps 10KB of production traffic, but easily 10MB of telemetry traffic. In effect that telemetry can be modeled as a huge key/val map of metadata, but you can't do it that way and remain efficient, you have to optimize for observability use cases. You have to increment a cardinality-bound set of metric counters for the request outcomes, and maybe emit some best-effort trace data for the request ID, and etc. etc. — _as separate things_!
The engineering costs dominate the design. But OpenTelemetry says this isn't the case. OpenTelemetry says that whatever requirements are on FX00 company CTO feature checklists are valid a priori, and commits to doing whatever is necessary to satisfy them. That's because OpenTelemetry is evaluated not on any technical merits, but on the adoption rate of the CNCF stack among those FX00 companies.
OpenTelemetry is explicitly and exclusively a thing meant to tick off an "observability" checklist item on the checklist of a FX00 CTO's due diligence form. That's it. That's the only goal. Nobody with a choice should be using it. Read the source code, it's abysmal.
Alternatives? Write code that leverages each pillar of observability directly. There's no short-cut. That's the whole point.
If your world is only time series databases that struggle with cardinality, then sure. Fortunately, there's a lot more tools out there, some of which do just fine with high cardinality data.
I don't really agree much with your entire comment. I think you're looking at observability through the lens of 2010s-era tools, and falling deeply into the trap of thinking that this is the only way to do things.
Assuming yes, everything else I'm claiming is noncontroversial.
most of your takes here sound like they're from somewhere around 2016-2018
That is what it would be nice to have you justify some more. After all, the same could be said of logs, or internal traffic when using microservices/DBMS (request cardinality/traffic will be multiplied).
I think I am not able to follow you correctly. Is your entire point that auto-instrumentation is too much and one should default to manual instrumentation instead?
Observability is not something that a vendor can provide without meaningful and deep integration in your infrastructure. You have to do some amount of work, and IMO few vendors deliver value beyond what a single engineer can produce with with a basic internal Prometheus infrastructure + short-term log aggregation.
The whole ball game for observability systems is optimization for specific consumption use cases. That _must_ occur at the point of origin, it _cannot_ be deferred. OTel says that it's possible to define a general-purpose exporter for arbitrary telemetry data, and that specialization and optimization of that generalized data can be done later, downstream. This is simply not true.
There are tons of startups and other not-fortune-x00 orgs benefitting from what OTel provides. Your claim that OTel is irrelevant outside fortune-x00 cos is very clearly not true.
That said, OTel is far from being easy to adopt still, and despite a lot of us trying hard to change that, it's got a long way to go. If you're using one of the "major" languages (Java, .NET, JS, Python) then it's pretty easy to set up automatic instrumentation and a Collector that you can tee off to your preferred backend analysis tool. But if you need more context from your apps, manual instrumentation outside of tracing is pretty hit-or-miss, and you need to build up a vocabulary around concepts (Resources, attributes, baggage, context, spans, etc.) that isn't easy unless you've got the time to sink into it. It's extremely powerful and has the building blocks to let you capture just about any data you need, pluggable processors and exporters, etc. -- but these building blocks are very unevenly composed into easy-to-use, turnkey-ish components that a lot of people ultimately want.
If you are stuck on a polyglot loving project, than RabbitMQ or Kafka can bolt on most functions with the standard AMQP services. Erlang/Elixir is weird, but is a single kind of weird... which even has built-in profiling tools without external dependencies.
Best of luck, =)
1. Actors can define incoming message queues with unbounded capacity
2. Actors can always crash in response to any given error condition
These assumptions produce a coherent system model, which is unfortunately incompatible with any physical system that can implement it -- unbounded queues are a fiction -- and which is effectively non-deterministic and unpredictable -- crash-only software makes it impossible to assert any guarantees on a callstack.
Erlang represents a sub-optimal local maximum. It's no panacea.
I'm also curious about this:
> ...crash-only software makes it impossible to assert any guarantees on a callstack.
What sort of guarantees are we talking about here?
Erlang, and its simplified Elixir wrapper language can allow a OTP child process to continually crash while remaining operational for other users.
Some tend to get irritated by the idea of graceful fail back code. Again, something a dead-letter queue in RabbitMQ handles rather trivially.
Have a wonderful day. =)
Admittedly, it probably isn't worth the effort if you have less than 48k concurrent users. But, the distributed consensus implementation still need added to many projects (RabbitMQ essentially packages these features for you).
From my perspective, it comes down to two choices:
1. Clown show Federation and Shovel for an extra dose of chaos
2. Monoculture Clustering with the OTP, and partitioning by region
Then again, I am not smart... so YMMV =)
Erlang and the OTP are interesting and worth studying and may be a good design choice under certain constraints. But in no way does the crash-only model of the Erlang OTP represent an optimal architectural model in general. Technology has advanced since 1960.
Cheers ;)
The approach that seems most common with serious Erlang systems is to identify the common errors and handle them without crashing.
Crashing on less common, or impossible to reproduce, errors allows you to gradually improve your error handling over time if it’s worthwhile, or just allow it to continue to crash into a saner state.
See the “bohrbug” vs “heisenbug” discussion: https://ferd.ca/the-zen-of-erlang.html
- use cloud offerings because it's easy to integrate with them and they're one of the better options out there; which isn't viable in some contexts, or when you don't have money allocated for this
- setup the full Elastic Stack or Sentry, or something enterprise like that and have your stack be composed of multiple interconnected pieces of software, that need constant maintenance or even people to constantly manage them, as well as a non-insignificant amount of resources
- go for a lightweight offering, like JavaMelody for Java applications, or some of the simpler fully featured stacks, like Apache Skywalking and try to make do with their more limited feature sets and possibly more limited documentation
For me, Apache Skywalking feels "good enough", although definitely not perfect: https://skywalking.apache.org/The Docker Compose stack for it doesn't look as complicated as that of Sentry, it's basically an almost monolithic piece of software like Zabbix is and it works okay. The UI is reasonably sane to navigate and you have agents that you can connect with most popular languages out there.
That said, the UI sometimes feels a bit janky, the documentation isn't exactly ideal and the community could definitely be bigger (niche language support). Also, ElasticSearch as the data store feels too resource intensive, I wonder if I could move to MySQL/MariaDB/PostgreSQL for smaller amounts of data.
Then again, if I could make monitoring and observability someone else's problem, I'd prosper more, so it depends on your circumstances.
It’s so useful that sometimes I’m tempted to reach for it in non-ETL contexts. My problem is that these tools generally don’t mesh well with real-time streaming requirements.
It always depends on the specific use case of course, but maybe it's also worth investigating if reducing the overall complexity of the system + amount of microservices could be a solution.