Domain-Oriented Observability
martinfowler.com
martinfowler.com
But no such luck.
However if anyone is creating a distributed product today in 2019, here is an excellent suggestion. Free of charge.
In whatever common code you have for your RPCs, add the ability for an RPC message to be traced. If it is traced, then you will log when it is received, when you reply to it, and also cause all RPC messages you send while handling it to also be traced.
At your front end you can then pick an arbitrary fraction of your traffic, say 1%, and trace it.
Tracing such a small fraction of traffic means that you can instrument it in real time, gather the logs, and create a complete view of how each web request cascaded through your system. And then whenever you run into a slow request that only happens for a small fraction of your traffic you can find a request that was both instrumented and slow, and then dig into those rare problems that happen here and there, then add up.
If you don't do this, then good luck tracking down the performance problems that happen 5% of the time, 3 RPC calls from the front.
I was researching it recently for work but haven’t had a chance to implement anything with it, yet.
Every cloud vendor has a product that does it (AWS X-Ray, Google Stackdriver Trace, etc), and there are plenty of open source products (Jaeger, OpenTracing), as well as third party vendors (we're using Datadog ourselves).
Having said that, I completely agree with this being very important, and it is the backbone of our ability to debug issues. I would add that it is important to have the ability to correlate things as well, specifically your logs and traces ("give me all the logs that relate to this trace X").
Because I realized API calls (which are just RPCs) need this (which is a defacto feature of programming language through exceptions and stacktraces), I now think it is better to start at the programming language level, write your function normally, and then compile it to call/respond as an API endpoint.
Actually I think it's best to write a monolith and then cut it s different pieces into independent microservices using code generation/manipulation, even though one's first intention is to build an ecosystem of microservices.
However I somewhat disagree with the article in the sense that instrumenting reporting logic is just table stakes bare minimum of software in a business. 90% of business software is about creating a report for somebody somewhere, often needing to be rapidly changed for ad hoc requests from product or management people. Reporting code is the code. It may even be the business logic in a truer sense than the application logic itself.
Another lesson along these lines applies to machine learning and modeling systems: the mathematical algorithm is always the easiest part. The hard part is situation-by-situation customized data ingestion (things that cannot be solved by standardizing on platform-specific data formatting, like Hadoop or Spark), and then also defining business metrics that capture the success of the model at a high level (as opposed to engineering metrics like accuracy, precision and recall, etc).
Model training and the 10% of the project spent tweaking an algorithm or optimizing things just pales in comparison to the type of systems you need to facilitate efficient custom data ingestion that can be arbitrarily different on a per project basis.
There's also a lot of mixed concerns here around in place logging and analytics collection for something that is a very business-logical focussed class.
If you're wanting to do things like this, I'd encourage you to have a look into a message driven Domain driven design structure. It's approach would be:
- remove all logging and analytics concerns from the class
- throw errors based on business rule violations, or alternatively publish a failure type message if it's logic path that had compensating actions
- publish all mutations of the shopping cart as immutable events
- subscribe to events related to analytics/logic and do that work in a handler, or do it in a higher level application service so the shopping cart just contains shopping cart concerns
A lot of this can be distilled down to basic DRY and SRP principals
The problem is in the engineering culture, not the particular tool used.
(shameless plug i recently wrote about how decorators can be used to keep the domain metric/plumbing free)
https://medium.com/dm03514-tech-blog/designpatterns-consider...
(also just disclaimer: Team's i've been on have had lots of success using the approach outlined in the article, injecting a metrics/instrumentation object, it cleans up the domain logic, is easy to provide a stub implementation during tests, and abstracts the from metric implementations.
------
A decorator version might compose the production code by wrapping the domain logic in metric adapters:
metrics = InstantiateMetricsClient()
ObservableShoppingCart(
ShoppingCart(
ObservableDiscountService(
DiscountService(
),
metrics=metrics,
)
)
metrics=metrics,
)
----Another option that I've seen be used successfully is to completely decouple domain events and the surfacing of domain events by having the domain generically emit events, and the logging/metrics subscribe and surface those however it decides.
----
I feel like for normal logging statements and metrics it only borders on being an issue, but when metrics include many timing wrappers, or when tracing (which involves much more reporting) is involved explicit adapters help to really keep things clean.
Pete Hodgson : Pete Hodgson is an independent software delivery consultant based in the San Francisco Bay Area. [...]
Zipkin looks like a library for collecting and inspecting distributed traces. This post is about how to keep your instrumentation code clean. Unless I missed something while I was reading it, it doesn't say anything about the underlying infrastructure.