Introduction to Event-Driven Architectures with RabbitMQ
blog.theodo.com
blog.theodo.com
Now that could have been the engineers/architects fault, but unless my team grows, I would hate having to deal with it again.
I feel like the more shortsighted/incentivized by sheer work volume a person is, the more they're into monorepos...
Can you debug a forest by inspecting a single tree?
Tracing will show you the distributed behavior, but the best log, at most, can only show you the current state of information on a single service on a single node even when you aggregate. Tracing also works hand in hand because the correlation id should associate with the relevant log message when you do decide to drill down.
Probably every technical decision maker pushing for microservices knows the perils of distributed debugging. They have some weight - just some.
That doesn't line up with my experience. I think humans have a pretty strong tendency to take for granted the advantages of a current situation when looking at something that solves a very specific new problem. "The grass is always greener"
"Hating" something of course doesn't matter, what does matter is response time in being able to debug production systems, and developer efficiency, and satisfaction / employee retention. So it's not orthogonal, imo, as it's very likely correlated with an efficiency loss.
The fact that there's great advantages to distributed architectures, yet they're harder to work with, is a signal that there's a market opportunity for better tooling around working with them.
We've got a couple of these kinds of applications where I work. All built on RabbitMq / MassTransit. Without smart logging it would be real frustrating to track down issues. You could put breakpoints everywhere locally, but that's a pain.
Tools like Seq are really helpful during development. ELK stack / splunk / log aggregator X are almost required on the non-local environments.
Every time I work somewhere I have to play shepherd and ask the very basics:
- Who's monitoring queue uptime, setting alerts on it if it goes down, waking up in the middle of the night to fix, patch it, setting it up in all test environments
- Have you thought about all the new problems that might happen: queue sending to dead endpoints, circular queue problem, queue being restarted somehow (e.g. deploys) and losing messages?
- If the app fails post-queue, not surfacing the message to the user, do you have a plan to ensure somebody in engineering sees and fixes that error? And then goes back and remediates the broken request(s)?
- Have you prepared code/logs to do distributed tracing?
- If there's a dispute a week from now whether Joe didn't get an email because of a problem BEFORE or AFTER the queue, will you be able to tell from the logs?
Many powerful engineering abstractions (threads, async, services) require one notch higher of engineering talent and allows for all sorts of new failure paths. The tradeoff must be taken very seriously. Most places I have worked at adopted complexity too soon.
> queue sending to dead endpoints
> circular queue problem
This is a subcategory of "app fails post-queue." Suppose you build a happy working queue, and someday somebody releases a new version of the queue consumer that isn't backward compatible and therefore is broken. How many messages will be lost on production before this caught? Do you have a way to recover those messages?
> circular queue problem
This is the queue version of an infinite loop. Suppose you have an infinite loop in your code, you'll crash the app but catch it very quickly.
Suppose queue message X calls a function which generates ANOTHER queue message X (infinite loop). This will be VERY HARD to catch and slow down the queue system progressively until its overwhelmed (likely only caught on prod).
Interesting point about things getting caught on prod, no amount of stress testing can sometimes reproduce these finicky bugs :)
Great explanations by the way but I try to avoid having too much logic spread out over different events.
At a previous gig, we had a processing workflow where an entity might move through any of 20 processes tied together through queues. It was very, very difficult to track down problems with things dropping out, things getting stuck, etc.
To me that's a design smell. The processes should be making requests of the entity via messaging. Each process should result in the entity emitting a change of state. That change of state should be the trigger for subsequent processes.
I'm currently working on solving a problem with an insanely over complicated setup (for the task at hand) that is build by another engineer who has since left the company.
It's a cluster of 3 virtualised machines running docker swarm where a RabbitMQ instance ties 40+ worker pods together. Once every 5 days or so the connection between RabbitMQ and (some of) the worker container stops working, causing the worker to crash and the queue message is lost.
We are talking like 5 layers of virtualization and/or abstraction. It's impossible to debug. I honestly don't know how to explain this to my customer.
We are partly in agreement regarding f/t. Logging comes into play when digging into recurring failure (bugs). Fault tolerance will not save you from bugs and bug hunting.
[p.s. to add, performance tuning/trouble-shooting also require instrumentation. Instrumentation would appear to be a fundamental (and useful) capability.]
Yes, you need to have monitoring etc, but the testing and debugging were substantially easier because of how we broke down the "services". For each major entity (or aggregate) there was a service that subscribed to a number of command and event topics, it produced output to an event topic for the aggregate.
We had FSMs for each of the aggregates, documenting the effect of each potential command or external event and the change of state and the resulting state changes (events) and/or commands.
The architectural constraint meant that the infrastructure was the same for each aggregate, the testing of each was independent and could be mocked easily using topic producers.
So as opposed to the "Big Ball Of Mud" we have a monitorable infrastructure (kafka + alerts/stats sent to a statsd integrated with AWS Cloudwatch), we have individual aggregate processing that only respond to incoming commands or events and have defined outputs for each potential incoming command/event.
Much much easier to design, develop and debug.
But the trick is at the start (like anything else). Analyzing the domain to determine the entities/aggregates, modelling the externally generated commands, modelling the FSMs for each aggregate etc.
I've always wanted to create some type of monitoring system that displays the entire system in that vein and then model or using control theory.
Has anybody seen a project that does this?