How do you cut a monolith in half?
programmingisterrible.com
programmingisterrible.com
There's a couple of nice ideas in here (the "message pump" for one at least) that I'm going to steal and use at work. It's also comforting to know that we're not completely crazy for using a DB to store tasks/processes that need to be run and retried if necessary instead of a message queue.
[1] https://en.wikipedia.org/wiki/Enterprise_Integration_Pattern...
I also like the idea, but the counterpoint is: "now you have to monitor the queue".
(I only work in small apps for my own use so its more like Postgres and forget...)
With careful management, it can be kept to a minimum, but I haven't seen that work for systems with a large numbers of developers working on it yet.
I feel like the article is attacking message brokering by discussing the disadvantages of bad use cases for them. A good use case for message brokers is when work needs to be done on an item, but not immediately.
My company uses them in a way I believe is quite effective. We pull in data from an external source, and send the ID of the item to about five different queues to do different tasks. Each time one of the queues finishes the work, it send a message to a validator that checks to see if all the work is done. If it isn't, it waits to get that message again from another worker. If it is, it marks that item as ready for the end user.
These are the kinds of use cases I think message brokers should be used for. Not to send a message and wait to get an answer back. Why not just use an HTTP request for that?
It doesn't matter, but it's more about debugging. If Service A does not work correctly, where is the bug? Is i Service A? Service B? The network?
The use at your company, is idiomatic to the paradigm. You have n different units of work, that can run seperately, so you do that and communicate with messages.
It's not that hard to do but it seems to take people by surprise and it's not helped by some poor defaults like almost everything in the RabbitMQ ecosystem silently blocking when the queue fills up (this probably happened transiently multiple times before it hit the level where it caused a visible outage but how many people will notice that if the default isn't to log or raise an error?).
I'm exaggerating a little bit, but not too much (actually, on further reflection, I might actually be downplaying it a bit. some of our services schedule tasks with like 10 different services for every item processed, and we do tens of thousands a day).
Debugging issues in this mess is not fun, because there's so many places you need to check to see if it's the source of the failure or not, and a failure in one service could really be in a different service so you have to test all the way up and down the chain. For every bug.
But then, I'm mostly against microservices because they lead to harder problems on every place. Documentation isn't even the worst.
(or, more often than not, because of the cost of doing this, you end up putting up with a slightly batshit-insane design...)
Yes, Oracle's AQ product is built on top of an RDBMS. And yes, early versions of Microsoft's MSMQ used SQL Server. But most dev teams use the database as a shortcut and don't invest the effort in making their hacktastic DB queue work like a real product.
And you use "shortcut" like some kind of negative attribute. Of course they use DB as a shortcut, and they should - it works well and even has additional features (transactions with your primary store) that you wouldn't get with separate queue product.
On the other hand, you can grow solution for those organically over time to fit your needs.
In the end, I'd say they were equally reliable, but from a data analyst's perspective I very much missed the postgres based queue. I think the mechanics inherent to a database solution prompted the original implementers to keep messages around over time vs the mechanics inherent to an MQ solution where they became ephemeral (on the senders side.) Having access to those messages was a treasure trove for troubleshooting and for analytics. That and from a pure cost benefit stance, sending 10-20k messages a day to the ESP definitely didn't necessitate the expensive rearchitecture, as the DB solution was more than capable of handling the load on cheap hardware.
* Postgresql and Oracle with skip locked
* MS SQL Server with readpast
* DB2 with skip locked data
If you use MySQL, it's going to be a bit more difficult.
This short screencast was also very helpful https://www.pgcasts.com/episodes/7/skip-locked/
First, if you're doing request/response using messaging, you're probably doing it wrong. Pub/sub and request/response are totally different animals. I for one, consider it both reasonable and necessary to use both side-by-side, in the same infrastructure. (Is this view uncommon?)
In our technology stack, which is a monolith-becoming-microservices, we use both pub/sub and request/response side-by-side. The general rule is that if service A calls service B, if the nature of that interaction is such that service B's response can preempt/interrupt/influence service A, the call needs to be done inline, in service A. If the nature of the interaction is more "advisory", use pub/sub.
Examples (from the hotel booking space): (a) When a reservation gets canceled, we publish a cancellation event. The reservation is then CANCELED, officially. A separate service sees the cancellation event and frees the associated held room inventory; that's a pub/sub interaction. (b) When a reservation wants to check in to a room, we check whether the room is already occupied. This has to be done using request/response (in our case, gRPC) because if the room is occupied, that's a hard gate on the success of the checkin.
Second, pub/sub != work queues.
Pub/sub is about distributing small bits of information all over the system and letting things be advised of stuff. Using a messaging system as a work queue is overall pretty stupid. I know it's common to use a message broker like RabbitMQ for task distribution, but it's silly. It's really silly when the tasks themselves contain huge binary objects inline, as part of the message payload. Store that shit in S3 or a proper system, and keep the message payloads light.
I guess that's all for now.
One model is synchronous, the other is asynchronous. I don't see why would anybody have any doubts.
> Second, pub/sub != work queues.
Publish/subscribe model doesn't say anything about preserving and acknowledging the messages, and with work queues usually only one worker takes the queued job. I don't see why would anybody mistake one for the other.
A pub/sub relationship should convey metadata of state change. Any state change in the system should be communicated via pub/sub.
A subscriber might need to use a request/response interaction with some other microservice in order to act on a pub/sub, but that's not a state change, that's just auxiliary data needed to do its job that is triggered on a state change.
> a message broker is a service that transforms network errors and machine failures into filled disks
...
> mark it as required in the database, and wait for something else to handle it. > > Assuming that something else isn’t a human who has been paged.
...
> Systems grow by pushing responsibilities to the edges
...
> A distributed system is something you can draw on a whiteboard pretty quickly, but it’ll take hours to explain how all the pieces interact.
I love when you can mix technicals and not taking yourself seriously.
Although be careful with using DB as a task queue. Concurrency is a b* and message brokers are very good at it. AMPQ has been created because the authors started with a BD message broker and it didn't work.
A task queue is message broker + persistance + status. Celery does that very well in the Python word, and works with rabbitmq, redis, postgres, etc.
What amaze me the most is autobahn + crossbar.io. It does PUB/SUB, RPC, load balancing and all the stuff for Python, JS, PHP, C#, Java... And it works even in the browser. Cool stuff.
I have to say. I spent way too much time trying to understand what b* trees have to do with concurrency, and what system you are using that implements them.
In a lot of cases, there are natural separation lines such as between groups of customers you can use to shard and scale things up. Unless you are building something like a social network where everything is connected to everything, you don't need a database that runs over a large cluster or clustered queues in between components. These are often just more moving parts that can break.
I suspect the advice to avoid it because of performance has become invalid for all but extreme use-cases. My company has dozens of high activity 1-100M item queues in single postgresql databases. It works great.
The way to do this, it to publish on queue _after_ transaction, but also be sure that the action won't be lost in the process.
https://martin.kleppmann.com/2015/04/23/bottled-water-real-t...
When I see such branching complexity, I often think that the architecture is somehow wrongly backwards and try a simple reversal of responsibilities. In this case, the services would be looking for a job to complete, in effect turning the afromentioned step into "job discovery" which is indeed the kind of architecture I've been applying the past decade.
Seems to be working well so far. Back pressure is handled at the front gate as none of the services pick up on the job, involuntary synchronization is still a problem, but avoidable by cleverly re-ordering the job queue. The job completion is communicated back to the front gate through pub/sub and the anecdotal evidence so far has been great.
Being able to send off a rich asynchronous message is nice because it means you do not need to have some co-owned table in a shared database that two different components are reading and writing from.
Or, worse, a widespread pattern of every service exposing a piece of its database to other services with an unsatisfying level of logging or control for what really happens.
My point is that a message-queue allows you decouple systems, whereas having two systems share the same database tables has its own kind of peril.
> In practice, a message broker is a service that transforms network errors and machine failures into filled disks.
When it's asychronous, it's easier to tolerate random errors or downtime, but the cost is that you have to store the message somewhere...
Good overview from my perspective of not knowing much about such systems.
It's always interesting to hear real reports from the trenches that aren't essentially ads for technology XYZ. I'm afraid that all too often, we make things more complicated for bad reasons; whether ignorance, chasing the latest trend, or resume-driven-development...
It's basically taking the concept of "integration" (or API) itself and creating a product for it, just as a database is a product for the concept of persistence. Thus just as not every application has to reinvent a database, with a messaging product not every product has to reinvent queueing up integration calls if the target system isn't available.
The article also mentions request-reply being "what you really want", I think that this is a) not true, as a lot of the time you can fire and forget, and b) when you need it the products generally provide a request-reply API on top of their lower-level APIs. No need to reinvent.
Both message brokers and request-response code have their places in distributed systems, but they really need to learn when each is appropriate.
I agree, saying that request-reply is "what you really want" was kind of silly, especially after the opening paragraph that states "it depends".