What I wish someone would have told me about using RabbitMQ (2020)
ryanrodemoyer.github.io
ryanrodemoyer.github.io
> I have personally experienced network partions happening in two ways: all nodes in the cluster being updated at the same time through Windows update and firewall rules.
I'd rather try to build a four story brick building with horse dung for masonry mortar, than run my backend on windows boxes. How many cumulative hours of horror has Windows Update unleashed on the human race?
+1
War story: an ex-client of mine was on a different continent demoing a product (involving a physical device talking to a software backend hosted by a partner company) at a neutral site to a potential new customer.
Months and months of talks and prep, dry run demos in the weeks before, everyone was quietly confident. Scope for selling lots and lots of these devices if the demo goes well, but the marketspace is fairly ruthless.
Minutes before the team are due to present, a sudden and total comms loss between the device and the backend. Cue frantic calls which due to the time zones even involved getting the CEO of the partner company out of bed. Minutes ticking by..
Potential customer not impressed, said: you have ten minutes, after that we're leaving.
As the potential customer was walking out of the building, the partner company finally tracked down the issue.
Windows Update, on a backend database server.
There will be situations where you can't, special cases where you're left with few options. If the system is that critical have 2 of everything (clusters, databases, etc.) in a way where the mirror can stay untouched and keep your business going.
And if it needs saying, something like updates or reboots should never happen truly automatically (as in the software decided when or if to do it). A human should decide how to automate that and have full control over when and if it happens.
Probably nobody consulted a sysadmin when they designed the system. *grin*
Windows 10 behaviour is another story, but I would always stand on what it's a self-inflicted wound - people blamed MS for not updating soon enough... while disabling or postponing updates for literally years. MS acted. Poorly, yes, because by that time it ascended to General Motors levels of bureaucracy[0], but again - this is another story.
[0] https://www.joelonsoftware.com/2006/06/16/my-first-billg-rev...
Sure, Windows update takes a while, but it's still so much better than whatever the hell Apple is doing. Despite the fact that I abhor removing the possibility to choose what your device does when, I think forced automatic updates are net win for computer users worldwide.
In fact, I suspect Microsoft may need to be even more forceful, because even though consumer Windows rebooted itself at the time, wannacry and friends still managed to spread across the globe using an exploit that had been patched in an important security updates months before.
There are so many other, better reasons to say it besides this or automatic updates, but Fuck Microsoft. Right in the ear.
If you're using the OS version which is designed to by used on Grandma's desktop as a server then you should expect a world of pain.
Me? I try to avoid using Windows hosts because .Net runs great under Linux and you don't need to have 5 different Windows VMs running on your dev machine to have an accurate development environment.
Mass-reboots or network outages happen for all sorts of reasons.
Power outages happen. Network switches fail. Firewalls get misconfigured. Scripts get mis-scheduled.
Windows Update is just one of many ways I've seen all-at-once reboots of servers.
I don't see how this would be better on Linux. Do Linux admins never make mistakes in cron jobs? Is Linux magically immune to switch hardware failures?
If RabbitMQ couldn't tolerate reboots, then that's on RabbitMQ, not Windows.
PS: SQL Server AlwaysOn Availability Groups will happily restart after an all-servers simultaneous reboot. ZooKeeper won't. By "design", apparently. It converts a temporary outage into a permanent one... on purpose.
When was the last time a switch just restarted to apply updates? i do not dislike Windows Update but I want to control it and apply things that are convenient to me.
When it's updated by a network team. It's a semi-regular occurrence in many workplaces.
Windows Update is not forced upon Server editions. It can be scheduled at will by administrators.
This complaint is the equivalent of saying that Linux is terrible because an admin scheduled "sudo apt-get update" to run on every node of a cluster at the same instant.
Why would admins running updates concurrently be a failure of Windows, but not a failure of Linux?
I seriously don't get this complaint, ESPECIALLY because the article makes it clear that:
1. The fault was in RabbitMQ, because I quote: "The default partition handling strategy is ignore which means to just enter the partitioned state and keep trucking along in this “split brain” mode thereby thrusting your cluster in to total chaos." [1]
2. Windows update wasn't inherently at fault: "The fix for Windows update was the ensure that nodes in the cluster are patched at different times."
I bet the default partition configuration setting of RabbitMQ is also wrong when running on Linux, and the behaviour would be identically bad if an admin scheduled updates to run concurrently on all cluster nodes.
[1] That's clown-shoes programming, which is why I'm never using RabbitMQ for anything ever. I've only heard bad things, like "the default settings result in corruption and data loss."
You can never, ever rely on people being infinitely knowledgeable and vigilant. Any attempt to do so is simply showing ignorance and indifference by yourself.
An unmarked pit in the middle of the road is a hazard. Sure, drivers in theory ought to be vigilant of such things, but fundamentally the fault is with the construction crew for not leaving out bright orange markers.
The default should be safe. The default should not be to optimised to win benchmarks to benefit the few, but to provide guarantees to provide for the many.
But the few have power, and power corrupts.
Hell, half the time that person is me.
There's a lot of valid criticism for Windows Update but I don't think this is one of them (anymore?). Part of being a system administrator dealing with Windows Server is managing updates through WSUS or something else. If that's your responsibility and Windows Servers unexpectedly "just restart" then you are not qualified for that job.
Maybe switching to Linux servers solely because of how updates are applied is reasonable. Personal preference is absolutely reason enough. But I think reaching the same level of customization+automation for updates there requires similar amounts of knowledge and time spent.
Having a server reboot on you unexpectedly is because either (1) you turned on automatic updates yourself or (2) you're running a server on a Windows client SKU, e.g.: Windows 10.
(TBH I am not sure if the author is using the right tool for the job. But everyone is an arm-chair expert and I only know MQTT from my smart home; if I reboot that daemon then a temporary loss of state simply doesn't matter)
Patching and rebooting all your nodes in a cluster at once is not highly available.
But hey, let’s focus on windows hate instead.
I think GP acutely points out that one of the problems the author has with RabbitMQ is actually a problem with (a misconfiguration of) their underlying OS, MS Windows.
If you've been bitten by this, I understand the hate completely.
Nope. Just ubuntu deploying docker update that had problem restarting the docker service again. Simple restart fixed all of them.
But I guess people do have some bias, since on PC, windows updates are really predatory.
We found out the hard way that RMQ does not behave like a transactional DB. Just because publishing worked does not mean the message will be delivered.
Our solution is to also write the message into an outbox table in the DB. We then publish the message using confirms[0]. RMQ asynchronously sends us a confirmation when it has really persisted the message. We then delete the outbox entry. If we do not receive the confirmation in time, a timer will re-publish the message.
Therefore I disagree with the suggestion of using a library wrapping the native RMQ one. We are using spring-amqp and this made it harder to understand what is going on. In the end, for a large project you will have to understand nuances of RMQ (and other infrastructure you are using). Using a leaky abstraction over it means you now have to understand both the underlying product and the abstraction.
[0] https://www.rabbitmq.com/confirms.html#publisher-confirms
We hoped it would be fast enough that we can just wait for the confirmation before committing the transaction.
The official documentation says
> This means that under a constant load, latency for basic.ack can reach a few hundred milliseconds
I never did statistics, just looked at the log. IIRC most were acceptable but > 3s occurred frequently enough (and we even had instances of messages never being confirmed, IIRC) that we abandoned that plan.
We considered using Debezium[0], but decided on the current solution as it could be solved entirely with the current services and infrastructure whereas Debezium would have required us to deploy (writing this from memory so this might be inaccurate/incomplete) Kafka, Zookeeper, and a connector service.
E.g. the ability to watch a ZK node for changes, which means in Kafka sans ZK, you can't detect changes to topics without continuously polling via the admin client.
A coworker is working to implement something like this for KRaft, but it really demonstrates how an IPO can cause a company that was the steward of a FOSS project to do things detrimental to that project to keep the share price up. (Was also interesting how many key Confluent people left right after the IPO)
The other very notable change is how Confluent's dev effort has switched from the open source project to the Enterprise Edition, but they still have the majority of PMC members, while not having the corporate blessing to spend time reviewing PRs.
Yikes, that sounds like an oversight! Aren't topic configs written to a system topic that you could consume from?
So my coworker's solution joins the KRaft quorum as an observer, then publishes metadata changes to a topic you can consume from.
KIP = Kafka Improvement Proposal. The mechanism for proposing big changes to Kafka and getting community feedback before core committers vote on adoption or not.
I'm not sure why Zookeeper is viewed so negatively in regards to Kafka, it's damn solid, and if I can quote Jepsen "Use Zookeeper". I know Confluent wanted to replace it because it struggled when you hit thousands of topics on a single cluster, but that feels like a very niche use case to me.
I agree re: zookeeper. It’s rarely been the part of the stack making me lose sleep. ~~Raft~~ seems like a way to “modernize” Kafka by detaching it from the Hadoop ecosystem — and I think that’s about it. :(
It's one of Kafka's strengths.
1) Adopt RabbitMQ without any experts on the team 2) Conceal Rabbit/AMQP functionality as much as possible behind a simplifying abstraction, often in multiple layers, often written by non-experts 3) Run into some intractable reliability or scaling problem 4) Have no idea how to solve it because you still don't have any experts 5) Throw a lot of money at the problem, fail 6) Decide to do a very expensive migration to a different system (SNS+SQS, Kafka, etc.)
At that point, you go back to step 1. If you're lucky, somebody has expertise in the new system and the migration can be pulled off successfully. Otherwise, you either end up repeating the whole process or everything goes off the rails when you're halfway migrated to the new system.
This same process happens for all kinds of stuff, not just RabbitMQ, of course.
For tracking state at scale (but still per-job, per-thing) a Cassandra-like system works best (but preferably a better implementation, eg. SkyllaDB or AeroSpike or some other KV store).
Where it bites people is that it’s not a queue and scaling is harder than just add more consumers.
This is one of the reasons I am a really, really big fan of Google Cloud's Task Queues. It allows the stupidest, simplest temporal execution of HTTP invocations.
Currently working on a project in AWS and it's stunning how complicated it is to achieve the same simple need of "I want to execute this HTTP call at this time in the future". It's either AmazonMQ -- using either ActiveMQ or RabbitMQ with plugins -- or hacking around SQS's 15 minute delay limit. In our case, we are going to end up wrapping our messages in an envelope with a delivery time and if it hasn't met the delivery time, we put it back into SQS.
GCP is highly underrated for how it simplifies control over execution of code. Pub/Sub and Task Queues both have HTTP delivery built in. Couple that with Google Cloud Run and it is a recipe for building almost any type of execution model with much less complexity and overhead
message bus: firehose of events, (by default) no ACK. usually multi producer multi consumer. (see also DBus which is more of an RPC layer + service discovery + pub/sub via event listeners)
message queue: usually between components, ACK, but no selective ACK, backpressure, might even have support for "dead letters" (letters not ACKed by any consumer)
job queue: selective ACK, retry, etc.
(there's also the "enterprise service bus", which is similar, but mostly implemented on things like IBM MQ)
In my case, a whole team of devs was using RMQ without knowing anything about it. Literally caused many sev-1 occurrences over the years until I resolved all of the issues. It took a datacenter migration that allowed me the opportunity to redesign the entire RMQ infrastructure before I was able to put the whole mess to bed.
Having this new failure mode added to all the other ones I’ve already met over the last few decades has colored my perception a bit, and I’m having opinions about how you shouldn’t try to wrap a wrapper, and maybe the best way to live with a bad API is to pass through the yucky bit as quickly as possible - preprocess to see if you can avoid calling it at all, and then avoid asking it to do anything extra the rest of the time.
That part doesn’t feel that transformative to me but maybe I’m wrong. What’s bigger and stickier for me is that I now have to think about some NIH code we wrote that deeply bothers me, and decide if I still don’t like it, or if the author had the same conclusion and this was their answer.
I ended up burnt out. Too much to learn, and I'm not really interested in most of the infra tooling we were using (I loved redis and got to know tons of it, on the other hand I never really cared about k8s or rabbitmq). Sadly, more often than not, companies out there still expect the mentality of "you build it you run it" which implies acquiring knowledge of all the little things in your infra that can go wrong in production. Damn it.
If I'd were in your shoes I would have ended up burned out too if I were given the responsibility of having to understand it all, but no mandate to simplify it as much as possible.
GraphQL is another toxic technology that relieves the responsibility of designing the API to users instead of the application owners. May as well just give folks access to your DB and have them run SQL queries themselves. Look! No need for a UI.
All this at the expense of the software engineer that has to manage all this complexity themselves.
I think it’s easy to say we as engineers should make things simpler. I have no idea how to do that in practice on a team though.
Once the complexity started to creep in, seems like a wave just kept sweeping everyone towards more and more complexity.
Having the default behavior during a network partition be "whatever, just chug along as if nothing is wrong" is bonkers. Yes, the person who first set up the cluster should have read the documentation and gone over the configuration file line by line to see what might need changing, but... damn, that's a terrible default. Sure, some people's applications might value availability over consistency, but that's not the safest choice that follows the principle of least surprise.
Using a higher-level library to interact with the cluster is really good advice, in general. We used Kafka at my last company, and colleagues who actually knew what they were doing wrote a (simple) wrapper library that set things up properly for our cluster so clueless people (such as myself) could write producers and consumers without having to understand what all the fiddly connection setup settings did, and how to handle various edge-case errors. Before that, quite a few outages were due to producer/consumer misconfigurations.
Also I can't imagine running this kind of infra on Windows servers. That sounds like a self-inflicted wound (by "self" I mean the company, not OP specifically). And the idea of Windows Update running on a prod server ruining your day... what? IMO infra should be as immutable as possible. Patching/updating software on a machine should be a matter of spinning up a new machine with an already-updated image (that you've built and tested elsewhere), bringing those new machines into a cluster, and then decommissioning the old ones. When colos and dedicated servers were all the rage, that was difficult (and sometimes impractical), but in this day and age, with on-demand cloud provisioning, there's no excuse for companies that can use that sort of infra.
While this is absolutely a "best practice" I really think there is a pretty big divide here between the best practice and what happens In The Real World (strictly by number of developers). Between $DAYJOB, consulting/freelancing, and the internet/tech twitter/HN/etc, I've seen basically three groups:
1. FAANG, startup, etc. that probably follows this best practice "cattle over pets" methodology more often than not. 90% of developers do not work here. For the same reason you can't compare FAANG salaries to really anything else, you can't compare their tooling and infrastructure either. They're playing a completely different game than most people.
2. Indie developers who want to mess around with this stuff.
3. Boring companies with IT departments where you have to request VM resources and wait for them to get spun up. I worked for a place within the last decade where the idea of getting a VM provisioned for you in less than 2 business days was laughable. This was a publicly traded, multi billion dollar company with probably 800 developers, IT, and support desk folks on staff and did not make any money from its software directly (we were not a profit center).
I am considerate of the "well you should just" mindset when it comes to dev ops and infrastructure, but most developers don't work in that environment. If the company has never had that sort of expertise or expense before, it's a hard sell especially if it's coming from a developer who has never worked in that environment before. And truthfully, the last thing I want to do as a software engineer is become "the devops guy" and now be responsible for the VM infrastructure of the entire team rather than actually writing code.
I didn't it before 3.8. Read the docs. All of them--three times.
So, don't do that. I don't use rabbit mq currently but we do use Elasticsearch, which has similar clustering capability and used to be more susceptible to split brain situations (been there, dealt with that)
These days, I recommend using Elastic Cloud and avoid self hosting it. It's only cheaper until the first time you have to deal with a split brain cluster because you botched an update, mis-configured it, etc. One of the nice features in Elastic cloud is that you can click an update button and it will orchestrate a rolling restart. If you don't know what that is, you should not be operating a cluster of any kind.
I'm sure there are similar cloud based services for rabbitmq. Probably well worth the money unless your in house ops team is super experienced with operating it (which clearly wasn't the case here). Such a team would cost you many hundreds of thousands of dollars per year. One person does not cut it. You need at least a 3 or 4 so you can afford them taking vacations, sick leave, leaving, or dying in some tragic accident. Half a million pays for some pretty nice cloud based clustering capacity. A good team will cost you more.
For reference, we pay about 70/month for a tiny Elastic cloud search cluster that is actually good enough. I can double the price and capacity with a simple slider and it would still be cheap. One hour of my time is more than that with my normal freelance rate. My monthly rate would pay for an enormous cluster that far exceeds anything we need and it would still be cheaper than making sure we have four people with my skill set in the team (we don't) able and willing to look after it at all hours. The largest cluster I've ever dealt with was a self hosted monster that could index billions of documents per hour (millions per second). You so much as looked wrong at it, all hell would break loose. That cluster was one of several managed by a very experienced ops team that probably cost millions per years. That's the price of doing business at scale when self hosting.
Most companies running into trouble with ES cut corners on doing it right and then pay the price with preventable outages, scaling issues, technical debt, etc. Cloud based services don't completely prevent this but if you know what you are doing, they provide a nice level of safety and risk mitigation.
We run a cluster for ingesting logs, and it's much cheaper. The clustering part is basically painless. I had to script a bunch of things to easily change instance sizes (adjust the memory according to available RAM), and that's it. Sure, if you don't want to have to deal with this, going the managed route is OK.
The most issues we've had, however, is with the actual documents sent for indexing, querying them, etc. So basically something with which ES Cloud, or other managed solution, wouldn't have helped.
I'd expect RMQ to be roughly the same. I've seen a bunch of people do weird things shoveling messages from left to right and not consuming them before the RAM exploded. All this on managed RMQ.
And I say this as somebody who really really dislikes managed services from a "but I want to have total control" point of view (this is not entirely rational on my part but true of me nonetheless.)
The other point is the implicated idea that, somewhat rightfully, operating the software without expert knowledge is "amateur hour" but apparently developing for the same software without knowing what you are doing is somehow ok? To me, that doesn't make sense. The same logic applies there, where you need a handful of people to counter sick leave, accidents and other unplanned human outages.
This is especially true with Elastic. Sure, running a production cluster on some developer laptop which suddenly reboots for updates isn't great, but neither are broken schemas, uncontrolled bucket growth, broken stemming, or any other of a million things that can go haywire logically without operations being an issue.
There's simply no getting away from requiring expertise. If it's your part of your day job, you need to really know the stuff. Yet not knowing what you're doing is so taken for granted that "throw it in the cloud so at least the operations people know what they're doing" is completely accepted logic.
That doesn't mean cloud services are bad per se, but the expertise-not-needed argument is. This in turn changes the economics considerably. Sorry for going on a tangent. I know I'm the odd one here but I can't bring myself to accept it.
I've seen what proper ops teams look like and if you can afford one they are a great asset. 24x7 uptime with five nines is a bit of a myth in our industry. But if you are honest about what that means, it means that you need people with a clue available at any point that you can rely on to be there and to be skilled ready to fix your system when it needs fixing. That's effectively what you pay for with managed services. You'll get a nice resilient setup, backups, monitoring, and support from their expert ops team to help you with whatever.
Most small companies are just faking being very dependable and end up spending more on preventable outages and figuring out all sorts of weird shit than they save by self-hosting. I see startups wasting a lot of time on all sorts of silly things that are clearly far out of their comfort zone. Hosted infrastructure is cheap and dependable these days. Unless you can compete with that level of dependability, you should not be self hosting anything. The cost savings rarely add up to being meaningful. The expenses on the other hand are very substantial if you stop and think about it and acknowledge that the per hour cost of your developers quickly adds up to whatever you might actually spend on managed services.
Most teams I've worked with spend more on a single developer per month than on their cloud hosting. That's how it should be. Once your bills hit tens of thousands per month, you might want to look at optimizing some of that. But before that, have your developers focus on developing rather than trying to be an amateur ops team.
I actually do a lot of consulting related to Elasticsearch and I have dealt with some pretty hairy and misguided setups. By the time they talk to me, they've already wasted a lot of time, money and effort. Most teams I deal with have one or two people that know only a little bit about Elasticsearch. Enough to use it and benefit from it but then they still need me to tell them how to use it properly. So, clearly not enough to operate it responsibly. Even if I come in for just a few days, they'll spend more on me than years of hosting with Elastic Cloud would cost them.
I like how on the surface they can be quite simple to use, but the optimisation and management of them can be fiendishly subtle.
Some of my most interesting projects have been trying to squeeze more messages through a pipe with lower latency, changing the way messages are sent and flow through these platforms, or digging into why one in a billion messages are dropped. There was also an interesting phase of trying to containerise and infra-as-code Kafka.
It’s all like plumbing for data infrastructure. An interesting corner of the IT industry.
- NATS Core[0] as an ephemeral message exchange (personally what I would use RabbitMQ for) - NATS Jeststream[1] as a persistent, queue-focused kafka alternative - Liftbridge as an alternate implementation of a persistent kafka alternative[2]
Liftbridge has a decent comparison page[3] but unfortunately it's still missing NATS JetStream[4].
I want to see more members of the community use and write about the NATS ecosystem -- I rarely hear complaints.
[0]: https://docs.nats.io/nats-concepts/core-nats
[1]: https://docs.nats.io/nats-concepts/jetstream
NATS is a lively project, has never had anything to do with mules or camels, and isn't hard-tied to concepts like ESB. It deserves to be mentioned in 2022.
But here are some projects I left out that do deserve to be mentioned I think:
- RedPanda[0]
- Apache Pulsar[1]
With the message bus the clients can (roughly) send requests as frequently as they want, while servers will handle them as fast as they can, without the danger of missing a request or having clients to wait. The queue is the "asynchronization" mechanism.
It's easier to handle certain situations, like if you have multiple recipients, you may want to notify all of them, at least one, at most one, etc.
It's also a way of handling consumers that may be unavailable for whatever reason and your source may not want to have to deal with that. For example, if the producer is a cron job, you may not want it to have to hang around for an arbitrary time if your consumers are already busy. Sure, this means that the queue is available, but in principle it's supposed to be more available than the consumers.
If your needs are fairly basic, that's probably the best approach. You also eliminate some overhead (piping messages through a middle man, operating said middle man, etc.).
But it's my understanding that if your needs are a bit more involved, the complexities of building a solid queue mean that it'll take time away from building your actual product. In that case, you may be better served by an off-the-shelf solution which already handles the corner cases and is good to go.
I think this is the usual decision between using an external product and rolling your own.
With a queue you don’t need to know, and the recipients can subscribe to information on the queue themselves.
Especially useful in event-driven architectures.
Push vs pull.
There’s also a difference in making a command to another service vs event or state notification. In the case of a command you often want to wait and find out if it went ok. Events or states published onto a queue are more of a case of “here is a thing that happened, deal with it how you wish”
It's not a bad idea at all, you probably use message based systems every day without thinking about it, culturally however the idea is irresistible for enterprise developers who live for abstractions. Just something to keep in mind.
Secondly, it's decoupled from your application. Trying to do things directly works for a while, then you do an in memory array, then your application crashes or gets updated and lost those messages.
2) Reliability. Imagine, if a customer places an order in a store, but the server process crashes and the order is lost. The customer won't be happy about this.
A mesasge queue might help in both cases. Note that you might implement a queue in a SQL database as well, but this still will be a message queue.
First, we use them to allow app users to run (large) data processing tasks without keeping connections open for a long time. Basically letting our users start a data processing job on the client side and then let them continue doing something else, and not forced to stay until in the same page until the data processing job is completed. We check for status updates in the database; once our data processing job is complete, it updates the status of the job.
Second, in the backend, we have multiple tasks (compute resources) running that are ready to take on these data processing requests based on messages in the queue. To allow for bursts of user requests for data processing, we make available multiple tasks whose job it is to run the data processing jobs. We place data processing request messages in the queue, which are then picked up by any of the available (non busy) data processing tasks. With a local/internal queue you suggested, we’d have a harder time figuring out and ensuring the right tasks picks up the data processing job, and that only one tasks does so to avoid duplicate effort. Or we might have to have local queues on each task, and somehow coordinate them. With the centralized queue, we don’t have to worry about this - whichever tasks is available picks up the message and runs the data processing job. When the tasks takes the message it pops it from the message queue and no other task picks up the same message (unless we want it to for some other reasons).
I'd probably have started out with -just- the outbox though and bolted the HTTP request sending onto the database, then looked at switching to a dedicated queueing system later. But that's a "because I already understand the care and feeding of PostgreSQL" choice as much as anything else.
then you notice that you can do much more than that, but most people start here.
The webpage I wrote twenty years ago would accept your POST request, wrap up everything into the email, send it, return a redirect to a GET to the same page, and finally send you a page saying "Thanks! I've sent you an email with a link to click on to confirm it's really you." You would not get this web page because your browser would have timed out, or maybe you would but you'd think "Jeez this site is so slow" not realising that it's just waiting for you.
The webpage I wrote ten years ago would accept your POST request, wrap up everything into an email, send it to a message queue, and so on except maybe now I'd say "I am sending you" rather than "I have sent you", which you'd see immediately. Around the same time, a queue worker notices a new email in the queue, sends it, hangs around for a minute or so waiting for your server. You get the email, click the link, wow what a nice fast responsive site!
The webpage I wrote last year would do much the same thing, but the queue worker might flip a flag in your account to say the email has been sent but the link has not been clicked. Meanwhile a wee bit of javascript is polling an endpoint watching for the flag, and flips the text round to say "Check your email, it's been sent".
Ugh. I hope OP is making a boatload of money.
>developer left, project dumped on him
>doesn't know the solution
not likely
> There’s this Network Partition thing, it’s kind of a big deal
How could someone using RabbitMQ cluster not consider how the cluster would behave during a partition? This is exactly the kind of thing that should be tested in a safe environment before running the cluster in production.
Testing for network partitions is not something one wishes someone else would have told. It is an essential responsibility for anyone in a software engineering role. Not doing some basic testing to understand partition scenarios before running a cluster (any type of cluster) in production is a disaster just waiting to happen.
This looks like an overengineering to me, unless I have missed something. For example, I don't understand this part: "the consumer gets the message and makes a HTTP call to another web service" - why cannot that web service pull the message directly?
I would implement it like this: when a client submits a request, it is added into an SQL database, and RabbitMQ is used only to notify the consumer. The client gets back a secure Job Token that it can use for polling to get the status of the job. The consumer reads the job from the database, executes it and updates status in the database. The client uses polling to know when the job is done.
Of course, if you don't like polling, then you can use a WebSockets daemon that would notify a client when the job is done.
It could e.g. be an external service run by a third-party, or an off-the-shelf service that's hard to add integration to, etc.
That's not a level of data volume that should require any kind of distributed messaging.
That's 100M messages per day or 1157 messages per second. That volume could be handled by a single DB table in an ACID way eliminating the need for a separate cluster
The MQ is great for scheduling jobs, passing data between different parts of a given process and generally detaching systems. However fast it is not, unless you're using something like Zero MQ for slow IPC.
What I see the issue is: We have exactly Nobody in our organization whose full-time job is to perform deep[0] testing.
[0] https://www.developsense.com/blog/2022/01/testing-deep-and-s...
"Risk coverage is how thoroughly we have examined the product with respect to some model of risk."
There had been no modeling of any kind of risks related to networking issues with this technology, so there is no surprise that something went terribly wrong.
distributed clusters are cool, especially messaging ones, but you have to know how to manage them or you're adding eventual points of failure, not removing them.
There's probably a way to write some code to drain a split brain's data but I suspect nobody has the time to write this code.
You could have a process where you block writers to the minority then drain messages and enqueue them on the majority.
So it seems that RabbitMQ wants to advertise how scalable it is by enabling clustering by default and accepting connections from anyone even though it is not secure. It is targeted at large corporations with giant clusters and doesn't care about developers who just need a single instance. Despite my guess that most developers actually do not have a volume of messages that would justify setting up a cluster.
And as RabbitMQ is written in an exotic language (Erlang) I cannot even read the code.
This is the problem not only with RabbitMQ, Elasticsearch also has clustering settings enabled by default, so when you try to run two independent instances, they connect to each other and start exchanging data and creating problems without you expecting it. So annoying and so difficult to disable. In earlier versions, as I remember, they would also try to scan local network and connect to any node they could find.
It doesn't make sense if you first open the port and then block it with a firewall.
Also, this is a bad idea because firewall config is kept separately from RabbitMQ config, it is easy to forget that you have a RabbitMQ instance exposed to the whole Internet and accidentally unblock the port.
if you arent a distributed system, why do you even need a multi node messaging system?
there are way better solutions for single node.
> 25672: used for inter-node and CLI tools communication [...] Unless external connections on these ports are really necessary (e.g. the cluster uses federation or CLI tools are used on machines outside the subnet), these ports should not be publicly exposed. See networking guide for details.
Individually binding each service to localhost to prevent remote access seems tedious and error prone. What happens if an update exposes a new port? I like to bind to localhost too, but only as a defense in depth measure.
I also have not seen RabbitMQ opening anything on my firewalls on installation or during the operation. There are some annoying applications that add firewall rules but RabbitMQ was not one.
If someone is solving issues by disabling whole firewall, his access should be revoked immediately from anything related to infra.
Firewall is a proper solution, learn how to use it properly instead of coming up with excuses.
In Debian by default all connections are allowed. And it makes sense: if you open a port, you probably want to be able to connect to it, otherwise why would you open it?
It doesn't make sense if you open a port (to accept connections) and then block this port with a firewall.
Of course this applies to cases when you have a single node, I understand that you might need firewall if you have several nodes.
/s
The firewall solution to a new/unexpected open port is a default deny rule. That's a commonly used best practice. There's nothing to forget.
In case with a single node, it doesn't make sense. If you open a port, you probably want to be able to connect to it, otherwise why would you open it? And if you don't want anyone to connect, then don't open a port. There is no need for a firewall in case with a single node, unless you want custom rules like "only host A can connect to host B".
There should not be "unexpected" open ports. Software should not open ports on external interfaces unless permitted by administrator.
RabbitMQ should just have a switch to disable all unnecessary clustering features, but for some reason they don't want to implement it.
Like a lot of highly complex Cool Tools they're marvellous right up until you hit an edge case or performance threshold or an odd failure state -- and then you find yourself copy-pasting increasingly trimmed-down log entries, desperately seeking people who've hit the same problem, or rather, people who've solved the same problem and thought to describe it on the Internet.
If you're on fresh software, a fresh version, or just doing something mildly off-label, this can be a despairing process.
It’s also that you don’t even state if you talk about DevOps as a philosophical approach to how software should be designed, written, and run. Or if you talk about a job title.
The main issue the article is talking about that you need to think about the future and day to day operations before you start running production workloads and even think about failure modes and handling of it before you can plan your system. And it really helps to simulate something like network partitioning before you write the bigger part of your application.