Amazon Launches New RabbitMQ Message Broker Service
aws.amazon.com
aws.amazon.com
bad - my job is really boring
It's only software engineers that bemoan their lives getting easier, so they can spend more time working on other problems higher up the abstraction chain.
Longterm isn't this bad? For example, if you can do it with 5 years experience. At 15 years experience you'll likely be too expensive for companies to want to hire you. They'll just hire more junior people.
The first is application complexity. A lot of real workloads are quite complex but do not need to scale out a lot. The cloud and all the hype is organised around simple workloads that need to scale out easily which is an easy win. So for example Netflix or a SaaS application with a few tens of endpoints and a React front end or something. The wide sprawling real businesses are a terrible fit and tend to get rather expensive rather quickly when you start putting their workloads into the cloud. There is marginal aggregate cost benefit over actually buying hardware ($4m a year SQL server clusters are a reality in the cloud), the real benefit being only agility.
The second is simply "hooking up components" sounds really easy. But it's not. I think perhaps 50% of my time is working out why X won't talk to Y or why Z is broken and finding some opaque abstraction which doesn't allow me to get to the bottom of the problem. It's very very easy to turn your deployment into a complete tangle of chaos and circular dependencies which are very hard to rationalise and automate even with state of the art automation tools (which I will say tend to melt in your hands). This is existing layering on top of the same concerns you had before rather than a different one.
Thirdly we have to work out the difference between mature products and hype. Nearly all solutions are described in little blog snippets that make things look really easy for a specific and narrow use case but realistically things are really fucking complicated and in some cases absolutely awfully described in documentation. In a lot of cases, including AWS, it's actually hard to find someone at the cloud vendor who knows how something works when you break it. And sometimes there are solutions which are just absolutely dire. Again pointing the finger here at Amazon's managed ElasticSearch.
Fourthly, you end up being perpetual bean counter afraid of the rube goldberg machine waking in the middle of the night due to some event you didn't anticipate and drinking the content of your credit card in a few minutes. Some of the cost management and spot instance management software automates this rather nicely into a whole cluster of new failure modes as well just as if the complexity wasn't enough already. A trite version of this is "saving money costs money and sometimes the benefits are less than the costs"
So what you end up doing is trading your original problems for a set of new and shiny ones which are possibly even more complicated.
But at least you only have one vendor to shout at, which is a net win if you've ever tried to get HPE and Cisco to work out what fucked up mess is going on between their two lumps of iron.
I digress but be careful with assumptions about it being magical unicorns. They poop and you have to shovel it.
Why are they better?
RabbitMQ is a smart play as Rabbit is very easy to use, understand, and troubleshoot at the low end (which is where I suspect the vast majority of queue systems live).
It also has a feature which is actually really hard to do (and sqs doesn't do). Guaranteed delivery of a message once.
That was THE reason we never migrated to SQS, there are scenarios where SQS can double deliver. Our codebase was built up from nothing over time and couldn't gracefully handle double delivery of messages in all scenarios. We could have refactored, but it wasn't worth the work when we were already doing a half billion in revenue without getting even close to the limitations of rabbit AND were close to selling (which we ultimately did).
AWS is great at selling multiple slight variations of the same product. If you look you can usually find ONE variation that works for you. The real test will be if the billing isn't garbage (garbage billing is why we didn't use their other AMQP service and part of the reason why we don't use things like EKS or Managed SFTP despite having the need).
That flies in the face of my distributed systems knowledge. It's not possible in some failure cases.
If your acknowledgement of a message gets lost (because either server involved or the pipes in-between fail) you've processed the message already but the queue server will think you haven't. It either has to resend it (duplicate delivery) or it ignores acknowledgements all together (drops messages that it sent you, but you didn't process - maybe because your server failed.) So the choice when there is a failure in the system is between at least once or at most once - exactly once cannot be guaranteed.
I'm not aware of any way around that predicament.
If I remember correctly SQS is hard limited to a fairly short timeout to requeue messages delivered but not acked. In rabbit it's much more configurable.
Also regular rabbit hosts support the kludge pattern of, 'just run one host and accept if it goes poof you can lose messages,' which is useful if you don't want to bother with the complexity of clustering or are on a shoe string budget.
Lastly you get a nice user interface with the management plugin and you can stand it up locally with docker compose (without depending on AWS for dev or any of the 'aws but on your laptop' solutions).
Though most people are just going to use a framework plugin to manage the messaging layer, so what's behind that is largely irrelevant.
Yes we could do that, but we had already been using rabbit in a bunch of places. It made no sense to change it.
If you have a Lambda function processing SQS messages they just get dumped in your handler method and it your function runs successfully they get automatically removed from the q. If your lambda fails, the message reappears after the visibility timeout out subject to your redrive policy
This is justified by “less maintenance, and easier deployment” but the reality of the situation is, it’s not worth giving your freedom up for, and to a lesser degree - if your platform becomes popular, you end up spending the same amount of time tweaking and optimising to match the idiosyncrasies of their implementation anyway.
But the most important part is vendor lock-in, it’s bad.
At the end of the day, you can still rewrite your code and switch in both cases. You can end up in a tough spot if the OSS community loses interest in the software you've already bet your complicated app on, as well
It also doesn't make sense to rewrite their current software which is probably abstracted for multi-cloud to support re-selling.
True -- I do think passing on the cost and taking a tiny margin with drastically reduced maintenance cost could be an attractive business model at scale though.
> It also doesn't make sense to rewrite their current software which is probably abstracted for multi-cloud to support re-selling.
I have no idea what their current software looks like, do you have any inside knowledge?
If they have abstracted, then they probably have multiple implementations of a similar API -- this is just changing one of them (or maybe even cloning it to reduce possibility of breakage). This might be as simple as just changing the AWS-specific provisioner to call out to AmazonMQ instead of EC2, or changing some code that generates terraform/pulumi scripts.
One thing I think they'd have to deal with is the fact that they support custom plugins that AmazonMQ may not.
Like other AWS products (RDS, Elasticache), there’s limitations since they provide protocol interoperability with proprietary tech behind the scenes.
https://aws.amazon.com/blogs/aws/firecracker-lightweight-vir...
Edit: In comparison to the scale of Amazon and the scale of contribution of other similarly-sized tech companies. Firecracker would rank as a more major contribution in my book if it wasn’t a cut-down fork of a pre-existing (and still active!) project.
If you're not using AWS you're not fully leveraging your software team, because that means you've got people spending time building and supporting these sorts of internal systems. That should only be done when you reach large scale, at which point people have leveraged other AWS synergies making it harder to exit the platform.
And then you're bleeding money!
And no, learning a specific AWS service is not a solid career investment. Tell that to anyone who tried to learn CloudFormation and then SAM and now forget everything because AWS pushes you to use CDK.
This is simply not true at all, and flies in the face of real world usage.
The only concrete and objective selling point of AWS is it's global coverage of data centers, and the infrastructure they have in place aimed at delivering reliable global-scale web services.
The problem is that the companies who actually operate at such a scale and with such tight operational requirements can be counted with your fingers. That count then drops down to a fraction once you start to do a cost/benefit analysis.
The rest of the world is quite honestly engaged in cargo cult software development.
And no, doing AWS is not simpler nor more efficient. You might launch an EC2 instance with a couple of clicks, but to navigate a service designed with global scale and multiple levels of redundancy across the same service and with tight integration and dependency across half a dozen AWS offers which may or may not be redundant or competing... No, that is not simple or allows for any type of time efficiency.
Hell, with AWS you do not learn how to manage or operate Infrastructure. With AWS you learn the AWS dashboard,and learn pavlov reactions to which button you press if you hear an alarm. You never fully grasp the impact or the reaction of pressing a button, and you have absolutely no idea what impact that click will have on your monthly bill.
In contrast, if you need to run microservices chatting through a message broker then your system on OVH or Hetzner or any other barebones system will be comprised of a bunch of nodes where one of them runs RabbitMQ and everyone else points to the RabbitMQ node. You can get everything running from scratch on a cluster managed by Docker Swarm in about 15 to 20 minutes. In the end you have a far simpler service running for a fraction of the cost and ina far more manageable environment.
AWS is resume-driven development fueled by cargo cult development.
I'm not taking a position on whether this is an acceptable level of complexity for the desired feature, of course, just pointing out how one might accomplish it if Rabbit is otherwise desirable.
I've seen growing levels of AWS contribution back to upstream projects over the past four years. Teams start out by operating a piece of software at scale, whether it is Redis, Kubernetes, etc. After they have operated it for a while they discover the bugs or performance issues, or customers of the service complain about something. At that point the team now has enough real world experience with that software to begin to contribute back to upstream.
It takes time: to learn the ins and outs of the software well enough to know where and what improvements should be made, to understand the software's design and history well enough not to make bad suggestions or contributions that were already determined to be dead ends in the past, and to earn the approval of the community and existing maintainers enough to get significant contributions accepted in the first place.
linux (master=) $ git shortlog -ns --author amazon
79 Arthur Kiyanovski
76 Gal Pressman
59 David Woodhouse
51 Sameeh Jubran
45 Netanel Belgazal
38 KarimAllah Ahmed
35 SeongJae Park
30 Jan H. Schönherr
26 Frank van der Linden
18 Andra Paraschiv
12 Paul Durrant
11 Talel Shenhar
10 Shay Agroskin
...I am in no way defending or attacking anyone,I just want to provide a data point.
That it feels the need to respond like that makes me see it in worse light over the matter rather than better.