Every other queue system I’ve encountered has terrible failure modes, buggy clients, weak semantics, and huge serialization overhead.
Every other queue system I’ve encountered has terrible failure modes, buggy clients, weak semantics, and huge serialization overhead.
http://mikehadlow.blogspot.com/2012/04/database-as-queue-ant...
Transactions make the problem worse with locking.
Read all the comments on that post.
Transactions make the problem easier, not “worse”.
But if they had to have queueing logic in the database, you would have to have locks on the table holding the messages and the statuses of the messages. Not to mention even more of a database load to delete the messages. Of course on top of that you are constantly polling the database.
Now compare that to a purpose built queueing system where you just read from the queue, it automatically marks the messages as unavailable- “in flight” - and that will automatically requeue the message after a certain amount of time of it isn’t deleted.
As opposed to a database system where either messages get stuck in “processing” status of a job fails or you have another process reading the queue and “fixing” the issue after a certain amount of time has passed.
And finally, what happens if you need a fan out type of queue? Where you have one producer and multiple types of consumers?
If it’s sane, it’s doing exactly the same transactional row locking an RDBMS is highly optimized to do.
I have apps with RDBMS queues doing thousands of work items per second with 16 workers. Very few applications need more than this.
Introducing a whole new subsystem and API for queueing is a bad engineering decision in most cases.
Fan-out is generally an anti-pattern, but if you need it, then explore other options
Note that I am talking about Kafka-style linear queues in an RDBMS, not pub-sub.
Triggers or app code do very fast transactional inserts into queue as part of the initial write; consumers fast-poll with “update top 1 ... where status = <unprocessed>” and back off polling exponentially when the queue is empty. This is just a few lines of code trivial to do correctly and cannot deadlock in READ COMMITTED or SERIALIZABLE modes.
So introducing a purpose built queueing system - that people have been doing for decades instead of using a database is technically bad?
A queueing system at most has to respond to a request, set a flag for “processing” and it’s done and since by definition, a queue only has to read the top most item, it is more efficient.
Besides, how do you handle a message that is sitting in a “processing” state because the process that originally read it crashed and didn’t change the status?
Triggers or app code do very fast transactional inserts into queue as part of the initial write; consumers fast-poll with “update top 1 ... where status = <unprocessed>” and back off polling exponentially when the queue is empty. This is just a few lines of code trivial to do correctly and cannot deadlock in READ COMMITTED or SERIALIZABLE modes.
Or instead of reinventing the wheel with your own bespoke database as a queue that has all of the maintenance overhead you could just use a queueing system that is already optimized for that use case.
Note that I am talking about Kafka-style linear queues in an RDBMS, not pub-sub.
In general parlance, most people would call Kafka style processing stream processing.
Even then, why write and maintain a pseudo streaming process that you have to maintain instead of just using Kafka where you can let it handle all of the fault tolerance, partitioning etc?
It’s about like a past company that I worked for where the “architect” had his own homegrown ORM, encryption scheme, and configuration system instead of just using an off the shelf solution because he thought his system was its own special snowflake.
95% of applications only need a queue for “do this thing asynchronously so the user doesn’t have to wait.” This is where a DB table queue is the best solution. Introducing RabbitMQ or any other service in such a common case is a terrible idea.
Every real-world message passing implementation I’ve encountered ends up with the complexity of a bespoke state database for each consumer to handle re-ordering and crashes, as well as poorly written code and more state in a database to handle message replay. Not all messages can be idempotent, and consumers never end up stateless in the real world.
All of these solutions lost events in production under various conditions.
My point is: think about if you really need the complexity of managing another service in production, when all you really need is “do this thing as soon as you can”.
So now the common definitions of things is wrong. So how do you propose that multiple systems that all care about a single event get notified? You realize that processing queues and using a fan-out pattern has been done for decades?
95% of applications only need a queue for “do this thing asynchronously so the user doesn’t have to wait.” This is where a DB table queue is the best solution. Introducing RabbitMQ or any other service in such a common case is a terrible idea.
Because based on your anecdotal experience you can say with confidence that “95%” of people are doing it wrong....
Every real-world message passing implementation I’ve encountered ends up with the complexity of a bespoke state database for each consumer to handle re-ordering and crashes, as well as poorly written code and more state in a database to handle message replay. Not all messages can be idempotent, and consumers never end up stateless in the real world.
Well, maybe “in your real world”, but people have been managing queues with idempotency, statelessness, and out of order execution for decades.
As far as handling crashes, there is nothing to do. Once the message is in process for a certain amount of time (“in flight”) and the consumer hasn’t acknowledged successful processing, the queueing system automatically puts the message back in the queue. After a certain number of retries it goes into a dead letter queue.
All of these solutions lost events in production under various conditions.
Don’t blame a poor implementation on the technology. I preach to people all of the time unless you are working at Google or even Twitter scale, you’re not a special snowflake that needs to reinvent the wheel and try to re-solve solved problems.
My point is: think about if you really need the complexity of managing another service in production, when all you really need is “do this thing as soon as you can”.
“Managing” RabbitMQ is not rocket science. But these days, I don’t deal with managing infrastructure. That’s what cloud providers are for.
We have many issues with RabbitMQ where most of our workload is background tasks with dozens of queues.
> That’s what cloud providers are for.
What cloud provider queueing system do you recommend?
Also since all queueing systems basically serve the same purpose, it’s easy to layer the AWS SDK calls under your own facade classes to reduce the dependency on AWS’s services.
All that being said:
Simple one consumer/one or multiple producers system:SQS
Multiple consumers/one or multiple producers: SNS/SQS
Kafka equivalent: AWS Kinesis or AWS MSK (Manager Kafka). I haven’t used Kafka but if you don’t want to use an AWS specific service and want easy portability, it couldn’t hurt to do a proof of concept.
With AWS SQS/SNS there are no servers to manage. You just create your queues from the web console (not recommended), use the CLI, CloudFormation, or Terraform.
The problem I have with SQS and SNS (and Celery) is you can not just throw tasks into it and eventually the system based on some hints scale up / down the workers. Of course you can rely on Lambdas but then you are locked with amazon (not need to mention that you can not control how much the lambdas will cost you).
Also, I disagree with you point that question tech/tool status-quo is NIH, hence is bad. I for instance, would like to be able to avoid vendor lock-in. Also, reinventing the wheel allows to stay in control. Using RabbitMQ and to some extent Celery or Kafka locks you up without much control since it's foreign code base with alien language.
The problem I have with SQS and SNS (and Celery) is you can not just throw tasks into it and eventually the system based on some hints scale up / down the workers. Of course you can rely on Lambdas but then you are locked with amazon (not need to mention that you can not control how much the lambdas will cost you).
Well two things:
Lambdas aren’t some magical thing that requires a lot of changes to your code. All lambda requires is one function added to your code base that takes a JSON Event and a lambda context. The only thing that your lambda handler should be doing is deserializing the event into your domain object and calling your business layer - the same thing that your regular entry point should be doing.
I have a C# solution that has three modules (assemblies). One has all of the AWS dependencies with the lambda entry point, one is a regular .Net executable with TopShelf integration to create a Windows service and the third is the actual business logic.
The lambda project takes the SQSEvent gets the message body, deserislizes it and sends it to the assembly with the business logic
The second, runs as a Windows service reads from the queue, deserializes the message and sends the object to the same assembly with the business logic.
When I push the code, AWS spins up a Linux Docker container using Code Build that builds both the Linux based Lambda and the Windows executable. There is no “lock-in” to lambda. We deploy the Windows service for QA testing.
Also, reinventing the wheel allows to stay in control. Using RabbitMQ and to some extent Celery or Kafka locks you up without much control since it's foreign code base with alien language.
We as software developers get paid to produce solutions that add business value and that allow the business to focus software development where it has a competitive advantage. You don’t add business value by reinventing the wheel. Besides that, no developer wants to come into a shop where all of the cross cutting concerns like logging, queue management, database access, etc are all some bespoke system where the architect thought they were a special snowflake. I would much rather go onto the Internet where if I have an issue, I can probably find someone else who had that same issue than trying to find the original creator of CustomQueueManager who may not be at the company any more.
In the case of AWS, I have an “easy button”. After I have gone through all of the obvious steps and something is still wonky with a managed service where I am using their SDK, I can just take advantages of our business support plan, open a ticket and start a chat. They will not rest until they figure out the issue.
As far as scaling out without using lambda. That’s easy. Just setup two alarms - one for when your queue is under a certain size and one when your queue is over a certain size and use the alarms to trigger autoscaling within an autoscaling group.