AWS SNS vs. SQS – Main Differences
blog.serverlessq.com
blog.serverlessq.com
There are a couple of things I‘m left wondering: When the SNS is used to notify through Email, Push or SMS then even in the fan-out pattern, no SQS is involved on these paths, right? (seems to me that only on A2A paths SQS would be involved, at least from the diagrams) So is there anything else to help with reliability there that the notifications actually do go out?
For workloads where you e.g. want to alert users through different means, not every user might have the same options selected, e.g. some only Push and others SMS and EMail; would that be modeled through different topics which have different combinations of subscribers attached (some only 1) or would it be better to skip SNS and push multiple messages to different queues directly?
Thanks for your kind words.
Regarding fanout. Yes exactly. Fanout doesn't mean it needs to involve SQS just it can involve SQS. It is also called a fanout pattern if you do a A2P and only notify Emails, SMS, etc.
To your second question how the architecture would look like for different preferences of architecture.
I think the main benefit of that architecture is that customer can subscribe to a topic. That means if your user A subscribes to the topic for Email and not in-app notification that is fine. It would be also just the one topic.
The consumer/subscriber has the power to subscribe and unsubscribe to topics (similar like you can to newsletters basically). That is one of the main benefits.
With a queue the producer would need to define which consumer will get the message and most probably it will be another application.
Does that help? :)
- SQS to send alert, all modes
- Lambda reading that queue, filtering on user settings
- SNS per mode of communication, eg email or text
My thinking is that you’d want to filter on user preference early, to prevent repeated work. A benefit of this approach is you prevent combinatorial complexity if you have both selection on kinds of alerts delivered and way to deliver alerts: the Lambda can handle all of that based on the user settings. And still a single SNS per communication channel.
Your gating Lambda can also implement other features, like volume aware decisions — where it eg, rejects “marketing” messages if there have been too many to a single customer recently while still allowing through “transaction” messages.
My understanding is that you can have multiple consumers of an SQS through the use of visibility timeouts[0]. Once a message is consumed it is as if that message doesn't exist for all other consumers until it reaches a timeout period or is marked done by that consumer. You can also manually mark a message as being ready for other consumers. This moves the message back into the queue for the other consumers to see.
I'm going to be linking this article to my team. We've been talking about moving to SNS/SQS/etc. and this article helps understand the use cases and distinctions better.
[0]: https://docs.aws.amazon.com/AWSSimpleQueueService/latest/SQS...
1. The main point of the visibility timeout is to handle failure. A message is read by a consumer; the visibility timeout starts; that consumer finishes some processing; then deletes the message from the queue. But, what happens if the consumer encounters a fault during processing which destroys its ability to even tell the queue it encountered a fault? The visibility timeout protects against that; the message just naturally reappears in the queue for processing by another consumer. If one overloaded the visibility timeout to also mean "other consumers should process this", you'd lose the ability to handle faults.
2. It also screws up deadletter redrive policies, which are primarily based on visibility timeout lapses (in addition to communicated failures). You basically could not reliably put a deadletter redrive on your queue, which again just means, you're protecting against fewer failure modes.
3. There would be natural, avoidable latency in waiting for the visibility timeout on every fan-out, whatever you set it to. 1 second? 100 consumers? That message is just clogging up the queue for over a minute as it gets fanned-out to everyone.
4. Consumer1 eats the first message, then times-out its visibility; its back in the queue; there's no way to ensure that message isn't just processed again by Consumer1 instead of Consumer2! You're basically tossing a coin and hoping that, eventually, Consumer2 gets its turn at the message, all the while having Consumer1 reprocess the message an indefinite number of times.
5. Someone has to delete the message. Who? The "last" component to touch it? Once all the other components are done? How do you coordinate that? Theres no guarantee of ordering on when each component sees the message. You'd need some kind of external state, and at that point, why are you even using SQS?
You could theoretically have each consumer read from a queue, process the message, delete that message from the queue, then redrive the message into a new queue for processing by another consumer. This may make sense if you have strict ordering needs for processing but still want the benefits of SQS. You could even have it redrive into N queues for N consumers at the same time. But, at that point, why? We're trying to put a square peg in a round hole; SQS is designed for single consumers. There are far better and simpler tools out there if what you're looking for is multi-consumer fan-out.
> there's no way to ensure that message isn't just processed again by Consumer1 instead of Consumer2!
Correct, but this isn't the job of the pipe. Smart endpoints, dumb pipes.
I built a data ingestion system that handled an average of 300 messages/sec, peaking at 1,000, and writing to a single R3 RDS instance. You can do a lot by pushing simple scaling strategies to their limit. Everyone thinks they need to handle web scale, but really you just need to handle your scale.
Also, two different queues being two different buffers that have durability issues can in an improperly conceived architecture amount to a distributed RAID0 of messages.
It really depends upon the tolerance to message duplication, SLA needs, and how prioritization should be handled. At a previous place we had multiple consumers for multiple SQS queues representing different priorities within the same region and it worked fine for many years with the primary headache being message de duplication handling being tricky.
The idea of having multiple homogenous consumers shouldn't be controversial; that's just horizontal scaling. And, well, at least until a few hours ago I also would have said that the idea of having multiple heterogenous consumers is also uncontroversially bad. But I guess everyone has "their way" of doing things.
Its also important to note that there's a third situation I see somewhat often: maybe call it homogenous delegated consumers, whereby you've got messages like '{"type":"SendDM", "content": {}}'. Or maybe: '{"type":"SendDM", "action": "UpdateDB", "content":{}}'. The consumers are still homogenous, they all run the same code, but they may internally delegate the message to do different things depending on enums within the message. This is pretty ok; its different because at least you'd never have a consumer hit message and be like "I don't want this take it back".
Though I'd caution against it; just understand that its something of a 'hack' to make one queue act like N queues, and that's ok if you're small and have a good grasp on the problem domain. The big issue it will inevitably run into is: some queue message "kinds" will take a lot longer to process than others; and so if you're e.g. overloading a queue to handle both a simple email send and a much more complex asynchronous database update, you'll inevitably get delayed emails. Absolutely inevitable. But, it can work for a time.
If you have multiple heterogenous consumers, do not use a single SQS queue.
I can't even comprehend how you would engineer around the issue of consumer re-processing. You can quote metaphors all day; if you love the idea of dumb pipes, why doesn't the city transport clean and gray water in the same pipe? Do you want to wash your hands using flushed toilet water?
Similarly, you can't engineer around heterogenous consumers grabbing a message, putting it back in the queue, then consuming it again. You can make them smart! You can have them say "woah hold on, I already saw that message I don't need to see it again put it back". Or, you can make them idempotent so reprocessing isn't undesirable. But its still reprocessing; its still a huge waste, and will probably require external state to manage. Moreover, there's literally no system guarantee that Consumer2 will ever see that message; it'll probably see it, fifty-fifty, well then again if one consumer is faster at accessing the AWS API than the second, who knows, anything could happen, but at least its convenient?
The city doesn't require every household to have gray water filtration. Because that would be insane. The pipes don't have to be "smart". We just build two pipes!
Can you technically use a Queue as a topic for pub/sub? Yes. But should you? Probably not. You're much better off not using SQS for that and instead using SNS.
The topic has "zero or more" subscribed queues, and when sending to the SNS topic, you don't need to know how many subscribers there are right now.
In many cases, it's more a flexible equivalent to write to a SNS topic; not directly to a SQS queue.
The article calls this "the fanout pattern"
Producer don't need to know to send it to which queues
Yes combining good content with advertising my product is the goal but I really focus on providing benefits with the content I create :)
Since I build on top of these great services I want to show at least how I am utilizing everything.
This is a great example of content marketing that uses a question to both answer and impart value in a very logical and compelling way. Through understanding the differences between SQS and SNS the reader will know how to use those in the future, get a clearer idea of the complexity, and be presented with a much simpler and still full-featured alternative at the end. Well done.
It's mostly important to remember when you're wondering why your mostly-empty queue is still costing you $0.26/month, even if you've got the receive message timeout set to 20 seconds.
But the whole point about costs with empty queues is important. It is definitely important to understand if you want to customize queues for example with long polling. This parameter changes the time your lambda will poll from your queue
Under the hood it might be polling. But it gives the illusion it’s push and it’s bloody useful.
Note that this can cause issues: say you have a time sensitive application that receives a batch of "bad" messages which cause failed lambda invocations. The poller will slow down and the throughput will drop drastically, even though your intention might be for the lambda to continue processing at the same rate and power through the bad messages.
This behavior can be disabled with a support request.
If you want, give ServerlessQ a try and let me know if you need any help! You find my contact info on my landing page: serverlessq.com or DM me on twitter: twitter.com/sandro_vol
I left that comment there for folks who didn't know that was possible, as going the SNS -> SQS -> Lambda route is what is most popular and most written about.
Used it recently, great improvement: https://aws.amazon.com/blogs/aws/enhanced-dlq-management-sqs...