S3 trickery: using it as a scheduler
hackernoon.com
hackernoon.com
Instead of doing the simple thing and running a server process that periodically checks some queue for tasks to execute, this "scheme" involves triggering a lambda function every minute, to do some IO operations on a distributed file store where the execution date metadata is embedded in the filenames to see if there is some task to execute. If there is, some more distributed IO operations are required to move the file to another part of the distributed file storage, where some listener will notice a new file was added and trigger the actual lambda function with the task-specific code...
But hey, you're not running any servers yourself...
As far as using a queue, lambda also supports triggering based on a queue.
Waking on items in queue would be better, though, and AWS has it built-in; the authors just did not use it for some reason.
Like Amazon's new SQS Lambda trigger you can then schedule a serverless function that's triggered by new items in the queue to have arbitrary scheduling of tasks.
Using their Python API it's pretty nice:
from datetime import datetime, timedelta
from azure.servicebus import ServiceBusService, Message
sbs = ServiceBusService('task-queue', shared_access_key_name=key_name, shared_access_key_value=key_value)
d = datetime.utcnow() + timedelta(minutes=1)
task = {"some": "object"}
sbs.send_queue_message(
"task-queue",
Message(
task,
broker_properties={'ScheduledEnqueueTimeUtc': d}
)
)- set a CloudWatch scheduled rule to schedule a lambda.
- or set a rule to send an sns, subscribe a queue to the sns topic and subscribe a lambda to the queue.
Agree regarding sns, but in our case sns == s3
Why use S3? S3 events aren’t guaranteed. SNS was specifically designed for this use case.
https://docs.aws.amazon.com/AmazonCloudWatch/latest/events/S...
The title should be "Using s3 as storage and some side thingy that runs every minute as a scheduler".
You can measure the impact and potentially automate recovery of missed events by:
1) Keep a track of events published.
2) Generate an S3 inventory daily.
3) Compare events received to objects listed in inventory.
You rarely end up with fewer events than objects, but it does occur.
I've personally only observed this with SNS target, but due to it being a problem with S3, I believe Lambda can fall afoul of this too.
A much more resilient approach would be:
S3 event -> SNS Topic -> SQS Queue -> lambda.
and set up a dead letter queue for the SQS queue.
It doesn’t help with the reliability of S3 events (and I’ve never seen that happen), but it does help if their is an error running your lambda.
Move the S3 object after processing it. As long as you move it to a bucket in the same region, there aren’t any charges.
Then if you are really paranoid, you can have a timed lambda that checks the source S3 bucket periodically and manually sends SNS messages to the same topic to force processing.
This seems overcomplicated compared to using a regular timed event to trigger a lambda and having it decide what to execute conditionally.
Mainly just calling out a gotcha where you might quietly miss out on scheduled events with no warning.
For example:
1. Object written to successfully to bucket.
2. `s3:ObjectCreated:Put` is _never_ delivered.
The possibility of duplicate events are warned about a lot in the AWS ecosystem, and this sets up an expectation of "at-least-once" delivery.
If only; this is AWS. You pretty much need premium support just to tell them their products are broken. Before someone says "forums", the forums are a joke.
All Cloudwatch would have to do is implement a recurrence = 1 feature but I’m guessing it's not a common enough use case.
So when doing one-time events (by specifying a year) - the triggered lambda need to delete it.
Or have a hourly scheduled lambda to delete already fired one-time events.
Or maybe CloudWatch is so smart it can delete them by itself?
And if you really need more than 50 rules, it’s just a soft limit. You send a request to support and they will raise it.
https://docs.aws.amazon.com/AmazonCloudWatch/latest/events/S...
In dynamodb you can set TTL for a record. When the record expires it will fire event to your designated lambda. That's it. You simply write a record, wait for the record to expire and you get notified on your lambda.
And as with everything overly flexible, people will abuse it and build over-engineered solutions on top of it, when simpler solution is already present, but possibly not obvious or poorly documented.
https://docs.aws.amazon.com/cli/latest/reference/events/put-...
The data you attach will be included in the SNS message or the data that gets sent to the lambda handler.
You can put static JSON as an event from either the AWS Console, the CLI, or CloudFormation.
Or you create a rule for a specific date / time. Any reason why that would not work?
Now I kinda want to try that myself.