No function runs forever.
Also if you want to reboot/repurpose the server, 15 minutes is max wait time.
And finally, you don't want people running long jobs here when you have a solution there.
Maybe you could let users indicate the operation will take a long time... but if the user knows the operation is long running in advance, why not just guide them to a more suitable system?
15 mins max runtime simplifies the resource management and avoid abuse. If you have workload for long running jobs, then that should go to something like AWS EKS/Batch/SageMaker.
That being said, things can change, if more and more people requires long running capacity for Lambda (though I am skeptical of that, as Lambda abstracts the underlying hardware away and is supposedly flexible to the requirements)
Chris Munns - Lead of Dev Advocacy for Serverless@AWS
What I care about is: * Scale to 0, and automatic scaling up without configuring it
* Automatic patching of the OS
* Fault isolation
Lambda gives me that. So each one runs for 15 minutes, processing all data in an SQS queue.
I do wonder if Fargate would be cheaper per millisecond? Dunno.
Someone mentioned step functions above. All of our steps run on spots. We also have some tasks like you have that read off queues and do processing, which also all run on spots.
If data doesn't have to be processed sequentially an option is to configure the AWS Lambda function to get invoked for new data in the SQS queue [1], so you don't have to care about manually fetching data from SQS at all.
[1]: https://docs.aws.amazon.com/lambda/latest/dg/with-sqs.html
https://aws.amazon.com/about-aws/whats-new/2020/11/aws-lambd...
But being real, we hear you on this one. I can't comment on API Gateway's roadmap here but this is something both teams is aware of. The reason it is the way it is today is for a valid reason. But this is def something we hear pretty often.
- Chris
;)
One example we've been wrestling with is a merge operation. Usually it's merging about 1000 records which completes in a few seconds. But every once in a while someone kicks off a job that tries to merge 1,000,000 records and it times out.
We want the benefits of serverless (scale down to zero, up to infinity at the drop of a hat) but these edge cases mean we're having to evaluate other options.
An hour or two would be a good start; then it'd cover 99.9% of requests. With a few hours we could add more nines :)
Was this recent, should I take a look at this again? My issue with Fargate as recent as a year ago was that running the same workload on ECS (if you can use your cluster nodes efficiently) was twice as cheap (even without reserved instances).
It's best suited for jobs that can be broken down into tons of small individual computations, or to respond directly to HTTP requests. If you can fit your pipeline / application into that model it's usually beneficial: Instant scaling, retries, reliable etc. Mixed with other concepts like SQS you can build pretty powerful things without having to pay when there's no load.