So lambda pricing scales down to cheaper than a VPC, but it also scales up a lot faster ;)
So lambda pricing scales down to cheaper than a VPC, but it also scales up a lot faster ;)
The above tracking method works for containerized things. For virtualized things, it’s different: When a Linux system has nothing to do (all processes sleeping or waiting), the kernel puts the CPU to sleep. The kernel will eventually be woken by an interrupt from something, at which point things continue. In a virtualized environment, the hypervisor can measure how long the entire VM is running or sleeping.
https://blog.cloudflare.com/cloud-computing-without-containe...
I've published LambdaFlex as an Infrastructure as Code (IaC) template. It automatically scales and manages traffic between AWS Lambda and AWS Fargate [1].
This setup leverages the strengths of both services: rapid scaling, scaling down to zero, and cost-effectiveness.
[1] GitHub: https://github.com/okigan/lambdaflex
To reduce the chances of being long-running, the api would be async, meaning: the client makes a request and gets a claim code result quickly, with the api host to process the request separately and then make a callback when done.
Lots of issues to consider.
Anyone have experience doing this? Would be great to hear about it.
If your hammer fails to drive a screw into wood, you don't need to contemplate a new type of low friction wood.
To go beyond my original question:
In terms of how I have thought about them before, webhooks are similar, but not quite the same thing. I've really only dealt with webhooks when the response was delayed because of a need for human interaction, for example: sending a document for e-signature and then later receiving notifications when the document is viewed and then signed or rejected.
I haven't experienced much need to make REST APIs async, only in cases where the processing was generally always very slow for reasons that couldn't be optimized away. I haven't seen much advocacy for it either.
However, if we think about lambda-based clients, then it makes a lot more sense to provide async apis. Why have both sides paying for the duration, even if the duration is relatively short?
update:
Whether or not it's cheaper for the client depends on the granularity of the pricing of the client's provider. For AWS, it seems to be per second of duration, so async would be more expensive for APIs that execute quickly.
I've seen this backfire in production though. E.g. suppose there was a 5 sec limit on an API call, but (perhaps unintentionally) that call was proportional to underlying data volumes or load. So, over time, that call time slowly creeps up. Then it gets to the 5 sec limit, and the call consistently times out. But some client process retries on failure. So, if there was no time limit, the call would complete in, say, 5.2 seconds. But now the call times out every time, so the client just continually retries. Even if the client does exponential backoff, the server is doing a ton of work (and you're getting charged a ton for lambda time), but the client never gets the data they need.
Better in my opinion to set alerts first.
3s would rarely be acceptable, much less 30s. Not that it never happens, but it should be rare enough that the cost of the lambda isn't the main concern. (Or if it's not rare, your focus should be on fixing the issue, not really the cost of the lambdas.)
Anyway, I think you'd typically limit the lambda run time limit to something a lot shorter than 30 sec.
The simpler solution is often to just lift the app from lambda to ECS without internal changes
If it can be avoided and use something like sqs to call the other lambda, die and then resume on a notification from the other lambda.
That can tricky, but if costs are getting bad it’s the way.
Any other logic is just normal SPA stuff.
If you set the min concurrency to 1 it's pretty snappy as well.
Running clusters of VM/container/instances to “scale up” to transient demand but leaving largely idle (a very, very common outcome) then an event oriented rewrite is going to save a big hunk of money.
But yeah, if you keep every instance busy all the time, it is definitely cheaper than lambda busy all the time.
If you're using k8s, HPA is handled for you. You just need to define policies/parameters. You're saying as if rewriting an entire app doesn't cost money.