AWS Lambda Web Adapter
github.com
github.com
So lambda pricing scales down to cheaper than a VPC, but it also scales up a lot faster ;)
Running clusters of VM/container/instances to “scale up” to transient demand but leaving largely idle (a very, very common outcome) then an event oriented rewrite is going to save a big hunk of money.
But yeah, if you keep every instance busy all the time, it is definitely cheaper than lambda busy all the time.
If you're using k8s, HPA is handled for you. You just need to define policies/parameters. You're saying as if rewriting an entire app doesn't cost money.
I've seen this backfire in production though. E.g. suppose there was a 5 sec limit on an API call, but (perhaps unintentionally) that call was proportional to underlying data volumes or load. So, over time, that call time slowly creeps up. Then it gets to the 5 sec limit, and the call consistently times out. But some client process retries on failure. So, if there was no time limit, the call would complete in, say, 5.2 seconds. But now the call times out every time, so the client just continually retries. Even if the client does exponential backoff, the server is doing a ton of work (and you're getting charged a ton for lambda time), but the client never gets the data they need.
Better in my opinion to set alerts first.
The above tracking method works for containerized things. For virtualized things, it’s different: When a Linux system has nothing to do (all processes sleeping or waiting), the kernel puts the CPU to sleep. The kernel will eventually be woken by an interrupt from something, at which point things continue. In a virtualized environment, the hypervisor can measure how long the entire VM is running or sleeping.
https://blog.cloudflare.com/cloud-computing-without-containe...
To reduce the chances of being long-running, the api would be async, meaning: the client makes a request and gets a claim code result quickly, with the api host to process the request separately and then make a callback when done.
Lots of issues to consider.
Anyone have experience doing this? Would be great to hear about it.
If your hammer fails to drive a screw into wood, you don't need to contemplate a new type of low friction wood.
To go beyond my original question:
In terms of how I have thought about them before, webhooks are similar, but not quite the same thing. I've really only dealt with webhooks when the response was delayed because of a need for human interaction, for example: sending a document for e-signature and then later receiving notifications when the document is viewed and then signed or rejected.
I haven't experienced much need to make REST APIs async, only in cases where the processing was generally always very slow for reasons that couldn't be optimized away. I haven't seen much advocacy for it either.
However, if we think about lambda-based clients, then it makes a lot more sense to provide async apis. Why have both sides paying for the duration, even if the duration is relatively short?
update:
Whether or not it's cheaper for the client depends on the granularity of the pricing of the client's provider. For AWS, it seems to be per second of duration, so async would be more expensive for APIs that execute quickly.
3s would rarely be acceptable, much less 30s. Not that it never happens, but it should be rare enough that the cost of the lambda isn't the main concern. (Or if it's not rare, your focus should be on fixing the issue, not really the cost of the lambdas.)
Anyway, I think you'd typically limit the lambda run time limit to something a lot shorter than 30 sec.
The simpler solution is often to just lift the app from lambda to ECS without internal changes
If it can be avoided and use something like sqs to call the other lambda, die and then resume on a notification from the other lambda.
That can tricky, but if costs are getting bad it’s the way.
Any other logic is just normal SPA stuff.
If you set the min concurrency to 1 it's pretty snappy as well.
I've published LambdaFlex as an Infrastructure as Code (IaC) template. It automatically scales and manages traffic between AWS Lambda and AWS Fargate [1].
This setup leverages the strengths of both services: rapid scaling, scaling down to zero, and cost-effectiveness.
[1] GitHub: https://github.com/okigan/lambdaflex
* We write our HTTP services and package them in containers.
* We add the Lambda Web Adapter into the Dockerfile.
* We push the image to ECR.
* There's a hook lambda that creates a service on ECS/Fargate (the first-party Kubernetes equivalent on AWS) and a lambda.
* Both are prepped to receive traffic from the ALB, but only one of them is activated.
For services that make sense on lambda, they ALB routes traffic to the lambda, otherwise to the service.
The other comments here have more detailed arguments over which service would do better where, but the decision making tree is a bit like this:
* is this a service with very few invocations? Probably use lambda.
* is there constant load on this service? Probably use the service.
* if load is mixed or if there's are lot of idle time in the request handling flow, figure out the inflection point at which a service would be cheaper. And run that.
While we wish there was a fully automated way to do this, this is working well for us.
The other is burst, we run a business where customers send us a couple hundred thousand things to do just once a week or once a day. And of course we don't know exactly when the orders will come in. Much easier to let the lambdas run than have large running servers waiting. Autoscaling is possible, of course, but laggy autoscaling won't work, and aggressive autoscaling is basically lambda.
It'd also be nice to have a better log aggregation strategy for lambda than scraping cloudwatch, but that feels less important.
What I mean by "programming the machine" is doing technical tasks that relates to making the cloud work in the way you want.
I've worked on projects in which tha vast majority of the work seems to be endless "programming the machine".
This repeats constantly because the underlying problem is a broken social environment where people aren’t aligned with what their users really need and that’s rewarded.
Can you be specific because that does not sound right to me.
I know both pretty well and seems to me there is vastly more things that need to be tuned and managed and configured in cloud.
I think it would be interesting to do a systematic study of the time required for each.
What about running code? If I deploy, say, Python Lambdas behind an API Gateway or CloudFront, they’ll run for years without any need to touch them. I don’t need to patch servers, schedule downtime, provision new servers for high load or turn them off when load is low, care that the EL7 generation servers need to be upgraded to EL8, nobody is thinking about load balancer or HTTPS certificate rotation, etc. What you’re paying for, and giving up customization for, is having someone else handle all of that.
So said Gary Bernhardt of WAT fame in 2015 https://x.com/garybernhardt/status/600783770925420546
> Consulting service: you bring your big data problems to me, I say "your data set fits in RAM", you pay me $10,000 for saving you $500,000.
Which, in turn, have inspired https://yourdatafitsinram.net/
A lot of web apps are just boring CRUD APIs. The business logic behind these is nothing special. The differentiating factor for a lot of them is how they scale, or how they integrate with other applications, or their performance, or how fast they are to iterate on, etc. The “machine” offers a lot of possibilities these days that can be taken advantage of, so customizing “the machine” to get what you need out of it is a big focus for many devs.
Writing and maintaining code that sets up and configures S3, SQS, SNS, DynamoDB, and Lambda is a non-zero cost, but if it then reduces the amount of code I have to write and maintain to "write the application," that's often a good thing. These services are typically solving the hard and/or tedious parts of the application, so it's actually quite valuable to just focus on developing the core differentiating points of the application instead of solving / maintaining highly available object storage or message queues.
I take a different perspective.
When using a cloud, i need to design my application to work within my selected clouds constraints.
A lot of the extra work then comes down to if using a cloud, not just a single VPS on a cloud provider, you are dealing with distributed computing and have to deal with all the issues of distributed computing. Consensus, error handling, etc.
I agree though, more often than not it's a lot of extra work but mostly due to dealing with distributed architecture/designing for being cost effective.
On personal projects I find designing to be cost effective has me jumping through a lot of hoops and doing novel things. It means I can run multi-region apps with performance better than most webapps for cents/dollars. The equivalent apps I work on in my day job can be costing tens to hundreds of thousands of dollars as they don't jump through the hoops I do. "It's not worth it", "let's look at it later", "we'll just do this, it's good enough for now". As it's someone else's money i just think "meh" and move on as isn't my fight when the majority of people don't want to do it.
Our approach to lambdas is to delegate all the real work to a "library" and just write a very thin adapter around it dealing with the Lambda details (event format etc.). Hopefully this can take care of as much of the wrapper as possible?
Indeed, I had to learn it ! Just as I had to learn how to code in Python (which makes me far more effective now), just as I had to learn LVM2 or iproute2 or C or whatever.
Once the learning curve is processed, AWS allows you to increase your productivity. Of course, and this is very important, you shall still use the right tool for the right job !
But why? The less it changes the the more vulnerable it is, the less compliant it is, the more expensive it is to operate (including insurance), the less it improves and add features, the less people like it.
Seems like a well-worn path to under performance as an organization.
I don’t follow this. Complexity is what leads to vulnerabilities. Reducing complexity by reusing the same, known code is better for security and for compliance. As a security person, if a team came to me and said they were reusing the same container in both contexts rather than creating a new code base, I would say “hell yeah, less work for all of us”.
There are other reasons why using the same container in both contexts might not be great (see other comments in this thread), but security and compliance aren’t at the top of my list at all (at least not for the reasons you listed).
It doesn't have to be about life support for neglected code.
All of that is true, but all of that cost is being paid regardless while the legacy system goes unmaintained. When the company decides to shut down a data center, the choice with legacy systems (especially niche ones) is often "lift and shift over to cloud" or "shut it down". Notably missing among the choices is "increase the maintenance budget".
All these hypotheticals aside, one actual bonus of shifting onto lambda is reduced attack surface. However crusty your app might be, you can't end up with a bitcoin miner on your server if there's not a permanent server to hack.
Especially when you don't have a good local AWS runner, vs using a HTTP service you can run and live reload on your laptop.
deploying to prod in less than 10 seconds, or even /live reloading/ the AWS service would be awesome.
CDK has a new(ish) feature called hotswap that also bypasses CloudFormation and makes deploys faster.
Pushing your image up to your registry, running `kubectl apply`, waiting for it to schedule your Pod and start your containers and exec'ing into them when shit breaks is WAY faster than zipping up your artifact and any layers that come with it, uploading it into S3, and re-deploying your Lambda functions (and any API Gateway apps associated with them), IMO.
When I use Serverless to do this, deploys take up to two minutes, which I sometimes do often because I have no real way of troubleshooting the application within Lambda (definitely necessary if you're using Lambda base images; less necessary if you're rolling your own; likely if you're integrating with API Gateway since documentation on the payloads sent from it is lacking IMO).
We have one service which has gone back to a Go program that runs locally using go run. It's then shipped in a container to ECR and then to EKS. The iteration cycle on a dev change is around 10 seconds and happens entirely on the engineer's laptop. A deployment takes around 30 seconds and happens entirely on the engineer's laptop. Apart from production, which takes 5 minutes and most of that is due to github actions being a pile of shite.
Is it literally just the scaling the 0? And you're willing to give up some number of requests that hit cold-start for that?
If I don't scale to 0 I'd prefer to work on dedicated hardware, anything in-between just doesn't give me enough benefit.
It just costs you the overall quality of your system to do it.
Lambda might be a decent fit for bursty CPU-intensive work that doesn't need to do IO and can take advantage of multiple cores for a single request, which is not many web applications.
Lambda could be a compelling offering for many use cases if they made it so that you could set concurrency on each invocation. e.g. only spin up one invocation if you have fewer then 1k requests in flight, and let that invocation process them concurrently. But as long as it can only do 1 request at a time per invocation, it's just a vastly worse version of spinning up 1 process per request, which we moved away from because of how poorly it scales.
If your response time is 100ms, that's 100k requests in 1 minute.
Lambda runs your code in a VM that's kept hot so repeated invocations aren't launching processes. AWS is eating the cost of keeping the infra idle for you (arguably passing it on).
Serving 10k concurrent connections/requests was an interesting problem in 1999. People were doing 10M on one server 10 years ago[0]. Lambda is traveling back in time 30 years.
[0] https://migratorydata.com/blog/migratorydata-solved-the-c10m...
Say your application uses 25ms real CPU time per request. That's 40 reqs/sec/cpu core. On a 4 core server, that's 160reqs/sec. That's 625 seconds to clear that backlog assuming a linear rate (it's probably sub linear unless you have good load shedding).
So that's 10 minutes to service 100k requests in your example. I'm ignoring any persistent storage (DB) since that would exist with our without Lambda so that would need its own design/architecture.
A 1 CPU container isn't going to handle that many "app" requests unless you have a trivial (almost no logic) or highly optimized (like c static web server) app
Lambda apps don't need to take advantage of multiple cores since you get a guaranteed fractional core per request
One example is retail fire sales. I interviewed with a company that had this exact use case. They produced anti bot software for limited product releases and needed to handle this load (they required their customers submit a request ahead of time to pre scale)
Also useful with telemetry systems where you might get a burst in logs or metrics and want to consume as fast as possible to avoid dropped data and buffering on the source but can dump to a queue for async processing
The 25 ms request you mentioned in another comment is what I'd categorize as extremely CPU intensive.
This adds a lot of cold start overhead? Instead of directly invoking a handler.
As a lambda lives for a period of time you can create a pool for that individual lambda. That pool lives for the life of the lambda.
Now for things like HTTP clients this is generally OK.
If you have a Postgres pool etc, it's not ok. Postgress doesn't like lots of connections and you will have POOL_SIZE * NUMBER_OF_ACTIVE_LAMBDA connections.
So unless you limit the number of Lambda's you have using reserved capacity Lambda's are simply not a good fit for running a PG pool locally inside the lambda for web stuff.
If you have a batch job, that runs once a day, with a single lambda, inserts a bunch of records in to postgres or similar, then a small pool in the lambda is fine.
So a lot depends on what you are doing in the lambda, how many you have, what your downstream service you are connecting to us.