Serverless at Scale: Lessons from 200M Lambda Invocations
insights.adadot.com
insights.adadot.com
The only thing I know that serverless architecture promises are big bills and a steady income for a cloud provider. I'd be happy to see a serverless setup that won't be blown away with a (way cheaper) small/medium-sized VM.
I’m asking not to seem snarky, I truly want to know what is making people hit high prices with lambda. Is it like functions that are super computationally intensive and require queued lambda functions?
The main advantage, though, is predictability of operations. The FaaS services "just work". If we accidentally make a change to one endpoint to consume too much resources, it doesn't affect anything else. It's great for allowing fast changes to new functionality without much risk of breaking mature features.
Ephemerality is a plus as well. Just from a security standpoint, having an ephemeral system means persistence is not possible.
> Just from a security standpoint, having an ephemeral system means persistence is not possible.
You still have to persist something somewhere, and there is a higher chance someone will figure out SQL injection or unfiltered POST request through your app than hack SSH access to the box. If someone wants to do any real damage, they'd just continuously DDoS that serverless setup, and the cloud provider will kill the company with the bill.
Is this something people are out there believing? That patching is something that's easy to automate? I find that kind of nuts, I thought everyone understood that this is, in fact, the opposite of easily automated...
> You still have to persist something somewhere, and there is a higher chance someone will figure out SQL injection or unfiltered POST request through your app than hack SSH access to the box.
"This entirely separate attack exists therefor completely removing an entire attack primitive haves no value" - how I read this comment.
It has value, but it's also true that trusting cloud providers serveless infrastructure introduces additional sets of vulnerabilities due to various reasons.
eg: https://sysdig.com/blog/exploit-mitigate-aws-lambdas-mitre/
Reading your comments, I get the impression that you are used to dealing with clients whose infrastructure management skills are lacking, and they are making a mess of things.
While serverless infrastructures certainly eliminate a range of vulnerability classes, it is adoption is unlikely to be sufficient to secure platforms that are inadequate for the threats they face.
At the end of the day, someone has to put in the work to ensure that things are patched, safe, and secure, whether the computing model is serverless or not.
I mean, I worked at Datadog when this happened: https://www.datadoghq.com/blog/engineering/2023-03-08-deep-d...
Multi-day outage because of an apt update.
Not the only one I've seen, and it's by no means the only issue that occurs with patching (extremely common that companies don't even know if they're patched for a given vuln).
Ansible on a cron, and the pipeline goes to prod if the test environment passes.
Or unattended upgrades in test, that fires a job to prod if it passes.
Or a continuous build process with Packer to replace running instances once they pass.
If you have certain things that can’t tolerate sudden downtime (a DB, etc.) then you need to know how to mask/hold those.
All this to say, it’s easy if you already know the footguns. But IMO, if you don’t know them, you don’t really have any business running Linux boxes in prod.
It will happen to you sooner or later also. Updates are always out of band for this reason.
That’s why everybody does builds and isolates isolates updates to that process.
I've seen this happen twice now, as well.
I am on your side, actually, I think managing machines is better than serverless, but it's not that easy.
Fly charges for a VM by the second, and when a VM is off, RAM and CPU are not charged (storage is still charged). They also allow you to easily configure shutting down machines when there are no active requests (right from their config file `fly.toml`), and support a more advanced method which involves your application essentially terminating itself when there's no work remaining, which kills the VM. When a new request arrives, it starts back up.
Here are the docs [0]. And here's a blog post on how to terminate the VM from inside a Phoenix app for example [1].
So essentially, you can write an app which processes multiple requests on the same VM (so not really serverless), but also saves costs when its not in use (essentially the promise of serverless).
[0] https://fly.io/docs/apps/autostart-stop/
[1] https://fly.io/phoenix-files/shut-down-idle-phoenix-app/
My current company is running entirely on Cloud Run. Not quite as "serverless" as pure Functions or Lambdas, but we have zero VMs or hardware that we manage, so I feel like it counts. We don't do huge amounts of traffic (and don't need to), but it's not trivial either and it's very spiky. The Cloud Run part of our setup is almost negligable (dominated by the database and storage/network costs by several orders of magnitude). With that we get easy deploys and rollbacks, auto-scaling, ephemeral preview environments for every PR, and a simple security story (when the security questionnaire spreadsheets come around, I get to just say "not applicable" and skip entire sections on host-based security, SSH keys, OS updates, etc). And it's basically just a standard Docker image for the app, so if we ever felt like it would be more cost effective to run it on a VM or K8s cluster, it wouldn't be that difficult.
I agree that not everything is better with serverless, but there are some things where it's just a vastly better fit.
It probably doesn't make sense economically, due to the cost of managing the infra vs the cost of more top-level VMs from a provider, but you can certainly segregate prod, staging, and independent spaces for multiple developers (or yourself on different projects) on a single rented VM if you want to.
Why make it sound so sensational? I did much more than that on a single xeon machine.
If a lambda is called 6 times per second, I suspect the underlying VMs/containers that power lambdas are rarely shut down (I don't know how AWS works but that's how it works in another cloud provider I'm familiar with - they wait for a little for new requests before shutting down the container). So might've as well just used an always-on server.
I also wonder why their calculations show that 6rps (that's what their "17 mln monthly lambda invocations" really means) would require 25 servers. We have a single mediocre VM which serves around 6 rps on average as well without issues... Although, of course, it all depends on what kind of load each request has. We don't do number crunching and most of the time is spent in the database.
I'm guessing that it has something to do with the average job taking 15 minutes. 6rps represents 6 jobs being created per second, but each one takes 15 minutes to run until completion. Another way to look at it is each second 90 minutes of lambda work is created.
If you consider 15 minutes the rolling window, which looks like a fair assumption based on the graph provided, there could be up to 5400 (15 min x 60 sec x 6 rps) functions running at once. Working backwards, 25 medium instance provide (25 instances x 4 GB mem) an 100gb memory pool, or 100,000mb. That leaves around ~18mb for each of those 5400 jobs, if you don't consider OS resource overhead.
Looking at averages in this situation very possibly can give a warped perception of reality, but 25 instances doesn't seem out of the realm of possibility. I'm sure they have much more relevant metrics to back up that number as well.
Whether the functions really need this much time to run is another issue entirely, and hard to answer with the information given.
IMO lambdas are most practical when someone else in the org is responsible for spinning up VMs and your boss is in a pissing match with their boss, preventing you from getting any new infrastructure to get work done. The technical merits rarely have anything to do with it.
Your IT org political problems may not match, and therefore their architecture may not either. Also there's the question of what problem you are trying to scale at what scale, with what resiliency/uptime requirements, etc...
And yet I've been at companies 80% smaller/poorer who didn't notice $60k development environment costs because they were blended into a gigantic AWS bill.
The incentives are very very bad.
But our outcome is: Outside of really heavyweight processes like model trainings, it's a lot of infrastructural effort to run something like this on our own systems, opposed to just sticking that code into a rarely called REST endpoint in some application we're already running anyway. We'd need a lot more volume of somewhat rarely executed tasks to make it worth running it.
You have to use lambda so you can overcome artificial engineering constraints so that you can write blog posts.
Per day. Per host. In 2009. On python.
How is everyone making everything so slow?
TL;DR your Xeon box is their always-on api box.
From what i can see its basically recreating mainframe batch processing but in the cloud. X happens which triggers Y Job which triggers Z job and so on.
The lesson I've learned instead is to start boring and traditional, then use serverless tech when you hit a problem the current setup cannot solve
Creating one repo per Lambda is going to make things messy of course, just as breaking every little internal library out into its own repo.
Regardless of the system or what it runs on, it's an easy trap to fall into but it's absolutely solvable with some technical leadership around standards.
I wish that was the case in real life. Unfortunately, the trend I've been noticing is to run anything that's an API call in Lambda, and then chaining multiple Lambdas in order to process whatever is needed for that API call.
For our region that limit is set to 1000. That might sound like a lot when you start, but you quickly realise it’s easy to reach once you have enough Lambdas and you scale up. We found ourselves hitting that limit a lot once our traffic and therefore our demands from our system started scaling.
You can file a support ticket to have that limit raised.https://docs.aws.amazon.com/servicequotas/latest/userguide/r...
The textbook example of this going wrong is a lambda that is invoked on uploading to S3 that writes the result to S3. There's even an AWS article on it - [0]
[0] https://aws.amazon.com/blogs/compute/avoiding-recursive-invo...
Specifically as it relates to Lambdas there's solid rationale behind these limits, but I agree that in many other cases the limits seem arbitrary and annoying.
- for limited resources like IPs, it avoids one customer eating all the stock. Yes he’s paying for them, but other customers wouldn’t be able to get some anymore, generating frustrated users and revenue loss - for most other "infinite stock" resources, it avoids the bill exploding. It’s good for the customer, but also for the provider as they’re sure to be paid and not take a billing decline or sucking up all of a startup’s money.
In addition to lambdas being a poor architectural choice in most cases, that is.
This should be called Fuction as a service. Acronym Faas.
Serverless is kinda confusing though. Like whenever I say I work with serverless functions I always have to explain there are unmanaged servers involved
This sounds way off to me. 10 GB to install a Python library?
Here is an unoptimized example, built on an M1 Mac:
$ cat <<EOF > Dockerfile
FROM python:3.12-slim-bookworm
RUN apt-get update && apt-get install -y python3-pip
RUN pip install pandas
LABEL "name"="python-pandas"
ENTRYPOINT ["python3"]
EOF
$ docker image ls -f 'label'='name'='python-pandas'
REPOSITORY TAG IMAGE ID CREATED SIZE
sgarland/python-pandas latest 89e31f6eb83d 9 minutes ago 764MB
A more optimized version: $ cat <<EOF > Dockerfile
FROM python:3.12-slim-bookworm
RUN apt-get update && \
apt-get install -y --no-install-recommends python3-pip && \
pip install pandas && \
apt-get purge -y --autoremove python3-pip && \
rm -rf /var/lib/apt/lists/*
LABEL "name"="python-pandas"
ENTRYPOINT ["python3"]
$ docker image ls -f 'label'='name'='python-pandas'
REPOSITORY TAG IMAGE ID CREATED SIZE
sgarland/python-pandas smaller 102308842b88 4 seconds ago 342MB
sgarland/python-pandas latest 89e31f6eb83d 27 minutes ago 764MB
Even adding in scipy didn't crack 500 MB: $ docker image ls -f 'label'='name'='python-pandas'
REPOSITORY TAG IMAGE ID CREATED SIZE
sgarland/python-pandas scipy 808535284f03 3 minutes ago 497MB
sgarland/python-pandas smaller 102308842b88 9 minutes ago 342MB
sgarland/python-pandas latest 89e31f6eb83d 36 minutes ago 764MB
I'm not sure how they managed 10 GB. Here's the non-slim version, with no optimizations (this is much larger because `python3-pip` has the system default Python interpreter as as dependency, so this installs Python3.11 into the image): $ cat <<EOF > Dockerfile
FROM python:3.12-bookworm
RUN apt-get update
RUN apt-get install -y python3-pip
RUN pip install pandas scipy
LABEL "name"="python-pandas"
ENTRYPOINT ["python3"]
EOF
$ docker image ls -f 'label'='name'='python-pandas'
REPOSITORY TAG IMAGE ID CREATED SIZE
sgarland/python-pandas bigger f8f98e9a241c 8 seconds ago 1.44GB
sgarland/python-pandas scipy 808535284f03 3 minutes ago 497MB
sgarland/python-pandas smaller 102308842b88 9 minutes ago 342MB
sgarland/python-pandas latest 89e31f6eb83d 36 minutes ago 764MBThis one speaks truth.