AWS launches ARM-powered Lambdas
aws.amazon.com
aws.amazon.com
Lambdas are small, isolated, stateless, and highly abstract. All of these things mean, most users will probably just change one line, and immediately see a 20-40% cost reduction.
Practically speaking, there is no reason for x86 lambdas to exist anymore (outside of, "we ship a go binary compiled to x86, need to change one line in the CI to compile to ARM, ok done").
More broadly speaking; as I type this on an M1 Macbook Air, a device with no fan, capable of matching the single-core performance of the brand new i7 11700KF and Ryzen 7 5900x, talking about how new AWS ARM processors will reduce lambda costs by up-to 40% while increasing performance by 20% over x86... we need consumer ARM chips, and we need broader non-Mac desktop application support. Seems like the writing is on the wall, and developers can either keep up with the curve or be left behind.
Right now, very specifically: Looking at you Autodesk. We're coming up on a year of M1. Where's the non-Rosetta release of, say, Fusion 360?
There are lots of companies running on Graviton 2 seeing more or less that kind of cost reduction with the switch. Including Twitter. To the point AWS are currently not meeting demands for those instances.
Not every workload will benefits, but it certainly isn't a niche market.
I'm sure they don't need me to tell them they'll be inundated with demand...
My take is that a 50% discount in price/performance and a 20% discount in price per execution unit that's only a toggle flip away is something that everyone loves and is eager to jump onboard.
The only exception is when your workload stubbornly stays within free tier limits, which in that case you have no compelling reason to do the work.
There's a pull request that went up about 40 minutes prior to my comment.
So it looks like it's coming, soon.
I believe the performance boost will not be meaningful and the only important factor is cost per GB*s, given that most Lambda workloads are IO bound or simple event handlers that barely take 1s to run anyway.
Nevertheless this is exciting stuff, specially given that NodeJS might be lambda's most popular runtime and thus we might only be a toggle flip away from a 20% discount.
But NodeJS is second!
It's certainly easier to figure out what's going on in a smaller container though! I've had to debug some nasty situation with layers in Python Lambdas before and it's not fun...
Also, for performance undoubtedly Golang is by far the optimal AWS Lambda runtime.
Consequently, I see NodeJS in every single application ever as pretty much the default runtime, with a bit of Java (renowned for being by far the worst AWS Lambda runtime) primarily to leverage code reuse and of course Golang.
If anyone reading this picked Python for your AWS Lambda runtime, do you mind telling why? I would love to hear your rationale as I'm sure I'm missing an important take.
I assume the answer's roughly the same for anyone - it's odd to me that you're talking about 'ideal Lambda runtime', I'd be surprised if many people care. You're going to choose your language by however you were going to choose your language otherwise, and then run it on Lambda. And python is more popular than node in general.
If you really wanted every last gram of performance, you'd be running an optimised compiled binary anyway, not using any of the provided runtimes.
That's a great point. I admit I followed the same exact rationale motivated by code reusie to go with Java runtimes in projects that consisted mainly of peeling responsibilities out of legacy Java services, and Java is notoriously by far the worst performing Lambda runtime.
Granted, this was before lambda's pricing granularity was updated from 100ms down to 1ms. Nowadays, a JDK invocation easily costs 50x the cost of a NodeJS lambda invocation, not to mention the concurrency consequences.
> I assume the answer's roughly the same for anyone - it's odd to me that you're talking about 'ideal Lambda runtime', I'd be surprised if many people care.
Sure, there are other constrains and requirements, but in the context of AWS Lambdas, and taking into account the way they are priced and the fact that they run on a single vCPU and that they are mainly IO-bound, NodeJS is very hard to beat. If you're starting out a project, you're not totally ignorant regarding AWS Lambda's pricing, and you are free to pick whatever runtime you want, it is very hard to justify any rational choice other than NodeJS or, alternatively, Golang.
> If you really wanted every last gram of performance, you'd be running an optimised compiled binary anyway
Hence Golang.
Meanwhile, NodeJS gets you about 90% where Golang takes you without taking a single step outside of the happy path.
> Hence Golang.
Sure. Or rust, C++, C, .. whatever.
Libraries and code reuse. Python has libraries for everything, and we have written a whole host of Python libraries of our own. We know how to do what we want to do in Python and many of our lambda functions started their life in either python CLI scripts or as part of a Flask app. "Copy pasting" working code into AWS Lambda is much quicker and easier than rewriting in Node or Go.
If we're 'moving' to anything it would be moving some functions to C(++)
At @dayjob we’ve got plenty of Lambda workloads where we’re butting up against the maximum execution time limit; while that's mostly a “not optimally architected for Lambda” problem, an extra ~20% capacity before that becomes a constraint rather than a concern would be welcome.
Disclosure as I'm Co-Founder and CEO of Vantage - but we'll likely be adding Graviton recommendations to our suite of cost saving recommendations on https://www.vantage.sh/ to give per resource views of potential savings (i.e. you can save X% on this Lambda function based upon the costs we saw by switching to Graviton, etc.)
Yay free money!
So reading the tea leaves... the first generation of EC2 ARM (a1) could have been almost as simple as two nitro cards hooked together.
Graviton2 is likely a bigger chip than what they're sticking on Nitro cards, but if Graviton1 was just the overhead for each Intel machine...
Where in the world is GCE and Azure with ARM processors to compete? It's been how many years now and nothing to show yet? Customers aren't going to wait much longer...
I can't speak for Azure, but Google Cloud looks like they've gone all-in with AMD's EPYC platform for their premium option, and offers the E2 instance family for those that want cheap compute and don't care what the underlying platform is.
> Customers aren't going to wait much longer...
I can't imagine anyone outside of hobbyists and tiny startups changing their entire hosting platform due to it not having an ARM offering.
Don’t the economics work the other way though? The bigger the workload the bigger the saving and the easier it will be to justify moving. Probably a degree of flexibility in pricing to keep the biggest customers though!
I mean Amazon has been preparing for this since 2015 when they acquired Annapurna Labs.
Considering TSMC somehow increased their planned capacity for 5nm by 50%. There may be a chance you will will see that going to Microsoft of Google coming in 2023.
Oddly I can’t see lambdas being enough of a cost to justify it to hobbyist such as myself.
But this is a great sign of things to come, so much energy is consumed by data centers. Then again, I wonder how much code will randomly break, plus AWS’s dependency management is literally bundle it up locally and upload a zip.
What happens if a ARM package can’t be built locally.
That's true and not great IMO either, I also wish the zip wasn't necessary - in the 'bootstrap' binary case just let me upload the binary! - but in case you're not aware of layers (just guessing from 'dependencies') it can be improved a bit:
https://docs.aws.amazon.com/lambda/latest/dg/configuration-l...
> What happens if a ARM package can’t be built locally.
An Arm package can only not be built locally if you decline to use tools that'd allow you to do so?
I much prefer using Docker containers since they added support for this. There's an established toolchain which supports things like this and it fits well into popular deployment pipelines from things like GitLab/Github.
If your concern is more about having access to ARM hardware, CodeBuild is a reasonable option. CodePipeline is more expensive than it should be given the feature set, but CodeBuild is cheap and easy to use directly.
Alternatively, I've been intending to move some of my build to Lambda itself. Easy to bootstrap if you use an architecture independent language like Python to script your build. And if your code takes more than 15 minutes to build, should it really be a Lambda function?
My main concern if I’m using something like Numpy, and I pull my dependencies locally on my 64-bit Machine I have no idea what’s going to work once I push it up to AWS ARM.
This doesn't sound very realistic. Cost-conscious organizations would hardly consider lambdas as a reasonable option unless you're well within the free tier limits. Otherwise lambdas are far more expensive than simply handling requests directly with a service running on EC2/ECS/Fargate.
We have another product built on lambdas where the compute is cheaper than the monitoring, security tooling, etc attached to the account.
For a lot of use cases, it is cheap.
You could also run a Digital Ocean droplet for less which might make sense if your on tinkerer budget and don’t need IAM access control or VPC access.
But, the x86-64 vCPU is not a "real core". It's a hyperthread (SMT), or 1/2 a core. So I'm curious if the scale is still 2-6 vCPUs for Graviton, where a vCPU == A real core...since there's no SMT on Graviton.
siblings : 2
cpu cores : 2
Which would imply no hyperthreads, to the degree that their hypervisor is telling the truth. So Lambdas are then one of the few places where an x86 core is a core. Thanks for the info.All Intel t instances, like t2, or t3 give you two hyperthreads, but cap your CPU utilization.
Lambda runs on EC2...
I feel in this case the distinction between EC2 and EC2 bare metal is very important. Based on the GP's use of "EC2"[0] its clear they meant "the EC2 containerization model/code", rather "EC2" to mean "the EC2 hardware and/or software together".
I'll correct my original response: "While Lambda does run on EC2 hardware, Lambda uses Firecracker rather than the EC2 containerization model/code. I would be very surprised if we can correctly infer anything about Lambda's containers based on EC2's containers.
Maybe one day in the future (or already?) EC2 will also use Firecraker. At which point this comment should be deprecated.
[0] "You can look at the EC2 t instance series to see how AWS might think about constraining resources."
> Workloads using multithreading and multiprocessing, or performing many I/O operations, can experience lower execution time and, as a consequence, even lower costs.
[1]: https://aws.amazon.com/blogs/aws/aws-lambda-functions-powere...
I know there are many reasons this will be difficult in real life, but it sounds like where it should be heading.
AWS Lambda supports about half a dozen runtimes, some based on interpreted languages, and some based on compiled languages. Thus you can't simply move golang or rust code between completely different architectures and expect things to work.
Also, one of AWS Lambda's usecases is having your runtime call precompiled binaries. I'd be pissed if my Lambda's ceased to work because AWS decided to run my trusty python-calling-C++ lib Lambdas in ARM just because they want to push people towards graviton.
https://twitter.com/braincode/status/1382940634093219840
Now, the SAM-CLI local developer experience is broken in Rust, those would be my next asks for AWS:
https://github.com/umccr/s3-rust-noodles-bam/blob/s3-server/...
hopefully we'll see ML chips on Lambda soon too
For ubuntu / python-slim etc docker images - what changes are needed to let them target ARM (if any).
Apple as their M1 and M2 (coming). MS has their Sq1, Sq2 Tesla has the FSD Chip AWS Graviton, Nvidia Snapdragon, Broadcom
I am sure there are many more.
I know Apple has not publicly documented their chips. I dont know for sure about the others.
Will all of this make building compilers magnitudes more complex? If you have a codebase and you want it to optimize for M2, with all of its goodies like the GPUs and neural whatnot, then I need the same codebase to run as fast as possible on the Graviton cpus that sounds difficult?
Maybe there is a subset of Opcodes shared across all or most of them but is that reasonable to limit compilers to them?
I am sure I am just confused but it seems to me to be a bit of a mess.
Can GCC be able to keep up?
The neural processors, GPU's, TPU's etc. aren't part of the CPU anyways (they're part of the SoC).
This fragmentation has some downsides but having a good compiler isn't one of them. It's more likely all the different ways these SoC's handle booting, connecting to peripherals, etc. This will create a huge burden on something like the Linux kernel if it attempts to support them all (for evidence, check out the recent work on Linux M1 support.. fascinating stuff).
I don't think it's shocking that folks who licensed a proprietary ISA, and then paid a stiff license fee to implement their own microarch off of a proprietary reference microarch, would keep their chips proprietary.
risc-v is the only hope in this arena, and luckily David Patterson is generally signal enough that something might be a good idea.
Can't I design a new uarch that implements the ISA and is compatible like Cyrix and AMD did back in the day?
On the other hand, the fact that Arm doesn't produce chips might be an important factor, if there's an argument that the Java API is incidental to Java IP but the ARM ISA is the substance of the IP.
I think this would also depend on the extent to which the ARM ISA specifically is protected by Arm's patents, since Google v. Oracle was strictly related to copyright. It might be the case that what you're suggesting is plainly illegal, or I could see it being a kind of grey area where it's fine per se, so long as no one actually puts it into practice without purchasing a patent license from Arm.
Is the ISA itself patented, or just the implementation? Or do they claim copyright over it?
However the cost of implementing the Arm ISA in a clean room is probably more than the licensing cost from Arm - they are spreading their costs over lots of customers and you are not - so why would you do it?
So at least with Node.js, I see no benefit with the arm64 architecture. YMMV.
Publishing my docker images to my registry as ARM images from my x86 based CI pipeline is a little weird though. Would be nice if docker baked in all the QEMU magic and let me use `docker build --target-arch arm64`.
Joked about using graviton.
Now this happened. I’m pretty sure I’m Rasputin.
Was among the earliest launched PCIe 4.0 platforms.
Note to self: RTFM.
Now if only I could run a custom kernel...
Can they really spin up quick enough to actually respond in real time? I don't really want to have to keep them warm as I'll end up paying through the nose for it.
In this way you can also run whatever version of the provided language runtimes you want, not limited to the official ones if you want newer python syntax or node stdlib features than are available, for example.
And lambda is more about changing the process lifecycle from traditional always using resources. Sort of the inetd for hyperscalers.
If your container runtime uses an emulator (say QEMU) you are able to run containers which do not match the host system, albeit at a slower speed than running the container using a matching architecture and namespaces.
If you can compile the single task your lambda will perform down to as close to bare metal as possible, without affecting your workflow, why not?
I use JavaScript in the few lambdas I have because of dev. ex. What little additional cost it would add offset by speed of development for me, and how important speed is to my tasks.
But in a similar vein, extracting interesting thumbnails would at scale, be one use case where lambda with assembly/compiled static bins may be a good choice.
Additionally, a lot of interpreted languages depends on native libraries. You need to build those libraries using the right arch. It happens a lot if you want, for example, deploy a Rails app in Lambda.
By the way, Lambda works great for low latency apps. The "startup cost" is negligible if you know how to implement it. You can serve ML models from a lambda, and in that case you want a math library optimized for the platform. Unless you have a lot of traffic a Lambda usually is cheaper.
Overall, I see that as a win: you can get started by putting code in something like Python or JavaScript and not caring but when you do hit more complex dependencies you have a path which is an extension of what you've already been using rather than having to switch to a completely different service like ECS.
Not saying AWS doesn't come with a premium, but lambda isn't exactly the poster-child for expensive compute.