AWS Lambda function URLs: Built-in HTTPS endpoints
aws.amazon.com
aws.amazon.com
Also an interesting note from the docs about how said URL is generated: "Because this process is deterministic, it may be possible for anyone to retrieve your account ID from the <url-id>." I don't know how much of a problem this could be, but it's worth being aware of.
/lambda/* could be routed to your functions
/* everything else your main app
Lambda has scale to zero and fast startup but its custom RPC interface (presumably an outgrowth of its batch processing origins) does not support streaming responses, has awkward response size limits, and prevents multiple requests from being executed concurrently on the same instance (so caches cannot be shared.)
Fargate provides the flexibility from simply running an HTTP server inside a container but at the cost of slower startup and no ability to scale to zero.
https://docs.aws.amazon.com/autoscaling/application/userguid...
It scales almost to zero (minimum cost is the memory dedicated to a single task).
Somewhat frustrating that the ephemeral storage mentioned in the documentation provides no clue as to the amount available. https://github.com/aws/apprunner-roadmap/issues/112
There are so many AWS services these days I have no idea what to use half the time…
It's implementing KNative API but on GCP Infra
While there are some nice benefits to serverless workloads on AWS, local development and reproing production bugs are major weak points.
https://aws.amazon.com/blogs/compute/accelerating-serverless...
Definitely cheaper than API gateways for even a moderate amount of traffic, but API gateway costs scale down to zero for unused or rarely used endpoints.
Note: The “official” way to work around this is to write your large payload to S3 and then create a pre-signed S3 URL that you return to the caller instead.
ALB with Lambda - 1MB
Lambda - 6MB
https://docs.aws.amazon.com/elasticloadbalancing/latest/appl...
Lambda doesn't charge for idle time or capacity, for example. Depending on your usage patterns and availability requirements, you may have a high cost in idle waste when using servers.
But I think where Lambda really shines is removing time from infra setup, maintenance, etc. These are hidden costs and, even when accounted for, frequently underestimated.
If you need even a part time DevOps or SRE engineer to make sure you'll keep things running smooth or will recover fast from a disaster, their salary alone already pays for shitloads of Lambda invocation and compute time.
You still need a way to test and deploy the code your pushing. You also need to set up logging and networking that ties it into the rest of your systems
We don't have an operations team, so each dev deploys their own services; this framework gives them all a consistent interface so we know how to test or deploy someone else's code.
Advantages from my POV:
* zero or simplified ops; no load balancer setup etc.
* we don't pay for idle hardware
* scales well by default (built with parallelism in mind)
If a certain game spikes in popularity then it COULD become expensive - but IMO that's better default behavior than a server falling over.Would you be interested in sharing your Serverless Story with our Serverless community? We are looking for guest blog contributors, community call showcases and case-studies.
Please fill out the form in the link below and I will follow up with you! https://formfacade.com/public/112252076977537904120/all/form...
Cheers!
You mention that serverless.com offers a sort of PaaS that simplifies your life. That's similar to Heroku, which does the same thing but with regular servers and similarly allows you to defer having a dedicated Ops person. Having worked at plenty of places that leveraged PaaS's for most of their needs, I'd say it's normal to not have an Ops team until you begin hitting issues around general maintenance and secure networking setups.
I was the fourth engineer at a smallish currebtly Series A startup that used lambdas for ~90% of their services and we ( the software devs) built the CICD, API gateway integrations, all of that stuff. We figured out a good solution early and then maintained it. Which made moving forward pretty trivial.
It only got a little bit more complicated as we wanted to use different authorization methods or VPC private link when we did eventually spin up a couple ECS clusters.
The company now has a full-time DevOps/SRE, but they don't really work on the code deploy CI stuff. They more deal with I am audits and security and stuff...
I'd argue that the need for someone focused on SRE/Devops comes from factors that are mostly unrelated to whether or not you go serverless.
huge understatement even as written :)
I'm not sure I agree, why Lambda and not AWS spot instances? Or heaven forbid, one beefy server / office warmer.
Lambda doesn't win on price it wins on flexibility.
But then, I have been running my SaaS on a single, rented beefy bare metal server in a data center with unmetered data and have had no more than 15 minutes of downtime per year for 3 years now.
Also nothing prevents you from using multi zoned auto scaling group with a min/max of 1, you'll fail over to another AZ gracefully. You don't magically get 9 nines with lambda in practice either, trust me.
There are more legit use cases. Lets say you are in charge of Eurovision.
You spike from 0 to xxx,xxx of votes on an irregular basis during a few nights a year.
You run an auction or ticket sales website of some sort.
Your blog gets 3 clicks a month except for when you actually post something and it briefly gets noticed by HN.
Source/Disclosure: I am Co-Founder and CEO at https://www.vantage.sh/ - a cloud cost platform tracking in the hundreds of millions of dollars of annualized cloud spend.
However, I'm afraid that I sadly would not touch your service, because SSO is only on your enterprise plan. I hope that eventually you might consider putting SSO on one of the priced tiers.
(https://sso.tax/ approximates my opinion)
I've mentioned this publicly before on HN: we are happy to offer anyone SAML SSO even if you fit into our Pro/Business tiers. We have configured "Enterprise" tiers for folks below $200/month and it is no sweat on our end. It just takes some pre-provisioning on our end that usually comes with more direct engagements with Enterprise customers.
Anecdotally, typically _very_ few customers in the self-service tiers make mention of this. It's something we need to message better on our marketing site but for onlookers here, know the option is available to you.
That constrains how you build architectures with lambda. Anything bloated or written in low efficiency languages, or does any amount of compute or latent network traffic will burn you. That includes dealing with slow clients and synchronous APIs.
As for rules, there aren’t really any catch-all ones. We build out on Kubernetes now though as that’s portable. The lambda developer story is quite horrible and that’s where our cost is.
Then again that’s how Bezos rolls.
With Rust on Lambda we found a good mix of start time and parallelism of each that made it much cheaper and faster.
EDIT: The workload times are only somewhat predictable so even doing something like scheduled scaling is a half-solution.
This may come off as arrogant, but it sounds like something that could be rewritten in Fortran to run on an old laptop with similar execution time.
We can definitely get by running the sampling on a dozen larger systems, but those systems still cost more than (one thousand Lambda executions * a few hundred times a week max).
Pre-computing is an interesting idea, but would certainly be an epic amount of work. The input set is also quite large (compiled motorsports data based on current state of an event).
I'm also totally willing to admit I'm probably in over my head in designing such a system, but I did try a few different ways of architecting this and tweaking the parallelization in the source and this was the best cross of time/money I could get! It's definitely been a learning process for me; although I've got some years experience with scaling out, the system and requirements were much different than I've been exposed to. I'm sure this first take will be changed or at least refined once it's been in use for awhile.
They report 3.3B function invocations for 2.3B requests served, a total of 61,000 hours of execution time and 1,500 concurrent executions at peak.
Using calculator.aws, simply entering the number of requests at 67ms each (61k hours/3.3B invocations) with 128 MB of memory each computes a cost of only $1,107.
Then of course there's bandwidth, databases, file storage, and a lot more… but 3.3B lambda executions for $1.1k is really cheap. Not to mention that a client like the BBC would benefit from negotiated pricing with significant discounts. Even without discounts, serverless is often pretty cheap.
Run it on Linode or something, they'll throw in 10TB of bandwidth which probably covers it.
Most of the serving is done by the CDN..
Certain kinds of organizational efficiencies frequently lead to "I'm bored lets use the new hotness that would look good on my resume for the next job", hence lambda.
Actually you don't need anything at all, this is a job completely for the CDN which has enough features they aren't using to do the same thing for no additional cost.
For what's it worth, I've worked in Ad Tech at 2 m qps and now in HFT so I'm not averse to efficiency, but org efficiency is (in my experience) very important if you're scaling the company to the $100+ m revenue range.
It really depends on what it is that you're actually, you know, doing.
Lambda lets you get going quicker but comes with its own idiosyncrasies, don't get lulled into a false sense of security. You're going to end up with a devops guy either way and they'll be able to handle either scenario.
Which one you should pick really really depends. But, all things being equal, it is better to shoot yourself in the foot than having AWS shoot you in the foot.
You gave your own counter-example. NomadList has a simple app, Lambda is the wrong choice for him. BBC link tells me their thing is so trivial it is moot - whatever floats their boat.
If I was doing some complicated EMR in BigCorp with many departments and stakeholders, slinging millions of files around S3 I'd pick lambda.
If I am setting up ad servers I would not pick Lambda because it is latency sensitive. 'etc.
Get yourself a devops guy and a proper architect and game it all out.
Think of your devops guy not as a server administrator but an advance AI that accelerates the velocity of your team to new peaks of productivity. They are not only debugging their code in production at 3am but your org processes.
Given that...you can't rely on the CDN being an optimization. You could still make the personalized calls through the CDN, if you wanted and had a clear caching policy, but something like "on load, assemble a page with real time updates on the stories this person expressed interest in" isn't going to benefit from a CDN.
They have a very finite amount of new articles a day and a little box on the page with personalized headlines. All these articles, in fact all the articles on BBC of all time ever, comfortably fit into ram if you are so inclined.
Those little linodes have fast enough SSDs, try playing around and use one of countless stress loaders to get a feel.
I've served hundreds of thousands of requests a second from a single server using Varnish. A well tuned beefy VPS would do tens of thousands.
How to "personalize"?
For example Varnish has this thing: https://en.wikipedia.org/wiki/Edge_Side_Includes https://varnish-cache.org/docs/6.0/reference/vcl.html
It will happily read whatever cookie you set and stitch together a remixed personalized page for all the BBC users in the world without a performance penalty much faster than your lambda cold starts.
CDNs sometimes let you hook into Varnish directly or expose something similar to essentially do exactly this.
The aim of the game of the guy who wrote that blog post is to reinvent something from the 90s, but worse, and blog about it.
I will 100% agree that what they describe in the link can be done trivially without lambda (they describe a single use case of regionalization, which doesn't even take anything special on the part of the CDN), but the point is I can easily see a future use case they're building towards, that I don't think those standards you mention will support.
It is a news site not Facebook. The feed would only ever be a subset of all the articles in cache. It can itself be cached, there are only so many permutations. Different people will have different feeds, but there will be people with the exact same feed as you.
Crunch those recommendations however you want wherever you want and spew them unto S3 at your leisure. Have your varnish cache sit on top of that. Set your personalized subscriptions via a cookie. Varnish on each request will decide which one out of n (probably n<10000) feeds to include in your html response, in under 2ms.
This is all stateless. Have two servers in two regions. Done.
Maybe you'll reboot them every couple of years.
It is much more fun though: don't be clever. Just cache-bust on change.
edit: the BBC review is horrifying:
> The page takes around 500ms to render and be delivered to the audience. In that timeframe we invoke around 30 functions. Around 150ms is spent running React to render the content to HTML
> we aim to personalise almost every page in some way — making it relevant for every user on every request
Good luck making perf numbers with all those cold cached personalized pages
This is a mostly read only web page. Half a second to load? You're barely going to lose anyone, if you lose anyone at all. Hacker News routinely runs me ~300ms to load and has zero personalization, Facebook takes over a second and a half before anything displays, as does Youtube (on a refresh, no less, so things should reside in cache locally!). Hitting a random person's LinkedIn page (once I've passed the verification, which is a whole different issues) takes 1.2 seconds. Etc.
Now, admittedly those are including the latency on my end, but the point is, no one is so meth addled that a page loading after half a second (or even a full second!) is going to have much effect on engagement. Even the studies that have been done (that I have some major issues with) only really start measuring anything significant well after a second or two.
It only takes a 3G connection a few miles outside a city to add another 500ms to that opening request. Say we're up to a second before some readable text appears, now we'd like to know how long the user will actually spend reading the text or waiting for images to load before navigating again.
Intuitively, I think that load time/dwell time ratio probably captures what those studies talk about better than just raw numbers. 1 second between 10 second TikTok video loads would be extremely noticeable, but barely worth mention if the user instead was spending 10 minutes reading e.g. a feature length news article.
My personal BBC reading habit regularly involves clicking into an article just to catch the opening paragraph and seeing which opening image they used (they rarely use the same for the thumbnail). The average is probably not as low as 10 seconds, but it's certainly something much less than 2 minutes.
Dwell time probably isn't a great way to capture it either. My pattern is quite "flicky" but I bet there is a spectrum all the way from "reads every last word" to "literally just loves to click". I guess latency becomes increasingly important for folk further along that spectrum
My point is just that there's a lot of room before it even hits the numbers quoted by the various 'studies', and given those studies all tend to be from perspectives of advertisers and click through rates (i.e., giving people time to wait makes them go "wait, do I even care about this?", whereas this is someone explicitly looking to a news source), it hardly is the problem the parent presents
Still interesting though and glad it works for them.
- Easier to maintain so less hours spent handling things like deployment and autoscaling. Payroll is likely the company's top expense so this is not insignificant.
- If the traffic is sporadic or unpredictable you have 100% efficient resource utilization, which is very difficult with traditional servers.
I have some microservices that would cost $10 per AZ / datacenter per month (so at least three to have High Availability) that cost effectively $0 per month on lambda.
At scale, it depends on how consistent your load is. If it is highly consistent a server may make sense. But even for high traffic apps the cost can be lower or -- if it is higher -- so negligibly higher that the saved maintenance cost pays for it.
I have many projects on Lambda, but for example: I run an entire small-startup of mine for < $1 per month and have high availability. The number of hours I would need to spend to get the cost that low would be way too expensive.
But some general info:
- API based product
- Frontend using Next.js hosted on S3 behind CloudFront with some CloudFront Edge functions
- Gets a few hundred thousand API calls a month
- I wrote my non-edge Lambdas in Go. I found it to be much faster cold start times (about 100ms for me) and much faster runtime (< 10ms) than Node. The Edge functions are Node though because Edge on AWS only supports the Node.js runtime.
- I use DynamoDB for my database.
You get billed for what you use on Lambda and on average a single API call bills me for 8 - 12ms each. Plus the cost of API gateway. Actual response time is higher since SSL negotiation and API Gateway adds overhead but it's usually < 80 - 100ms which is within my SLA.
But could scale up to a few million API calls a month without much added cost... about 20 - 30 million API calls per month before I hit the cost of the (minimum) 3 servers + load balancer I would need to do High Availability using servers.
I'd argue that you should just package and deploy your Lambda as a Docker image and when you need consistency head over to Fargate. It's costs are reasonably comparable with EC2 and you get rid of most Lambda limitations.
But I'm not a fan of Fargate. I find it easier to use vanilla EC2, and it's cheaper than Fargate. But I'm also very comfortable with EC2 and Fargate takes care of a lot of the ops stuff so to each their own.
What I didn't mention too is Lambda has "Reserved Concurrency" pricing which for extremely consistent workloads lowers the gap... I've never had a product with that consistent a workload though.
"100% efficient resource utilization" is a bit over stating it, but it sure is a lot easier to ensure you don't over provision.
But the granularity is much finer with Lambda so you're closer to full efficiency. No having a server with Gigs of RAM and multi-cores sitting around for hours a day with 10% utilization.
Google cloud-run can handle multiple requests at a time, but still suspends the instance while no requests are being processed, and is billed to the nearest 100ms
To flip (because I'm generally pro lambda) the one call per instance also encourages global state (since you don't need to worry about two calls running in parallel using the same memory), which is pretty bad coding practice.
I'm not sure why but I found with pre-compiled languages like Go it's not as big an issue (as long as your app is I/O bound). With Node.js I've found that increasing the size of the Lambda helps even with I/O bound functions. I assumed it was because the JIT compiling of the JS takes CPU but it also seems slower on subsequent runs and Lambda sleeps the apps it doesn't stop them (until you hit the end of the 5 minute reuse window). I give my Node.js a min of 1 Gig of RAM whereas I've been able to get some of my Go functions down to 128MB with no performance hit.
Which yes, means the Node.js functions are about twice as expensive per MS before you even consider that the Go function runs for less time. But in both cases I've still found it cheaper than servers due to efficiency even though strictly speaking it costs more than a server per GB/Ghz.
However, it appears to not only lower the amount of cores available, but the performance of those cores.
So even if your code is very single threaded, and has low memory requirements, you still might want to provision a larger lambda memory size.
The frustrating thing with this, is that your single threaded code might only be using one of up to 6 CPU cores made available to it.
Kinda. Amazon is the one getting 100% efficient resource utilization (or close to it). You will be billed based on what's on their rate chart, not utilization.
We have a large dev team (~12-16 devs) working on the project and our dev account costs are only $100-250/mo. We even deploy each PR to the cloud to run automated tests against. Our production costs have been very manageable too.
Lambda/serverless has been great for us. Our organization is relatively new to building SaaS cloud apps and doesn’t have the most mature devops practices (we’re growing there). Building this app on Lambda and DynamoDB and letting AWS help us with most of the scaling has really been a win for the team.
Beyond that, depends on your use case. I love using Lambda for webhooks that may be called anywhere from a few times a day to thousands of times per day. Once they get to the point of being "oh wow, this one lambda is expensive", you can clearly afford the budget to move it to a real server. But below that $20/mo mark (which is more than 5 million invocations of a small function) - you're golden.
Something seems off compared to your numbers.
Maybe the length of a single execution?
If your lambda bill is $4k a month, either your functions run multiple seconds each or you're using multiple gigs of memory.
You still can of course but it's one less thing you have to do.
Like another comment said you can expose it through IAM/SDK but then you're getting random permissions credentials whatever out to the world as well.
I'd love to use this with a custom domain, without having to use an API Gateway.
If not -> get prepared for huge bill of ddosed function invocations.
I hope we can at least attach something to those urls.
If you want ddos protection and proper rate limiting amazon will happily charge you for it several different ways. Maybe you can hide these urls behind cloudflare if you're penny pinching..
For all production workloads we typically have cloudfront in front of everything.
But api gw is 10mb, not for lambda tho, confuse: https://docs.aws.amazon.com/apigateway/latest/developerguide...
Use s3 object lambda and have no limit to push to s3 with logic, adds a bit of latency tho https://aws.amazon.com/blogs/aws/introducing-amazon-s3-objec...
All historical but hard to understand the limits as a poor user..
Ended up going with cloudflare workers for s3 uploads with extra logic just to avoid the unknown, workers are great btw.
If you want the developer experience of Django with the benefits of Serverless Compute platforms check it out!
I have a function triggered by Cron once a day that goes wrong about once a month. I trigger it again using the debugging tools, but it would be nice if I could just hit a URL to trigger it again.
It's a nice way to do cron/scheduled tasks without any extra work. Just deploy another serverless endpoint and have something else hit that on a schedule. If anything goes wrong, EasyCron notifies me and I just hit the URL directly to re-run the task. Simple and crude, but highly effective and near zero effort.
[0] https://docs.microsoft.com/en-us/azure/azure-functions/funct...
If you can find an example for doing otherwise I'd be delighted!
> You can mix and match different bindings to suit your needs. Bindings are optional and a function might have one or multiple input and/or output bindings.[0]
However, there is some documentation explaining how to execute a function that does not have a HTTP trigger via HTTP[1]. The example uses the function app's master key though, it'd be interesting to see if that's a requirement or if you could use a key scoped only for invocation of the specific function.
[0]https://docs.microsoft.com/en-us/azure/azure-functions/funct...
[1]https://docs.microsoft.com/en-us/azure/azure-functions/funct...
I wonder if it's worth changing my current API Gateway endpoints to the built in Lambda URL's, since I haven't launched yet.
Seems like AWS is actually launching new endpoints with IPv6 support by default now.
Also, I'm not sure what you're referring to, having read it twice.
I don't have advanced settings. Instead I have to go "Configuration->Function URL" to find this.
Every time I peek over into the AWS ecosystem I’m very glad I don’t have to work in it. This seems like it’s multiple years behind what GCP has unless I’m missing something obvious?
Workers has 1MB script size limit (post compression), so that's there, too, and can run WASM or JS workloads (which CloudFront Functions can't, but Lambda@Edge can).
As for AWS Lambda Function URLs: Well, it isn't comparable to Workers at all. But if my use case fits Workers, then that's what I'd would prefer. In fact, I've gone many lengths to make my workload fit Workers. Deno Deploy is another viable alternative.
[0] https://docs.aws.amazon.com/AmazonCloudFront/latest/Develope...
[1] dated, but relevant: https://medium.com/@zackbloom/serverless-pricing-and-costs-a...
If you need more than 1MB of script size, please reach out to Cloudflare support.
I thought our Unbound Workers are supposed to also be cheaper as well but I need to double-check that piece.
Bundled and Unbound Workers are equally fast.
For most workloads, I'd reckon that Unbound Workers are about the same cost as Bundled. In fact, Unbound will be ~2x cheaper than Bundled if your average workload completes within 50ms IO or 10ms CPU.
As for cost, a Bundled Worker definitely had a price advantage if you have CPU-light but IO-heavy workload. If my math is right then Unbound is cheaper up to roughly 220 ms of wall clock (I used 100M requests as an example). So if it takes > 220ms of time to send the response fully, Unbound will be the same price as Bundled and only get more expensive the longer the response takes. This isn't the RTT time to your origin. It's the total request time. So if you're doing lots of round-trips to origins, proxying WebSocket messages back and forth over a long time, proxying a large response body from somewhere else etc. This gets more complicated since we put in an important optimization that makes Unbound much cheaper if you're just proxying a response without modifying it since billing will stop once you return the Response so now the Unbound Worker has to actually be meaningfully involved in generating the Response body for it to bill until the response finishes sending to your client [2].
[1] https://stackoverflow.com/questions/68720436/what-is-cpu-tim... [2] https://blog.cloudflare.com/workers-optimization-reduces-you...
* If you use waitUntil() to schedule async work that completes after the HTTP response has been sent, this work is limited to 30 seconds.
* In general, if a request runs longer than 30 seconds, the chance of random cancellation increases a lot. For example, when we upgrade the Workers Runtime to a new version, we will give in-flight requests 30 seconds to finish before the process exits, which will cancel all remaining requests. (Of course, any application that relies on long-running connections needs to handle random disconnects regardless, due to the general unreliability of networks.)
I've been using Workers since 2019 and quite haven't kept up with Cloudflare's pace of innovation ever since. It has been dizzying. Looking forward to handling TCP and WebRTC workloads (announced last year) with Workers next.
So TLDR Workers is faster, can run longer, and costs less than Lambda@Edge.