The Unfulfilled Promise of Serverless
lastweekinaws.com
lastweekinaws.com
Serverless is purely about CAPEX, vs OPEX.
Companies are so loath to make capital expenditures (CAPEX) that they will willingly let their employees waste thousands of extra hours learning a new development model so they can pay Amazon using operating expenses instead.
And it's funny because even though amazon has no choice but to use capital expenditures to populate their datacentres with machines, they are offering lambda as a way to monetize the unused cycles.
So lambdas are being twice used as a kind of compromise to fill in the gaps on an accounting sheet.
I believe the joke is on us as technical people. We twist ourselves into knots to promote this new development model as a technical innovation, which it really isn't.
We have just made the time-sharing service (they used to have them in the 60s and 70s) fashionable again. (Along with the long feedback cycle which makes working on these systems so frustrating.)
(Cheaper that is if you don't value your employees time, which most large companies don't.)
Once you have enough load to justify going “unserverless” the prices drop to approximately 1/4 that of Cloud Functions for the same burst performance, and that’s before playing billing games like long-term commitments or using preemptible instances.
What’s truly scalable about “serverless compute” versus VMs is the line item on your bill. Sure, “they manage the auto-scaling” but for the per-unit price you rapidly hit a point where you might as well set up a Kubernetes load balancer and eat the setup costs. The pricing model only works out in your favor if you _don’t_ have load.
It would not surprise me to learn that the same is true of AWS prices.
* monitor lots of things like filesystem usage and CPU usage
* build new machine images (as well as possibly new container images!) to keep up with security updates
* tune autoscaling at multiple levels based on whichever server-based systems you're using
* maintain a secure way for production support folks to log into servers to see what is going wrong or perform emergency fixes
* build additional failover automation as well as what you'd already need for serverless
Serverless is a bit more expensive overall and less tunable, but there's just less to go wrong: much of the underlying server SRE work is handled by provider automation and engineers for whom that is their full-time job. So, I'm still pretty happy with it.Not only that but the joke is on technical people for promoting a model of severe vendor lock-in. Its doubly funny when you often see the same techies complain bitterly about "walled gardens" on their smartphones !
If $cloud_vendor doubled their prices tomorrow, what are you going to do ?
If $cloud_vendor decided to kill-off a product tomorrow and replace it with another one with a different API, what are you going to do ?
In most cases, the answer to the above and other scenarios would be "suck it up and swallow the expense".
Sure you can shift the expenses from CAPEX to OPEX and construct numerous business cases to convince the boss that the cloud is the next best thing to sliced bread. But at what cost ?
Of course there are some use-cases which are genuinely well suited to the cloud model, but they are in the minority. For many business the cloud is a case of "if the only thing you have is a hammer, everything looks like a nail".
I am most familiar with AWS, so all I can say is their prices have gone down over times, the availability has gone up, and the total number of services is amazing.
With other cloud vendors I think you might be at more risk of them shutting down services, but if they start doing that with out giving you multi-year notice it means they are going out of business.
No. Just...no. There are more than the two possible positions, its not either "ignore technological dependencies completely and blindly give your balls to a single vendor" or "create everything yourself from the ground up". You can look at your stack and consider how vulnerable you are compared to how much you gain. Sometimes you will find that it's worth it going all in with one vendor, but it should be a conscious decision. Business have to consider their vulnerabilities all the time and it boils down to compromises and tradeoffs.
> If $cloud_vendor decided to kill-off a product tomorrow and replace it with another one with a different API, what are you going to do ?
Been using AWS for close to ~10 years now and I don't recall them doing anything remotely similar to above. Of course, they could potentially do it in the coming days. But going by their past behaviour I'd say it's a low probability event.
When the cloud/AWS was new these fears were indeed valid. But this line of reasoning is getting old and tiring; especially when it's not backed by even anecdotal evidence and is purely speculative-fear in nature.
I do know that Google, Twitter and others have pulled down some fairly popular APIs/products but please don't conflate that with GCP/AWS/cloud offerings.
For most of the startups it's not even a question of if cloud now. Early days (2010-2015ish) it used to be the case that engineers had to convince management about the advantages of cloud. Primarily because one had to mostly go hybrid as many of the services weren't in cloud yet (such as Route 53, RDS etc.,). It was a pain to maintain hybrid. Now it's possible to run 100% of tech stack in cloud. So it's the other way round; VCs and business folks will frown upon tech if they run on bare-metal/data-centre.
For a typical startup trying to figure out PMF and launch an MVP or figure out customer retention I'd be really surprised if off-cloud contingency is even on top-50 priority list.
But for someone with a proven business model with a good revenue stream cloud dependency or vendor lock-in or single-vendor-point of failure does indeed become an action item to be worked upon.
Or when a regulator requires setting up a DR site which they did for us back in 2015. We had no option but to set it up in a data-centre.
Or let's say someone like Mighty (1) whose product is sort of competitor to AWS's (2). They wouldn't want to go anywhere near AWS or any cloud for that matter. And last I heard they were indeed ordering their own physical server machines.
There are valid reasons to look for either hybrid or purely on-prem solutions as I listed a few above. But please don't forward "what if cost doubles" type of reasoning. All that said, I'd hypothesise not going full on cloud is an exception for a typical startup now a days.
(1) https://www.mightyapp.com (2) https://aws.amazon.com/workspaces/
Yes, it's a simpler change. But other serverless tech does exists on mostly the same dev language.
Additionally, price changes should be warned way upfront.
They can't do this because of reputational damage. No one would consider that provider for the next 5 years. It also wouldn't squeeze out as much as you'd think because medium size and larger accounts are all on long-term contracts, so they have time to migrate.
"Engineers buy, managers oversee" is the killer app.
Quite the opposite? We can be paid for migrating current applications to "serverless", and when the tides of tech fashion change we can be paid again to migrate to the new fashionable tech.
If the joke is on anyone, it is on the shareholders of the companies getting locked in. But, it is their money and if they want to trade some CAPEX for OPEX that is not really my problem.
What if we made it your problem by granting you equity in the business? I feel like this is the #1 reason to funnel a portion of shares to your employees. Making the technical people give even 1% of a shit about the cost of doing business is infinitely better than 0%.
The #1 benefit of options for startups is that they shift engineering costs from now to the future, so that you can hire any engineers at all while operating on a small budget. If the business works out, you already have so much money that the cost of the options will be negligible. If the business doesn't work out, the cost of the options was zero.
A popular vesting schedule is 25/25/25/25, which means you would be able to sell 25% one year after getting the shares, another 25% after the second year and so on. Typically they would keep giving you more shares as your shares vest so that you always have some shares that you can't sell yet.
But leasing servers and getting hands on with the metal and containers are also OPEX.
Accounting doesn't care about the technical architecture. They just don't want a data center of idle servers running Pentium's from 2010.
Instead I think the promise is, "Don't worry your pretty little heads about what code is running where. Trust us to handle that!" Which surely seems seductive until people realized that they still have to think about those things.
I heard about one project where consultants decided to use the sparkletastic magic of AWS Lambda so they didn't have to worry about how to scale up their crawler. But they didn't really think it through, so they just had a zillion invocations sitting around waiting for HTTP responses. First-month bill was $12k when you could have spent ~$100 to get the same results with a low-end virtual server running a basic Scrapy setup. So for Amazon I think it's less about filling in a utilization gap and more about attracting people who don't know how to optimize.
Sad.
FTFY!
Ed: fixed autocorrect error
What it's really about is managing cash flow and not spending your entire round until your team has some inkling as to what it costs for the revenues it's bringing in. This is considerably easier to do with cloud vendors, serverless or not.
It's also about operational velocity, and supporting engineering teams in getting things done without planning overhead that's almost always going to be wrong. Serverless is a tool, and if used right can significantly reduce costs. I can use a hammer to build a house, but I can also smash my fingers with it. Tools be be used improperly.
So once the company is mature, such that meaningful future predictions can be made from past data, then it finds itself in a position to make significant upfront outlays in response to current or future business needs. This usually involves the CFO & FP&A team building a model to see how such a change would impact the balance sheet over the next X years. If the numbers add up, it would absolutely make sense to spend the money. But that's not always the case, even for companies at scale.
1. We have a bunch of services running on App Engine NodeJS Flexible version. We have extremely minimal lock-in because we basically just have a normal Express app serving GraphQL.
2. In fact, we migrated some of our services to Cloud Run for lower costs. Cloud Run is basically serverless Docker containers, and the migration was very easy. Again, our apps are for the most part platform agnostic apps on Node.
3. We also make use of Google Cloud Functions for our asynchronous event handling. This has also been a great choice.
Going with serverless tech has easily saved us ~2 FTEs in a team of fewer than 10 engineers.
I guess at the end of the day it's just JS / Python code, but is there an "export to / import from a common format" API / web portal button?
* Cloud Run is just running a Docker image. It's basically as platform agnostic as Docker is.
* Similarly, App Engine Flexible is just running a plain version of NodeJS. No specific lockin if you don't access GCP-specific services.
* With respect to Google Cloud Functions, again we're just running a simple Node service in response to some PubSub events. There would be work to migrate this, for example, to AWS Lambda, but it would be a very straightforward mapping.
You can, for example, containerize your monolith and expose task triggers via API endpoints so your web server doesn't have to do that work. I really like it.
We similarly run a start-up on AWS serverless, and there are zero regrets.
The portion of code that actually interacts with the Lambda-specific JSON requests is trivial, and just routes to code that is unaware of environment.
As for lock in: we habitually put abstractions between things like key-value stores, messaging and queuing services and the business logic -- so moving this stuff to, say, fargate/dynamo/sns/sqs would not be a lot of work -- probably a few weeks to rework the terraform scripts, figure out a good approach to roles, and thats it.
For our app, I'm in the process of removing API Gateway and calling lambdas directly from our back end servers instead, as it added a lot of complexity and the 30 seconds timeout imposed by API Gateway is causing issues.
1. https://www.reddit.com/r/aws/comments/qxrubv/lambda_function...
Instead, I've used ALB -> Lambda (1) to good effect; much simpler to setup and no maintenance issues so far.
(1) https://aws.amazon.com/blogs/networking-and-content-delivery...
Just like with microservices, the benefits come from not using it as a giant hammer for everything, but in isolated use cases where it's actually a good fit. These are few and far between for serverless and the first challenge is making the right decision of when to use it.
I do agree about vendor lock-in, but at the end of the day this is inescapable because of its very nature. It was always meant to be "run this code on someone else's infrastructure without me having to worry about maintaining it", so vendors are free to choose how that actually happens.
What we need to fix this are standards that can be adopted by all vendors, so that it's easier to migrate. Something like OCI for serverless that isn't tied to specific runtimes would be ideal.
I was pretty excited by Lambda initially, but I gave up on it for some of the reasons the author talks about. Google Cloud Run actually achieves what I want from serverless: I never have to think about the number of servers and can pay for less than one instance during low-traffic periods, but I could migrate my app elsewhere easily if I wanted.
Figuring out how to finagle various proprietary services and wrangle them together can take a ton of time and effort.
The crazy thing to me is that this is how things run at par. As the article points out, the downside risk with serverless is massive compared to using more traditional tools. If AWS jacks up its prices, deprecates services we rely on, or has a major issue with my company for any reason -- these could easily become existential events for us.
* being able to scale on a per-request basis and understand your burst usage of resources is pretty useful
* the top-level comment regarding CAPEX vs OPEX was spot on. Our higher-ups always had their eyes on the AWS bill and not having to pay for instances (either spot or elastic) appeased them.
* at the time at least with Java-based lambdas, the local tooling was really clunky and slow with that nasty startup. You might have better success with node or python in that area.
* otherwise, like any other tech, there's gonna be tradeoffs. Something good to keep in your toolbox, but like most things they're no silver bullet.
Running a server or a cluster of servers is elementary, kids. Do not believe the cloud hype. There is no need to spend the cloud dollars when a small fraction of ones' cloud expense gets one exponentially more compute. All you need to do is get over the fear that running a server is difficult. it is not.
Talking about a cluster being elementary is frankly absurd.
I think you might be confusing “got it working!” with “hardened in production.“
Yeah, I've found that when we recently tried serverless, we ended up wasting tons of time bikeshedding multiple CI/CD pipelines.
I completely disagree with this point, as it goes totally against the experience that I and my team have piled up after a couple of years of running a couple of serverless apps that rely heavily on a few AWS Lambdas. The API Gateway and ALB code was pretty much a one-and-done, with the bulk of the work consisting of setting up TLS termination, and the bulk of the work was on writing and testing all the business logic.
The only exception I've seen to this rule is if your serverless apps consists of a bunch of event handler from a large bunch of AWS services, like S3 triggers and message queues, that are not much than one-liners that don't do much beyond plumbing around events. Still, I don't feel it's right to describe this sort of application around lambdas that don't do much by design.
> I’m apparently atypical here! Folks don’t like to spend an order of magnitude more to monitor a system than the system itself costs to run. (...)
I also don't believe this point is fair or reasonable. It makes no sense to complain about serverless because it can be dirt-cheap (or even free to use) but your choice of monitoring service, coupled with the way you chose to use it, ends up costing more. You pick what you use and decide how you use it, and if your personal choices lead to a price tag greater than zero then that's the outcome of your own design decisions.
This complain is particularly eggregious given that AWS CloudWatch has a free tier that's very clear and included in basic intro to AWS tutorials.
> (...) It turns out that while it’s super easy to find folks who know WordPress, you’re in trouble if both of the freelance developers who understand serverless are out sick that day — not to mention that they cost roughly as much as an anesthesiologist.
Again, this is hardly a serverless issue. You'd experience the exact same problem if you ran a Spring monolith.
Perhaps I am just jaded, but I take it for granted that writing application logic is only ever going to be a significant minority of “your time”—in larger orgs there are teams specializing on the different monitoring, platform, infra, etc taxes that are a reality of running software at scale.
Five smart hackers that know aws and python well are going to get much further on lambda-all-the-things than an org of several thousand would, probably.
Lambda is also outrageously expensive at scale; but at the end of the day I’d chalk this up mostly to “OP not very good serverless compared to other patterns, fivea good at serverless”
An outage "at scale" can also be extremely expensive, so there's no way around the costs of "scale"...
I use lambdas exclusively to process other events generated by S3 or DynamoDB that happen because my web service stored something or directly wants something to happen.
This way, I didn't have to write a scheduling system or a batch system for my software, I just use AWS's perfectly good implementations. I don't really see the advantage in specifying my API, tying myself to the AWS authentication infrastructure or trying to encapsulate small bits of my API into Lambdas although I can see why you might do that in a bigger system. For me, the scaling occurs in ECS, not in distributing lambda functions.
The better part of "serverless" is the containers running the web system, S3 and the managed service running the database. The lambda functions make the job of longer running tasks much easier to manage and they are terrific for that.
I'm trying to figure out what this means because my experience has been the opposite. I find the logs to be generally very code, AWS X-Ray gives you good visibility into the whole process. And you can use third party telemetry / tracing providers too.
Also AWS Cloudwatch is actually pretty powerful for monitoring and supports custom metrics and alarms. It's not the most powerful log system in the world but it is pretty painless to use.
Also, Lambda are just short functions written in the language of your choice (of which there are many) or in a container. What exactly about that is not portable? If I write node.js code it is trivial to run in another environment that runs node.js.
I'm no an AWS super-fan but I feel this article is a rant (and one written without knowing all the actual capabilities of AWS) and not fact based.
I personally use Lambda a lot. The operations effort is near nonexistent saving me countless hours. As I said, I actually like cloudwatch. And economically, if my microservice is unused it uses zero resources and scales up near instantly.
Edit: to add to that, it sounds like the author's real issue is with microservices not serverless.
https://observablehq.com/@endpointservices/webcode
I used to work on Google's FaaS and I intend to take those lessons for generation 2.
> you’ll spend most of your time figuring out how to mate these functions with other services from that cloud provider
That's just as true if you move the code from serverless to a VM in that cloud, and keep using all the other services.
Vendor lock-in for the serverless function itself seems to mostly be an AWS thing (Lamdba). Elsewhere, the trend is more running any old container will function as serverless.
If you choose to consume dozens of proprietary services from a regular container, that is on you.
All my serverless code always has a wrapper function to decide the request object so porting it to another cloud provider (even from Lambda) is pretty easy.
Edit: The article seems to state the issue is that things like step functions are proprietary but:
a. That's not lambda
b. Don't use step functions, no one is forcing you. There is nothing step functions does that can't he handled other ways.
Other providers started out with "just serve HTTP from a container".
The perform is quick enough for our purposes. Our TLAs are basically do it before the 30 second window times out for the API Gateway. If we need better performance or if we hit issues like payload size limitations, we could deploy our apps on EC2 with no loss of features or rework.
Hard disagree. If you define lucrative as having a successful business model around serverless tooling/architecture or if your business is simply powered with severless services. Lots of folks have been had lucrative success without question.
Not only that, serverless has allowed full-stack web dev to flourish, so Id argue it’s helped some individuals build lucrative careers as well.
As much as I enjoy Corey's tweets, he is wrong on this one use case.