The Serverless Revolution Has Stalled
infoq.com
infoq.com
I went to a webdev convention, and it ended up being a serverless hype train. Industry experts with a financial incentive to promote serverless went on stage and told me they can't debug their code, or run it on their machine. They showed me comically large system diagrams for very simple use cases, then spent an hour explaining how to do not-quite-ACID transactions. Oh yeah and you can only use 3-5 languages, and each import statement has a dollar amount tied to it.
More importantly, all those skills you develop are tied to Amazon, or some other giant. All the code you write is at their mercy. Any problem you have depends on their support.
Am I supposed to bet hundreds or thousands of man hours on that?
But don't let it encourage you to think serverless has no value, or it can't be done portably or cheaply. It has its sweet spots just like everything else.
https://aws.amazon.com/solutions/implementations/aws-perspec...
https://news.ycombinator.com/item?id=24552779
(regarding costs, this setup ends up being +$500/month, again, to draw graphs of your architecture https://twitter.com/der_rehan/status/1308242717307174912)
This field used to be inspiring, but now I see the ideia of having a server being sold as the plague and lots of negativity towards people who are good at servers. They are not seen as another human being but the "other".
Also I can't understand why one would prefer to pay that much for such complexity.
It seems unsustainable for me, not to mention the new generation being spoon fed that that is way to go makes me concerned about the future of open computing.
To make an analogy, it's like hiring a carpenter based on the number of tools they have. NO. Show me your skill with a hammer and chisel, and I'll assume you can figure out the rest.
If there is one thing that dealing with AWS reps has taught me, it's that this was 100% by customers. I swear to god, AWS doesn't do anything without customers asking for it.
If you are wondering why products are build in AWS, it's because people wanted to give them money for this. Say what you want, but this isn't something are pushing on us. This is something "we" push on them to provide.
Just because it doesn't work for your situation doesn't mean it doesn't work at all.
It's a tool in your toolbox.
The new generation is always joining one cargo cult or another, that's why competent technical leadership is important. Remember when noSQL was the best thing since sliced bread?
Serverless can be a good option if you have large and unpredictable transient loads.
Like any architectural choice, you need to consider the tradeoffs and suitability for your use case. A TODO app probably doesn't need to be built with a serverless SOA.
I have not observed enough uses of serverless I can draw conclusions from, but if it's anywhere like hadoop style scaling, 95% of the users will pay 10x-100x on every metric without ever actually deriving any benefit compared to a non-distributed reasonably defined system, and 5% will actually benefit -- but everyone would want to put it on their resume and buzzword-bingo cards.
You might want to give a link next time to avoid ambiguity.
Still think it is mostly awesome and the services are solid. But such infrastructure comes with its own caveats.
EBS now bugs me about some 'HealthCheckAuthEnabled'. I don't know what it wants, just that I have limited time to react. Cannot understand the clearly auto-translated text. Still didn't get what it wants from me when I read it in English. I banned the health check from my logs and said where it can get a response from the service I run. I hoped our friendship ended there. Maybe it did and this is some form of retaliation.
The load balancer AWS set up for me now suddenly costs money. And I thought I could just use it to grab the free TLS cert that comes with it... well, the clients pay for the additional costs anyway and I would be surprised if anyone even noticed the price spike. The traffic is minimal but I am surprised it is actually that expensive. I could probably cut the costs in half if I actually had motivation to look through all the settings...
Microsoft just cancelled some features of some online service my colleague works with because they want to market an alternative for their BI solution. Did cost him 3 weeks of work at the minimum. The stuff cannot be ported to an on-premise version.
Hosting your own server is a lot of work and isn't fun. There is still a lot of work to do if you host on something on such an infrastructure.
Somehow the amount of work you have to put into software hasn't decreased. Honestly the main argument for it is that you can give away responsibilities.
edit: Just noticed that EBS probably counts as PaaS instead of serverless. I run "stuff" (very simple functions) on Lambda too, but to think it would replace apps like the article suggested seems a bit much. It has its purpose, but nearly everything I have on Lambda is AWS cloud specific. It comes in handy for Alexa integrations for example. I think the container won this fight to be honest.
For me it's easy because i just maintain my ansible roles.
It's fun because I get to keep up with OSS and be part of an amazing community.
And I get much better hardware for the price.
The tool doesn't matter, it's the practice that matters ;)
So true.
I really want to know how many of these "I host my own stuff"-people keep all there systems up to date. Updates, security patches etc.
Running a server has been a non-issue for years for me.
Sometimes I think it is a great trick some big players are pulling off, in convincing the world to subsidize their infrastructure costs (and more than that) by using their resources because it is "simpler" or more effective or alternatives are hard, etc ...
... to where the ability to competently setup and maintain those services will vanish, in due time. (Which, of course would only help to consolidate those -already- running such systems ...)
People who don't put time into learning development should not do it.
However, it seems detached from your first point.
If one's role is a front-end developer, is it necessary that they know about back-end development? If it is outside their intended job function, why would they need to know about it, if it doesn't get in the way of performing their job? If you are a backend developer, do you need to know about how to host your own infrastructure? Handle your own networking? Chip design? Your logic could be applied to any job function. Each level of the stack benefits from the levels below it being abstracted. We all stand on the shoulders of giants, and we're all much better for it.
Overall, I do think it's better that one has a good understanding about the various components one interacts with. Having a grasp of the overall system will come in handy. A curiosity into other parts of the system is beneficial, and likely is one of many indicators of success. However, if job functions can be simplified and superfluous context removed, why should we fault those for taking advantage of that?
This issue with "lowering the bar" and being glad that a simplification has "failed" (which, is yet to be determined), reeks to me like gatekeeping. The same logic could be applied to any job role which benefits from simplification. In an extreme example, this logic could be extrapolated to support the notion that anyone who cannot build their own machine from the ground up should never work in a programming position. What height is appropriate for the "bar"?
But developing with "serverless" as base stack? Nope.
Each example that I've seen could have been replaced with a simple single process job. You should not need to go over the network because you need a process to run for more than 10 minutes. You should be able to check the status of resource creation via for loop.
An argument can be made for not needing a server, but I can easily fire off a container in fargate/gks/ecs and get similar benefit.
Futures are basically trivial to serialize, so the cost of involving a second node is as little as it could be; and node consolidation can be generalized because each node's dispatcher knows how many times it has accessed a remote process, which means that consensus on who should lead its log can be reached by simply sending it to the node that has the most use for it (with load balancing achieved by nodes simply not bidding for logs when their load factor or total service time is over some threshold, or say the 60th percentile of all the nodes, or whichever of those two is higher).
† In some cases it's faster to just catch up the log, rather than sending the aggregate.
We can't run it locally because the emulator won't run 7/10 of the code, there are memory limit errors that are totally opaque to us since we can't step through and watch it break. We are getting unexplained timing issues that were being worked on via logs in dev. The costs are insane.
If we want to say serverless is the future, the future isn't here yet. There are many tools that would need to be built to make serverless viable, and the tools that are built are immature and bad. Not to mention, you're tied to the companies support, so Google was outdated from the LTS for Node for a while. The reality is going serverless means you don't have control over what happens to your code.
Serverless (specifically AWS Lambda) plays a part in the cloud, but most of your objections are true of the cloud and/or AWS more broadly, and have little to do with, specifically, serverless. (Despite whatever the webdev convention was trying to sell you.) There's a whole world of EC2 instances which are mostly VPSes, to know about that definitively aren't serverless (though there are some really neat cattle-level features that serverless learns from)
If the computer is a bicycle for the mind, then the cloud is a car for the programmer's mind. You can definitely build your own stack with your own team that manages to avoids the cloud, but at some point it becomes more than one person can handle. As you scale, there's a points where the cloud does, and does not make sense, but the cloud gets you access to far more resources as a tiny team than would otherwise be available to you.
In the end, you are just renting well maintained server farms (well, a specific percentage of operational time of some of the servers in them). There is absolutely no appeal for large technology-based companies do this once they can (the following is the lowest scale example) afford to maintain their own servers, while potentially renting a few offsite backups in other areas of the world (again, this is just the lowest-scale architecture starting at which Serverless is always the worse option).
Existing solutions are extremely over-engineered. It can be excused with "officially planning for the use-case where maximum scalability is required", but it's almost certainly a pretense to sell "certifications", aka "explaining our own convoluted badly documented mess". What this actually means is that many SWE's who are good enough to learn to use Serverless effectively, can learn any other framework that allows building distributed systems across server nodes with equal effort. Why would I base my whole business on your vendor locked, sub-optimized dumpster when I can do the same on an infinitely scalable VM networks, that can be ported to literally any vendor who supports $5/mo VMs (or, you know, self hosted if my company is large enough)?
Vendor lock-in is not a big consideration for us. I increasingly think about cloud vendors like Operating systems. I really dont care which Linux distro we are on. Pick a vendor and run 'natively' on it to go as fast as possible.
I'm happy paying AWS for maintaining the servers and getting out of that messy business. There are other concerns about serverless around observability and managability, but vendor lockin and cost is not part of my equation
If you are interested in these kind of stuff, maybe take a look at Wardley Maps: https://medium.com/wardleymaps
At least on AWS, the "SAM" experience has been probably the worst development experience I've ever had in ~20 years of web development.
It's so slow (iteration speed) and you need to jump through a billion hoops of complexity all over the place. Even dealing with something as simple as loading environment variables for both local and "real" function invokes required way too much effort.
Note: I'm not working with this tech by choice. It's for a bit of client work. I think their use case for Serverless makes sense (calling something very infrequently that glues together a few AWS resources).
Is the experience better on other platforms?
I tried AWS, and then IBM's offering which is based on an open source (Apache OpenWhisk) project, thinking that it might be easier to work with, but that was also a pain.
I just lost interest as I was only checking it out. For something constantly marketed on the ease of not having to manage servers, it fell a long way short of "easy".
Look into Firebase functions. Drop some JS in a folder, export them from an index.js file and you have yourself some endpoints.
> exports.helloWorld = functions.https.onRequest((request, response) => { response.send("Hello from Firebase!"); });
The amount of work AWS has put in front of Lambdas confuses me. Firebase does it right. You can go from "never having written a REST endpoint" to "real code" in less than 20 minutes. New endpoints can be created as fast as you can export functions from an index.js file.
Being able to throw up a new REST endpoint in under 10 minutes with 0 config is really cool though.
And Firebase Functions are priced to work as daily drivers, they can front an entire application and not cost an insane amount of $, per single ms pricing. Lambda's are a lot more complicated.
Firebase is magic... but I never recommend it for anyone, until there's some sort of migration path.
With Cloud Firestore (our serverless DB) that’s the case as well. And Firebase Auth can be seamlessly upgraded to Google Cloud Identity Platform with a click.
However you’re right that for many Firebase products (Real-time Database, Hosting) there’s no relation to Cloud Resources.
To be fair, Firebase recently released a local development tool which alleviates the need to deploy on every change, but I haven't used it yet.
Super useful though!
It's a little incomplete, missing some of the AWS IAM automation that makes local development smooth, environment management for testing and promoting changes, and some sort of visualization to make architecture easier to design as a team.
I work for a platform company called Stackery which aims to provide an end-to-end workflow and DX for serverless & CloudFormation. Thanks for comments like these that help identify pain points that need attention.
What I have now is a loosely coupled, serverless, frontend+backend monorepo that wraps AWS SAM and CloudFormation. At the end of the day it is just a handful of scripts and some foundational conventions.
I just (this morning!) started to put together notes and docs for myself on how I can generalize and open source this to make it available for others.
stack is vue/python/s3/lambda/dynamodb/stripe but the tooling I developed is generic enough to directly support any lambda runtime for any sub-namespace of your project so it would also support a react/rails application just as well.
Sure, you can write a bunch of tools to work around the crufty, terrible development environment's shortcomings. But ultimately, you are just locking yourself further & further & further in to the hostile, hard to work with environment, bending yourself into the bizarre abnormal geometry the serverless environment has demanded of you.
To me, as a developer who values being able to understand & comprehend & try, I would prefer staying far far far away from any serverless environment that is vendor locked. I would be willing & interested to try serverless environments that give me, the developer, the traditional great & vast powers of running as root locally that I expect. Short of a local dev environment, one meets both vendor lock in, & faces ongoing difficulties trying to understand what is happening, with what performance profiles/costs. I'd rather not invest my creativity & effort in trying to eek more & more signals out of the vendor's black box. Especially if trouble is knocking, then I would very much like to be able to fall back on the amazing toolkits I know & love.
That isn't a compliment.
Time will tell.
but lambda gets to the point where there is no local parity, where it's detached, is no longer an easier managed (remotely operated) parallel to what we know & do, but is a system entirely into itself, playing by different rules. one must trust the cloud-native experience it brings entirely, versus the historical past where the cloud offered native local parallels.
If you are working on a project where all of your infrastructure will live on AWS I would definitely urge you to give it a second look. The amount of infrastructure I manage right now with a single .yaml file is really killer.
Personally for simple projects I've had pretty good experiences writing a Yaml-based template directly, and for more complex projects I use Troposphere to generate Cloudformation template Yaml in Python.
Did you chose python for backend dev? Any framework like flask or django?
Also no preference of single-language for both backend and front-end, by using node.js backend?
Just trying to get into web-dev.
If I could do it all over, I would still choose Python. That being said, I have been working professionally (building apps like this) for almost 14 years so my willingness to bite off a homebrew Python framework endeavor as I did here is a lot different than someone just getting into the field.
Django: avoid unless you have a highly compelling (read: $$$$) reason to learn and use this tool. I cannot think of one, honestly.
Flask: fantastic, but be conscientious about your project structure early on and try to keep businees-logic out of your handler functions (the ones that you decorate with @app...)
Sophisticated or more sugary Node.js backends are not something I have ever explored, aside from the tried-n-true express.js. I tend to leverage Python for all of my backend tasks because I haven't found a compelling reason not to.
You have to roll your own way too often in Flask et al, so much so that I don't see any reason to use Flask for anything other than ad-hoc servers with only a few endpoints.
All-in-all, Django is not bad software. I have a bad taste in my mouth though because as I learned and developed new approaches to solving problems in my career I feel like Django got in the way of that.
For instance, there are some really killer ways you can model certain problems in a database by using things like single table inheritance or polymorphism. These are sorta possible in Django's ORM, but you are usually going against the grain and bending it to do things it wasn't really supposed to. Some might look at me and go: ok dude well don't do that! But there are plenty of times where it makes sense to deviate from convention.
That is just one example, but I feel like I hit those road blocks all the time with Django. The benefit of Django is it is pre-assembled and you can basically hit the ground running immediately. The alternative is to use a microframework like Flask which is very lightweight and requires you to make conscious choices about integrating your data layer and other components.
For some this is a real burden - because you are overwhelmed by choice as far as how you lay out your codebase as well as the specific libraries and tools you use.
After your 20th API or website backend you will start to have some strong preferences about how you want to build things and that is why I tend to go for the compose-tiny-pieces approach versus the ready-to-run Django appraoch.
It's really a trade-off. If you are content with the Django ORM and everything else that is presented, it is not so bad. If you know better, you know better. Only time and experience will get you there.
Minimalist frameworks are great for either very small (since they don’t need much of anything) or very large projects (since they will need a bunch of customization regardless).
In that regard, I think Django is kind of like the Wordpress of Python.
I don’t think it’s as bad as people make it out to be.
I've been pretty happy with Cloudflare Workers.
You can easily define environments with variables via a toml file. The DX is great and iteration speed is very fast. When using `wrangler dev` your new version is ready in a second or two after saving.
I was planning on trying out FastAPI + Zappa, but have not gotten round to it yet.
Disclaimer: I work at AWS, however not for any service or marketing team. Opinions are my own.
Of interest, I've spent some free time crunching on CNCF survey data over the past few months. Some of the strongest correlations are between particular serverless offerings and particular delivery offerings. If you use Azure Functions then I know you are more likely to use Azure Devops than anything else. Same for Lambda + CodePipeline and Google Cloud Functions + Cloud Build.
This is what I tried to do initially after experiencing the dev pain for only a few minutes.
But unfortunately this doesn't work very well in anything but the most trivial case because as soon as your lambda has a 3rd party package dependency you need to install that dependency somehow.
For example, let's say you have a Python lambda that does some stuff, writes the result to postgres and then sends a webhook out using the requests library.
That means your code needs access to a postgres database library and the requests library to send a webhook response.
Suddenly you need to pollute your dev environment with these dependencies to even run the thing outside of lambda and every dev needs to follow a 100 step README file to get these dependencies installed and now we're back to pre-Docker days.
Or you spin up your own Docker container with a volume mount and manage all of the complexity on your own. It seems criminal to create your own Dockerfile just to develop the business logic of a lambda where you only use that Dockerfile for development.
Then there's the whole problem of running your split out business logic without it being triggered from a lambda. Do you just write boiler plate scripts that read the same JSON files, and set up command line parsing code in the same way as sam local invoke does to pass in params?
Then there's also the problem of wanting one of your non-Serverless services to invoke a lambda in development so you can actually test what happens when you call it in your main web app but instead of calling sam local invoke, you really want that service's code to be more like how it would run in production where it's triggered by an SNS publish message. Now you need to somehow figure out how to mock out SNS in development.
Serverless is madness from start to finish.
SAM already provides a way to mock out what SNS would send to your function so that the function can use the same code path in both cases. Basically mocking the signature. This is good to make sure your function is running the same code in both dev and prod and lets you trigger a function in development without needing SNS.
But the problem is locally invoking the function with the SAM CLI tool is the trigger mechanism where you pass in that mocked out SNS event, but in reality that only works for running that function in complete isolation in development.
In practice, what you'd really likely want to do is call it from another local service so you can test how your web app works (the thing really calling your lambda at the end of the day). This involves calling SNS publish in your service's code base to trigger the lambda. That means really setting up an SNS topic and deploying your lambda to AWS or calling some API compatible mock of SNS because if you execute a different code path then you have no means to test the most important part of your code in dev.
> In the case of Postgres, you could use an ORM that supports SQLite for dependency-free development but at a compatibility cos
The DB is mostly easy. You can throw it into a docker-compose.yml file and use the same version as you run on RDS with like 5 lines of yaml and little system requirements. Then use the same code in both dev and prod while changing the connection string with an environment variable.
> That’s not exactly a knock against serverless itself, is it?
It is for everything surrounding how lambdas are triggered and run. But yes, you'd run into the DB, S3, etc. issues with any tech choice.
And yes, I think the problem here are APIs you use in multiple places but can’t easily run yourself in a production-friendly way. Until AWS builds and supports Docker containers to run their APIs locally, I don’t see how this improves... end to end testing of AWS requires AWS? ;-)
It’s hard to do this when surrounding code doesn’t accommodate the approach, but it’s great way to design an application if you have the choice. I really love sandboxing features before moving them into the application. Everything from design to testing can be so much faster and fun without the distractions of the build system and the rest of the application.
There's a couple of sharp edges still but in general it just 'makes sense'. If you don't like TypeScript there are also bindings for Python and Java, among others, although TypeScript is really the preferred language.
Currently using some CDK in a production app and finally I found a way of doing IaC I actually enjoy.
CDK is moving quite fast and not all parts are out of the experimental phase, so there are breaking changes shipped often. I think in a couple of years it will stabilize and mature and become a very productive way of working with infrastructure.
Have you tried Amplify?
Unfortunately, I quickly hit the limits of its configurability (particularly with Cloudfront) and had to move off it within a few months.
> serverless invoke local
All the time for doing testing. You can pass in a json file that contains a simulated event to test different scenarios. Works very well for us.
I believe they decided to re-create their own credentials chain and a lot of things are not working with MFA (like support for credential_process).
Edit: CloudFormation was also painful for me, the docs were sparse and there were very few examples that helped me out.
Over 5000 pages of documentation
https://docs.aws.amazon.com/AWSCloudFormation/latest/UserGui...
> very few examples that helped me out
AWS provide an example for most common services, plus there are thousands of other community supplied examples out there.
Yes there are examples, but there wasn’t one at the time that mapped to what I was trying to accomplish. Because, again, SAM templates are not one-for-one CloudFormation templates.
I found the community around SAM to be very limited. One of the many reasons I’ve moved to the Kubernetes ecosystem.
Honestly, it reminds me of PHP development years ago: running it locally sucked, so you need to upload it to the server and test your work. It. Sucked.
With SAM we're talking ~6-7 seconds with an SSD to build + LOCALLY invoke a new copy of a fairly simple function where you're changing 1 line of code, and you need to do this every time you change your code.
That's even with creating a custom Makefile command to roll up the build + invoke into 1 human action. The wait time is purely waiting for SAM to do what it needs to do.
With a more traditional non-Serverless set up, with or without Docker (using volumes) the turn around time is effectively instant. You can make your code change and by the time you reload your browser it's all good to go. This is speaking from a Python, Ruby, Elixir and Node POV.
(I'm still doing the "upload to server to test" thing.... I've tried MAMP and Vagrant/VirtualBox for local dev but both of them seem horribly complex compared to what we can do with local dev with node.js/mongo and so on.)
Of course this won't fix the fact that you have a lambda behind api gateway that does some heic->jpg conversion and can't be hit outside the DMZ, or some esoteric SQS queue that you can't mimmic locally - but it should get you almost there.
sudo apt install nginx php mysql, point the www directory of nginx to something on your computer (/mnt/c/projects/xyz) and you've got a running setup. Or run Linux in general, that's what most people I've seen work on backends seem to do. You can run the same flavour of software that your deployment server runs so it'll save you time testing and comparing API changes or version incompatibilities as well.
I don't know any solution for macOS but you can probably get the necessary software via Brew if you're so inclined. Then run the built-in PHP web server (php -s) and you get the same effect.
Apache is almost as easy, download zip, unpack & configure apache conf to use your php.
MySQL is somewhat more complicated because you need to run an setup script after unpacking the zip.
I tried AWS Lambdas a few years back and it felt way more primitive.
Yes, SAM sucks. Boto3 sucks. I mean, yes, the AWS APIs are not the most obvious or ergonomics thing.
TJ abstracts the Lambda nonsense so you just get to use a standard HTTP res/resp interface. https://github.com/apex/up-examples/
Deployment iterations for me in Singapore can be as low as 2s, but tbh I use https://github.com/codegangsta/gin locally for development iterations.
What does it mention?
Still very actively working on this but our goal is that nobody ever has to test via deploy again.
Slightly OT... perhaps cheeky... any idea why firebase doesn’t provide automated backups for firestore and storage out of the box? Seems like a no brainer and a valuable service people would pay for.
DISCLAIMER: I used to be a developer of this product (it's from the previous company I've worked)
My only nitpick, and only specifically relating to dotnet, is that config files and env vars differ between Functions and regular ASP.NET Core web apps. I think there is some work going on to fix that, but it's taking forever.
“Serverless models don’t require users to maintain their own operating systems, or even to build applications that are compatible with particular OSs. Instead, developers can produce generic code, and then upload it to the serverless framework, and watch it run.”
... is utterly compelling and is why serverless will not just win, but leave renting a server a tiny niche market that few developers will have an experience of post 2030.
Maintaining your own server is completely nuts. If that isn’t obvious now, it will be in another decade. It’s massively inefficient. Like running your own power plant to serve your factory, except you also have to worry about security and constant maintenance, along with all the moving parts that surround a server.
Almost all the objections in the article can be rephrased as “serverless is not mature enough yet”, and that’s accurate, but I suspect there’s also a bias against giving up control to the cloud companies, and some wishful thinking as a result.
The future of software development is going to be defined by cloud providers. They’re going to define the language ecosystem, the canonical architectures for apps etc... it’s just early days and cloud is really very primitive. Just clicking around Azure or GC or AWS illustrates how piecemeal everything is. But they have a lot to do, and just keeping pace with growth is probably hard enough. I’m not sure I’m super happy with this outcome, but I’m pretty certain the trend line is unmissable.
We need self-hosted serverless...
TFA is clickbait
There is nothing generic about the code that runs on serverless services. It’s the ultimate lock in.
Is there anything stopping an organisation from defining some standard types of serverless environments?
Is there anything stopping someone from turning that standard into implementations to help cloud providers offer it, or even be a fallback option that could be deployed on any generic cloud infrastructure?
I think those are the way forward from here.
Yes. Basic economics. There is nothing _technical_ stopping 'an organisation' from making a federated twitter or facebook. But there are (evidently) insurmountable non-technical reasons: It hasn't happened / there have been attempts which have all effectively failed (in the sense that they have made no significant dent in these services' user numbers).
Why would e.g. Amazon (AWS) attempt to form a consortium or otherwise work together or follow a standard, relegating their offerings to the ultimate in elasticity? Economically speaking, selling grain is a bad business: If someone else sells it for 1ct less per kilo then the vast majority of your customers will go buy from someone else, there's no product differentiation.
Serverless lockin (such as GAE or AWS Lambda) is the opposite. No matter how expensive you make the service, your users will stay for quite a while. But make a universal standard and you fly in one fell swoop to the other end of the spectrum. If I have a serverless deployment and the warts of serverless are fixed (which would, presumably, involve the ability to go to my source repo, run a single command, give it some credentials, and my service is now live after some compilation and uploading occurs) - then if someone else offers it 1ct cheaper tomorrow I'll probably just switch for the month. Why not?
This cycle can be broken; but you're going to have to paint me a picture on how this happens. Government intervention? A social movement amongst CEOs (After the war, there was a lot of this going around)? A social movement amongst users so aggressive they demand it? Possible, but that would require that we all NOT go to serverless until the services offering it come up with a workable standard and make commitments to it.
I think cloud customers are savvy to the lock-in. That we're having this conversation in evidence of that. Perhaps AWS can achieve adoption of Lambda without needing to cater to customers who are cautious about getting locked in, but any challenger might find that it's much easier to gain customers if they also provide some form of an escape hatch.
As Jeff Bezos would say about retail, "your margin is my opportunity."
We have some already, e.g. RFC3875 and its descendents https://tools.ietf.org/html/rfc3875
* Well, obviously we use a hosted database
* And obviously, AWS provides our logging and all our analytics.
* Obviously when people call our lambda functions, they do so either through an AWS-specific API, or one constrained to a very limited set of forms.
* Of course, we can't blindly let everything access everything, so naturally we have IAM roles and permissions for every lambda function.
* Well, the cloud provider will look after secrets and things like that for us, no need for us to worry about database passwords.
* Naturally, with all these functions and IAM roles to look after, and we need tagging for billing. We should define it all with CloudFormation scripting.
* Well, the nosql database the they provide comes with their specific library. And as it shards things like this, and doesn't let you index things like that, you've got to structure your data this specific way if you want to avoid performance problems.
* You don't want your function to take 200ms+ to respond, your users will notice how slow it is. So no installing things with apt-get or pip for you, let me get you a guide on how to repackage those into the vendor-specific bundle format.
* You want to test your functions locally, with access to an interactive debugger? You're living in the past, modern developers deploy to a beta environment and debug exclusively with print statements.
* And so on.
In this case, a lot of the 'complexity' one hoped to eliminate has just been moved into XML files and weird console GUIs.
It was a complete omnishambles - to the point I avoided the fancy lunch and went to the Pub for a ploughman's lunch, In case I suddenly blurted out "this is all S*&T" and caused a political row with the mobile side of the company I worked for.
Of the actual function code, sure.
Of course, if you aren't manually configuring everything and are doing IaC, that isn't generic. And if you are supporting the Lambda with any other AWS serverless services, the code interfacing with them is pretty AWS specific.
Sure they do.
Because if you need a DB for persistence you are either setting up a server for it (and therefore not serverless, even if part of your system uses a Lambda) or consuming a serverless DB service. And so on for other parts of the stack. On AWS, for DB, that might be Dynamo (lock-in heavy) or, say, Aurora Serverless (no real lock-in if you don't use the RDS Data API but instead use the normal MySQL or Postgres API, but that's higher friction—more involved VPC setup—to use from Lambda than RDS Data API is, so the path of least resistance leads to lockin.)
Lambda or other FaaS is often part of a serverless solution, but is rarely a serverless solution by itself.
Like, today, given the political uncertainty in the USA, any large company would be nuts to bet on hosting their critical services on US-dependant infrastructure without having a huge plan B already in the works.
That has always been my sticking point. With AWS in particular. I don't trust AWS to exist forever, and it's definitely not without it own ongoing maintenance issues. Locking my entire business into their ecosystem seems risky at best.
Something I see really often in startups is huge dependencies in the form of SaaS. It should be no secret to those in tech that many of these businesses will not be around in 3 years. Even the likelihood of their service staying the same for 3 years is pretty low. I have been bitten by enough deprecated services, APIs and Incredible Journeys that I am wary.
Instead of maintaining software compatible with the right operating system, you're maintaining software compatible with the right flavor of serverless by Cloud Provider. Now we're back to square one on at least one front.
On the control aspect, the bias against giving up control is not an unwarranted one. Maintaining control of critical infrastructure is extremely important. And in fact outsourcing your critical infrastructure is an existential one, and not just in an academic sense. When you give up control you give up the ability to prevent your infrastructure being hostile, but incompetent. In these cases it reduces the quality of your product for your customers.
I won't even go into the anti-competitive tactics Amazon themselves get into that make them not a good choice for your infrastructure. Instead I'll draw upon a recent experience that illustrates why outsourcing infrastructure, even at a higher level, is a bad idea.
My girlfriend recently was taking her NLN exam remotely. They weren't allowed to use calculators of their own, they had to use a virtual calculator provided by the company administering the test. Like most of these companies they are doing remote proctoring of the exams. During her exam this virtual calculator flat out wasn't available. The proctor told her to simply click through the exam and that once she submitted it she'd be able to call into customer service to get the exam re-scheduled due to the technical difficulties.
Well, that wasn't the case. After doing some deep digging for her, here is what I found. The testmaker NLN contracted out the test administration to a third party, Questionmark. Questionmark in turn contracted out yet another 3rd party, Examity to handle the proctoring. Examity proctors don't have access to Questionmark's systems. Questionmark doesn't have access to NLN's systems, etc. So how did we get this resolved? I had to track down the CEO of Questionmark, the CEO of Examity, and the head of testing for NLN. I had to reach out through LinkedIn inmail to get this on their radar. And then it was handled(quickly and efficiently I might add!). However, frontline support for each of these companies could do nothing. They just had to offload blame onto the support staff of each other. Another aspect of this is that each handoff between 3rd parties creates a communication barrier. In this case the communication barrier seems to have kept Questionmark from configuring this specific test correctly. I wouldn't blame any of these companies for this specific failure mode because it's just the nature of what happens when you offload your work to 3rd parties.
When you say, "Oh it's great we don't have to worry about X because Y can do it." The implication is that you lose all of the power of a vertically integrated company by essentially spinning off tons of subsidiaries and creating a communication overhead both before and when problems DO arise.
What is the future of software development going to look like when it reaches consumers and you have to say "Sorry, we can't fix that issue because Cloud Provider has to get back to me, and then in the background Cloud Provider has to say sorry we can't get back to you because we have to wait for our spunoff hardware division to get back to us?"
Maybe this type of business is fine for fun apps, but it's not fine for a lot of businesses. Even SLAs and disclaiming responsibility in your own contracts won't save your reputation. All it does is protect you financially!
Except that it is not. The security and constant maintenance is needed but it is worth in many cases. And large companies can not really off load all ownership of data and applications. The interest in servers actually went in reverse direction due to the cloud effect.
I would say having own server or application hosting capacity is very similar to why you produce your own solar power and store it in batteries - it is simple? No. But the technology is improving and thus it makes it easier for people to adopt this paradigm.
When I see what is happening with WebAssembly/WASI in particular, I see a great future for self-hosting again. Software written in any programming language (as long as it targets WASM) is a lot easier to host than existing models. Also there is interoperability of software coming from different languages at the WebAssemply level as I understand.
For the past 20 years you have been able to target x86 gnu/linux and have it running without modification on a readily available server, either your own hardware or rented/public cloud. How does switching from one binary format to another (x86 to WASM) change anything (except maybe slowing down your code)? As I understand it, the main draw of WASM is running non-JS code in a web browser.
Simplified deployment + automatic sandboxing?
AFAIK with x86 you can't just write a client app and have it automatically run on any computer that visits your website.
That's... quite depressing to consider, actually. I long for a return to the internet of yore when Native apps were still king and not everything was as-a-service.
I doubt cloud providers can dictate environments, other providers would quickly fill the gap to meet developer preferences. I also think that more developers care about lock in these days.
It's called time-sharing and it existed in the 1960's. [1]
[1] https://en.wikipedia.org/wiki/Time-sharing
The difference between time-sharing and serverless is that the former solved the issue of expensive personal computing, until cheap personal computers took over that market. The latter solves perceived expensive computing on the "server" side.
But what does it serverless solve exactly? It doesn't solve a technical problem, rather addresses concerns on the business side. Serverless solves a cost problem.
First, computing needs aren't linear, they fluctuate. And so, there's a problem of under- and over-utilization vs availability of resources. Serverless approaches computing power like tapwater: you're essentially paying for the CPU time you end up using.
Second, elasticity. Instead of having staff struggle - losing time - with the fine intricacies of autobalancers, sharding and what not; you outsource that entirely to a cloud provider. Just like a tap, if you need more power, you just turn the tap open a bit more.
Finally, serverless services abstract any and all low level concepts away. Developers just throw functions in an abstraction. The actual processing is entirely black box. No need to worry about the inner details of the box.
Sounds like a good deal, right?
> Like running your own power plant to serve your factory, except you also have to worry about security and constant maintenance, along with all the moving parts that surround a server.
Well... no. Outsourcing all of that to a third party cloud computing vendor doesn't dismiss you from your responsibility. All it does is shift accountability to the cloud provider who agreed to take you on as their customer. Securing your factory still very much includes deploying a secure digital solution to manage your machinery and process lines.
Plenty of industries wouldn't even remotely consider outsourcing critical parts of their operations, and this would include digital infrastructure. And this is regardless of the maturity of serverless technology. Risk management is a vast field in that regard.
Then there's legal compliance. There are plenty of industry specific regulations that simply don't even allow data to be processed by third party cloud services unless stringent conditions are adhered to. Medicine, banking and insurance come to mind.
Finally, when it comes to business critical processes, businesses aren't interested in upgrading to the latest technology for the sake of it being cutting edge. They want a solution that solves their problem and keeps solving that problem for many years to come. Without having to re-invest year after year in upgrades, migrations and changes because API's and services keep shifting.
Does that mean that there isn't a market for serverless computing? Of course there is. Serverless computing is a JIT solution. It's an excellent solution for businesses in a particular stage of their growth. And it closes the gap for plenty of fields where there really is a good match. I just feel that "maintaining your own server is completely nuts" is a bit overconfident here;
It is nuts to run one server. Then you're wasting money with a server/VM. That's what serverless is ideal for: stuff no one uses. That's a real niche. Who's going to use that? Not profitable companies.
Often I think for most cases where you reach for serverless, you should reconsider the choice of a client-server architecture. An AWS Lambda isn't a server anymore; it's not "listening" to anything. Why can't the "client" do whatever the Lambda/RPC is doing?
Maybe what you want is just a convenient way to upload code and have it "just work" without thinking about system administration. The types of problems where you don't care about the OS is once again a niche. You probably don't even need new software for these kinds of things. You can just use SaaS products like Wordpress, Shopify, etc.
Serverless won't be profitable because the people who need it don't make money.
You seem to implying only applications with huge numbers of users can be profitable. This statement ignores a tremendous amount of (typically B2B) applications that provide enormous value for their users but don't see a lot of traffic.
I have worked on applications that are at the core of profitable businesses yet they can go days, in some cases weeks without any usage. Serverless architecture will be a real benefit there once it matures.
I've in this area professionally for some time now, and I've never "maintained" servers in any reasonable sense. There are kernel people who maintain the kernel and there are Debian devs who maintain the operating system. The server may be mine (but more often that not, it isn't) but only in very specific circumstances do I ever concern myself with maintaining any part of this stack.
A vanilla Linux VM is a platform to build on. Just like AWS or anything else. It is the environment in which my software runs.
Thus far, something like Debian has been more stable and much less of a moving target than any proprietary platform has been, cloud or non-cloud. Should a client wish to minimize maintenance costs of the software over the coming decade, it is most cost effective to not depend on specialized, proprietary, platform components.
That may change in the future, but right now there is no indication that is the case.
There's also the long-time guarantee; if you write an application that runs on Ubuntu Server or Windows Server now, you can bet your ass that it will still run, unchanged, for another 10 years. The only maintenance you need to do is to fix your own bugs and maybe help out with some database stuff. If you deploy a Lambda app now, you have nothing to guarantee compatibility for such a long time other than "Amazon probably won't change their API, I think".
Using lambdas tie you to AWS because as soon as you use a few step functions, or you have a few lambdas interacting, changing to Azure or GCP becomes a huge pile of dev work and QA.
Having everything "just run" on linux instances let you be perfectly portable and now you can actually shop around.
Had you deployed an application to a proprietary cloud platform ten years ago, a handful of those services would have had their APIs changed or even been sunset by now.
I agree that it almost never happens, and that's why I run Debian as well. However, if you run production then things happen.
As an experienced back-end developer and linux user I would pull my hair out if I was completely helpless in fixing an issue or implementing some side thing that requires shell access. I don't want to wait for some guy in Philippines who will be online 12 hours later to come try to fix it.
Instead, developers can build applications that are compatible with particular serverless frameworks.
On the ops side even with a platform like Lambda, training an operations team to take over maintenance of a nested spaghetti of random interlinked services and bits of YAML trapped in random parts of the cloud is a total nightmare. The amount of documentation required and even the simple overhead of enumerating every dependency is a long term management burden in its own right. "The app is down" -> escalate to the developers every single time.
Compare that to "the app is down", "this app has basically no ops documentation", "try rebooting the instances", "ah wonderful, it came back"
I'm pro-cloud in many ways and even pro-serverless for certain problems, but let's not close our eyes and pretend dumping everything into these services is anything like a universal win.
Also - and I'm an infra guy so I'm probably biased - I don't really get all this developer anxiety to outsource infra needs. Yeah if you are 2 devs working on your startup it makes sense, but when you scale up the team/company, even with serverless, you WILL need to dedicate time to infra/operations and do not dedicate it to strictly-business-related code. Having somebody dedicated to this is good for both.
Yes, Ops will still have stuff to do, it will just be at another level.
It's inevitable, that's why we're not making our own memory cores anymore.
I haven't done anything with serverless, but surely the class of problems that would be fixed by an instance restart don't happen in the first place on serverless
What we have now is very primitive compared to how app development might work in the future; serverless is laying the foundation for a completely different way of thinking about software development.
It's more back to the mainframe model of software development. I did this back in the 90s and I never had to think about scaling. Granted these were just simple crud / back-office apps.
But I can see how it would work for most modern software.
Upon learning about it some time ago, this was exactly my conception of what a Lambda-like serverless architecture would yield.
And it would seem difficult, if not impossible, for any dev to maintain a mental map of the architecture.
TANSTAAFL. Let’s say serverless becomes commoditised the same way as electricity. What are the margins in that business? What are AWS etc margins now?
There are very strong reasons to believe that serverless will offer convenience at a premium price, forever.
You don't patch your OS, or your apps, you define their versions and configuration in code and it gets built for you. And typically snapshotted at that point and made into a restartable instance image that can be simply thrown away if it's misbehaving, and rerun from the known good image.
My first experience with serverless architecture was back in 2007 or so when trying to port Google News to App Engine. That was a thoroughly painful experience, and things haven't exactly gotten much easier since. If you go back in time a decade, Google's strategy for selling compute capacity was App Engine. Amazon went the EC2 route. Reality suggests AWS made the better choice.
I can understand the superficial notion that having idling virtual machines is inefficient (because it is). But this reminds me a bit of tricks we did to increase disk throughput in systems with lots of SCSI disks back in the day. Our observation was that if we could keep the operation queues on every controller full as much of the time as possible, we'd get performance gains when controllers were given some leeway to re-order operations, resulting in slightly better throughput. Overall you got higher throughput, and thus higher efficiency, but from the perspective of each process in the system, the response was sluggish and unpredictable. Meaning that if you were to try to do any online transactions, it would perform very poorly.
For a solution to be desirable it has to be matched with a problem that needs solving.
As the article points out, there are some scenarios where serverless architectures might be a good design paradigm. But extrapolating this to the assumption that this paradigm is a universal solution requires not only a leap of faith, but it also requires us to ignore observed reality.
So you owe me some homework. Tell me what needs to happen for serverless architectures to reach "maturity".
None of that requires OS maintenance. My house plants require far more maintenance than my software. I sometimes forget where some things run because I haven't touched them in years.
Serverless code runs on the cloud you built it for. I don't want that. I don't want to invest years of my life becoming an Amazon developer, writing Amazon code for Amazon servers.
That's without delving into the extra work serverless requires. There isn't a dollar amount tied to my import statements. I don't need software to help me graph and make sense of my infrastructure. I can run my code on my computer and debug it, even offline.
Your hosts are still packed with a bunch libraries and services (sshd for example) that should probably be updated with regularity.
I echo a lot of what you say here regarding run anywhere and not marrying some giant vendor.
On hosts I manage professionally, I update/upgrade weekly after reading the notes - it takes a few minutes, I know I'm up to date and if there is anything I should be wary of.
On a personal debian server, I have an update/dist-upgrade -y nightly on a cron job, and I reboot if I read on HN/slashdot/reddit/lwn about an important kernel fix; Never had an issue, and I suspect it's about as secure and trouble free as whatever is underlying lambda -- with the exception that every 3-4 years I have to do an OS upgrade.
That hard drive event showed me how disposable the machine itself has become thanks to docker.
How much would you even pay for hosting? $2 a month?
My point was about having portable code on generic hardware, which in my opinion is a better bet than writing Amazon software for Amazon servers, and praying their prices don't change much.
Then how do you know they are still secure and even working?
Yes, deploying servers is very easy, maintaining and securing them is the hard part. Sure, you can automate the updates and it will work with a good OS-Distribution for some years. But no system is perfect, exploits are everywhere, even in your own configuration. And then it becomes tricky to protect your data.
Also no need to be that scared about servers.
With newer power technologies becoming more affordable and effective, solar, wind, & storage are increasingly being used to power factories and other businesses.
It's all about control of your product and operations. If it is economically feasible, it's always better to control your own stack all the way down.
Does serverleds actually deliver better control over your development, portability, reliability, security, etc. for your application & situation, or not?
This sounds a bit like the "When will they turn off the last mainframe?" arguments a while back - I wouldn't expect servers to disappear either...
That's probably true of web development, but "software development" writ large is much more than just cloud providers and webapps. "Software" encompasses everything from embedded microcontrollers to big iron mainframes that drive (yes, even today, even in 2030) much of the world's energy, transportation, financial and governmental infrastructure.
But long-term I think the cloud providers will assimilate so much power that the embedded side will follow its diktats.
Big iron mainframes have longevity, but will absolutely die out close to complete extinction - I've worked on those systems, I understand their strengths and the legacy issues, and there's no way that cloud isn't going to gobble up that market, it's just going to take a long time (as you say, beyond 2030).
It allows me to stand up a practically maintenance-free endpoint in a matter of hours (usually to glue separate services together):
* Want to run a quick ETL process to refresh some tables or pull some data down from a third-party endpoint?
* Want to expose a quick usage report for the project manager?
* Want to send a slack message when a build finishes from TeamCity so QA can jump on it?
* Want to monitor the rate of Crashlytics logs and send a warning to slack if it breeches a threshold?
There are tons of "one-off" little processes that now just automatically run and don't require server upkeep. Most of them are functionally independent, so having them as separate functions makes sense (when we do need to coordinate, that's done through Step Functions).
Note: I still don't think that they are a great fit for production services (and more fits one-off and cron usage patterns)
At least some people mention "but I can scale this lambda x1000" and that's one advantage... But you can do all those tasks on a hetzner server in same amount of type, just python scripts.
I don't want to have to manage security updates, disk usage, or any of the other multiple things that come with owning a server.
I also could probably debate on the friendly-ness of just dumping all of these one-off services onto a single server (seems like it could get messy pretty quick), but I think you could probably go either way on that.
We are also an AWS shop for our "production" stack already, so there is no added overhead for owning an AWS account.
Let's say that Hetzner server goes completely toast at 2am and you get a pingdom message. You log in, realize you have to restore from backup. You fire up a new instance, and go through your restore procedure, and everything works great. You're lucky to be offline for an hour.
I cannot even tell you what would happen if an underlying instance running a Lambda goes offline, because it never happens. But let's say that an underlying machine dies in us-east-1. Lambda just starts running more instances of your code on another underlying group of cpus, and you aren't offline.
I wish there was some easier way to quantify the “soft” aspects of everything that surrounds us, e.g. the long-term impact of beautiful and usable UI design, the reduction in “ambient” psychological stress, the impacts of chaos/consistency on how we feel, any robust measures of happiness, excitement, relaxation, etc.
They seem to be such important attributes, but not as easily measured as a latency. Which in turn makes reaching consensus harder—and probably why design by committee doesn't work?
My entire goal is for my services to never wake me up. And if they do it had better because I clearly messed up and am the sole human alive able to fix it.
There's no point in waking up when everything is blowing up if you don't have the access to fix the underlying issue. AWS doesn't actually guarantee 99.9% uptime. What they guarantee is they'll give you some company currency[1] if they can't meet or exceed 99.9% uptime.
So the next time the ship is on fire, don't worry, stay asleep and your Amazon account manager will be by shortly with a thimble of water to throw on you :)
[1] - https://aws.amazon.com/about-aws/whats-new/2019/03/aws-syste...
Bravo.
Having been on call for some truly monstrous amount of things and seen things break in all kinds of ways... I can't agree more with this principle.
While the cloud isn't cheap, if your org can afford it, there is no question about it: go with managed services. Focus on building things, leave the OPS to AWS/GCP/Azure.
Are you really sure that it all re-ran? No. You have to query your latest batch - which means you have to have arbitrary boundaries for when failures might have occurred like say, last 24 hours. If was a weekend, just query the whole database looking for any entry that's missing the "final" status success. Try not to run anything while you're doing this, unless your system is tolerant to the DB being locked for a full-table scan.
This is all very fragile and change-averse. Who would choose to run an ETL like this? I have worked for large companies who all end up with this rickety system. Every one, every time. God forbid you want to run tests with a mock set of serverless AWS services.
Our use of AWS Lambda for ETL is usually just one-off small processes that don't have dependencies.
As soon as you start getting into "This ETL depends on X, Y, and Z being run first" I think you're out of "small utility" territory.
Meanwhile, I don't have to do things like /make sure the ETL process is running/ or /clear out a bunch of logs when a disk inadvertently fills up/ or any of those other stupid last-century compute problems. It's just a far, far easier way of thinking about application development.
I can appreciate that. A lot more people than you might think, use Kafka or SQS because of various requirements (like that's all their infrastructure supports) and you end up with a massive "source" DB which is a duplication of all-time messages. Such is the reality of a larger byzantine organization. If it can't be audited by someone else, at their own pace, it's not approved.
Does it make sense for FAANG to run their own datacenters? Yes of course. Does it make sense for you? No. This is what I don't get about arguments for running a bunch of VPSs-- for literally any environment that isn't thousands of engineers, it is more cost effective to let someone manage your underlying infrastructure for you, and any workload at that scale, you aren't contracting with some random bare metal hosting shop.
I believe he is comparing aws bills vs. engineer time+other hosting bills.
This is how we actually end up having centralized services. The larger companies who have all this infrastructure, and have the size to justify building their own.. well they inevitably ask themselves "Hey! we have all this infra, we have to pay these people anyways, why aren't we selling excess capacity?"
Then you get AWS. Then as time progresses people at even large institutions that could benefit from running their own infra just go and use another company's infra. Where the other company has evolved into being an infra provider instead of a car parts store, book store(amazon), or search engine.
To answer your question directly. Yes I think this can be done within one engineer salary even if I'm including my own salary. There would be a startup cost that eats into that salary in year 1. But it would be manageable. And then iterative upgrades/replacement parts would be a negligible cost ongoing after year 1. With some financial tricks that startup cost can be spread out over a long enough period that it wouldn't even impact year 1.
And we're approaching that because as these types of skills are outsourced to these institutions the amount of people with the skills necessary to carry them out are diminished in number from attrition. Amazon doesn't need 10 company's worth of engineers to have the requisite skills. If everyone buys into this model then eventually the market will only bear the number of engineers having that knowledge as Amazon needs.
You'd be spending more than 1/2 your time just on the phone talking to people.
If we want to talk about running a CDN company in terms of lots of additional organizational complexity in terms of customer service it can't be done on a single engineer's salary. My perception of a homegrown CDN is one that rivals uptime and deliverability metrics of existing commercial CDN providers. Within an existing org it can be done. If we're talking outside of an existing org think http://www.nosupportlinuxhosting.com/ to shave costs.
If we're willing to define parameters on what level of service this CDN needs to meet and ramp up time I'd be willing to take it on.
This is my point-- I am not in the hosting business. I am in the application creation business. Every minute that I even have to think about modifying /etc/localtime in order to make sure that the instance is running in UTC is a minute I'm not being productive. Multiply that by those million little things that you have to do in order to keep even an ec2 instance running, and it's such a colossal waste of time and effort. Similarly, I am not in the database tuning business. I plan out my data relationships, create the required Dynamo table or tables, and I never need to worry about replication or backups or any of those ancillary tasks that keep me from being productive.
I'm not a woodworker, machinist, farmer, or doing a PhD in literature. Nevertheless, I engage in all of these things. I enjoy crafting things and learning new techniques. I enjoy learning how to optimize plant growth. I enjoy reading old literature and gathering what I can from it about the historical context, linguistics, etc. All of these types of things, while not "productive" give me additional knowledge that I can pull from in "unrelated" tasks.
In the tech world, I'd say that realistically all abstractions are leaky. So engaging in these tasks(you mentioned) are productive. But even more than that because abstractions are leaky this is knowledge that is useful. If you don't have the knowledge of the systems underlying the abstractions you're using it's a footgun waiting to happen.
I can give an example. I had a boss that was doing some serverless work and he made an assumption about JSON. I told him in review that JSON does not respect ordering of keys. Well, he assumed it would not be a problem. It was pushed to production anyways. And a few weeks later it came back to bite him.
Now what happens when you don't dig into your abstractions? You lack the knowledge to fix problems that happen at a lower level. When you're deploying these "relationships" to your database provider. What happens when the machinery breaks down for replication, backups, etc? You have to sit around and wait for an expert to fix your app. Meanwhile you're staring at your users and stonewalling.
No, you're not in the hosting business, you're not in the database tuning business, you're not in the sysadmin business, and you're not in the Ecmascript working group. Until you are. Then when you're foisted into this role you're a fish out of water. Specialization is for ants, but we're humans, not ants.
Have fun with an army of one-trick pony hires. One old-school diehard computer nerd knows enough to supplant 10 of those fools, but only costs 2x.
About staff, maybe consider finding a good consulting/support firm that can look after your servers for <10% the price of hiring one person, are friendly and on-call, and will likely stick around even as their people change.
That solves the worry about replacing people, as well as the cost.
But as with people, it can take some luck to find a good one :-)
(I used to provide that kind of consulting/support service. Not as much now because there isn't much demand, and development & research work is more satisfying. But still a little, ticking along in the background.)
That's some ridiculous propaganda. AWS has whole regions going down for hours.
That's a big selling point for DigitalOcean: you sleep easy knowing what your bill will be. If you have an unexpected spike in traffic--whether a great opportunity, a mistaken test run wild, or a DDoS--it doesn't increase your bill. AWS offers the opposite: no matter what happens, we'll keep you online and just send you the bill.
Edit: add RDS to the parenthetical list of AWS services above.
Nothing a server template, code repo, and daily db dump cronjob can't solve
Dealing with server infrastructure isn't business differentiating. Doing it poorly can certainly be business terminating, though.
I agree, for most businesses no it isn't.
However, saving on cloud hosting costs to replace them with a handful of dedicated servers in different data centres can make a large difference to costs, depending on what the cloud usage is. For a funded rocketship startup with lots of free AWS credits it makes sense to use AWS; for a steady state or long-term low income business, not so much.
Those costs can make all the difference to a business if the cloud rental is creeping up to look comparable with salaries.
I've used a combination of dedicated and cloud hosting for a long time. In general the dedicated systems are so much cheaper for the same amount of bandwidth and compute, and you can run VMs on them with great performance, so you can run everything interesting inside VMs and a lot of the arguments around upgrade downtime have completely gone away.
With some cloud to assist, you can even use cloud temporarily to cover the downtime for major upgrades such as replacing hosts or reconfiguring the cluster network.
99.9% uptime with excellent bandwidth is easily achievable this way, and you can completely avoid user-facing downtime due to planned upgrades, with a bit of planning around DNS and IP transitions, VM migration, overlay networks, things like that.
Sometimes at much lower opex cost (100x) than the modern cloud-recommended equivalents, although it takes knowledge and time to ensure it.
If you have enough redundancy and occasional infrastructure stress testing, a similar level of "optimise for sleep" is possible as with cloud deployments. If you care about this (not every business does), it does take some work and knowledge to produce and verify hot redundancy, and of course there are costs.
Things will go wrong, but they'll tend to be at the level of things running inside VMs and containers, rather than hosts failing. The same things which go wrong with cloud deployments anyway.
The cloud can be used alongside dedicated to provide some of that redundancy and extra scaling. Cloud is rented in small units of time, so you can use cloud as "free most of the time" cold-backup service, and if there's a burst in traffic beyond what the dedicated systems can provide, or if you want to temporarily improve geographic locality for a service.
It might seem wasteful to rent larger dedicated services rather than cloud-as-needed, but the cost differential has changed in the last 10 years to favour the former more than it used to.
Even assuming that you save 100% of your hosting costs, you have to be spending quite a lot to even cover the salary of one full-time employee.
I can believe that larger businesses can save quite a bit, but I don't see how it makes sense for anyone spending less than seven figures a year on AWS.
For a small business, it makes sense to outsource things like that, even you if you could conceivably do it in-house.
I dunno, that's a problem for the ops team. It was in the 20th century too. The only reason you would have to worry about it is if you're pulling double duty as under-trained ops. If that were the case the right answer to every "build or buy" question is "buy" because you don't have anyone with the domain knowledge to bring more architectural options to the table than 'half-ass server management with internet tutorials.'
Then don't have backups at all, let's see how that goes.
One to run the application.
One to run the application while the first is being upgraded.
Now you need three servers, one to run the application, one as backup and one as load balancer.
99.99% uptime means you can be down no more than an hour a year. Which is really easy to overshoot when you're dicking around with a dist upgrade on three linux boxes.
So that's why you end up with a server farm starting from your one small server.
Drain a server, upgrade it, boot it, check that it works, let it accept traffic. Repeat once more. Done.
It may be true that most businesses don’t need the uptime, but I’d say it’s usually easier to give them high availability rather than field complaints when some random user tries to look something up during a maintenance window.
I don’t think of this as particularly fancy or expensive, considering how much you have to pay for engineers and technicians. It was critical enough for >3 nines, and that means having a technician on-call.
I saw this setup at two different companies I worked at, even though the tech stacks were completely different. Whenever I price out setups for similar requirements in cloud I almost always end up with basically the same setup, just on cloud VMs instead of dedicated servers.
My bank has regular scheduled downtime, Steam/Dota has regular downtime, entire Microsoft datacenters had issues and had around a day of downtime, Lime, Uber, my mobile carrier - all had downtime or bugs or glitches that were equivalent to downtime.
If you really need crazy uptime, then you must make sure you have no bugs, because bugs cause more downtime than any hardware failures.
That means an approach to software development that isn't relying on flavour of the month js library.
That 99.99% uptime is achievable with a single dedicated server, plus cloud as a cold backup and some monitoring to trigger spinning it up.
You will spin up the cloud server when the dedicate server is down. It will cost, but only for brief periods of time.
That's sometimes cheaper than two dedicated servers.
I prefer more than 1 dedicated server, but it's definitely possible to get the 99.99% uptime with just 1.
And in some ways the extra diversity of having a different kind of backup is useful, e.g. for burst scaling as well, and for peace of mind not relying on a single provider.
Even better, you can (and should) run the applications in VMs, or in containers inside VMs. They perform well these days, and even a severe application crash won't take down the host.
Then regular application upgrades don't need a separate host server. Instead you're spinning up and shutting down VMs.
You can also migrate VMs between hosts, and this can be done live if you want. It takes some setting up to have working distributed storage and movable networking, but it's possible.
However, you might decide not to bother with that, as starting and stopping instances and having fleets of instances managed by something like K8S is modern practice for other comforting reasons anyway.
Then the only "big" upgrades are limited to host server upgrades and major network reconfigurations, both of which you will do rarely. You might do host server security upgrades (kernel etc) more often, but that costs the time of a reboot unless it goes wrong.
If those are planned or automated, those "big" upgrades can be done with zero user-facing downtime except in the event of failure, in which case it's limited to the cold-spare startup time
So provided you have the backups in place, maybe on the cloud as a cold-spare so you're not paying for it except when it's needed, you can still have that 99.99% uptime for approximately the cost of one server.
You do not need a separate load balancer device for uptime specifically; only if load balancing is something you want anyway for other reasons. For seamless transitions between backend servers, if you have multiple servers running at a particular time, there are other methods.
On occasions you might instantiate a cloud load balancer temporarily to cover downtime during a planned host upgrade that can't be routed around another way. Like with a cloud cold-spare server, you don't pay for it except when it's temporarily instantiated.
But not good for production REST API endpoints in my experience. In the context of AWS Lambda, you are talking a 29 second timeout in API Gateway. API Gateway is fairly expensive. Also, is connection pooling a solved problem?
Many times you can do better with an EC2 instance running PM2 and Express, or whatever. Like, what does Serverless Framework do locally anyway? Express? Why not just run Express?
The best part is, I never have to worry about scaling the underlying resources, if sales brings in large clients. Worst case is I ask AWS for a greater Lambda concurrency limit.
It’s great but you have to put management of usage in place first.
How does a build trigger the serverless function to notify all of you? I kind of assume serverless means it activates on a timer or when you visit a specific URL. So a build script executes it by visiting a URL when it's finished? And the script executes from any internet connected machine?
Why is it better than having a VPS? I currently use a VPS for a few one off scripts. Cron does the timer ones and the URL ones are entries in a nginx config file. Actually that's half true I actually config an app I wrote to do it cause I didn't want to bother finding out how to do shellexecute on nginx
You can do the same with a VPS, but I think VPS cost more because you pay for all the time you're not using them. With serverless functions you only pay for execution time. And with one-off, daily, hourly jobs you end up paying far less, unless your jobs are compute intensive (cron jobs tend to not be...?)
I have seen some VPS these days come in at like $2 a month though, so it's worth comparing the cost.
That framework is separate from the class of services regarded as serverless. The star that kicked it off was AWS Lambda. Serverless colloquially means you are dealing with an abstract service contract rather than a server (e.g. the oldest is S3). This removes patching and other maintenance that usually does not directly support business value. More formally, serverless includes auto-scaling to zero, paying only for what you use, high availability, and other design patterns most outside the large tech houses cannot use at low to no cost.
Your VPS is always on and always paid for. It can crash or get in a bad state. It can go out of date and need patching. It is mutable and more vulnerable. It is limited in resources.
Nike reported AWS Lambda scaling in production at 20K RPS/S (0RPS@0s, 20K RPS@1s, 60K@2s, 120K@3s, ...).
serverless-artillery, with loosened account "safety" limits can scale from nothing to producing billions of requests per second on target systems.
Full disclosure: I contributed to the serverless framework and serverless-artillery. I'm a biased fanboy.
Man1: I'll work on your report.
Man2: I'll watch you work on my report.
To me it's just full of risk and costs where as adding a new function, getting it reviewed and deploying takes all of a few minutes with none of the cost.
However in general what you're doing isn't devops, its 'devs creating a big old mess for someone else to cleanup later'.
All you're doing is trading long term maintainability and quality for 'getting it done right now'.
Sure; it works. I get it. ...but I've also been on the receiving end when that 2-person team has 1 person leave and gets scaled up to a 5-person team, and you have to a) bin everything, and b) write the entire thing from scratch.
I personally consider it a very selfish way to build things. Me first. Me now. ...someone else's problem later.
There are scalable solutions in this space (eg. pipelines, airflow, azure data factory, glue); doing it using lambdas is... not. ideal.
That's not a problem with severless in general; but what you've described isn't any of the things which is good about serverless either: what you've described is the problem with the low barrier to entry with serverless, resulting in production serverless code that skips past 'quality control' to 'let's deploy straight to production!'.
All of our serverless projects are located in a single repository, and each little service gets its own little google doc that explains what it does, what functions it has, and any resources that it uses.
As mentioned below, as soon as an ETL gets multiple steps I agree that pure-lambda is not the right tool for the job (Airflow would be my preference, though we've also successfully used Step Functions)
Doing it 'quick and dirty right now with the tools I happen to know' instead of 'doing some research, then doing it in a way that is maintainable using standard tools and taking a bit longer' isn't 'endangering the survival of the company', it's professional level practice.
Failing to do so is, a) negligent, and b) incompetent.
I'm certainly not suggesting the OP isn't competent; they seem on the ball, and maybe serverless is the best choice for them because (insert reasons here)...but I think its fair to say, if you're using your one tool (ie. serverless) for all your tasks, you probably haven't done your due diligence.
Specifically in this case, using a standard ETL tool to do ETL type tasks... would probably be suitable.
Anyhow; no. Long term thinking isn't what this is about; it's about behaving like a professional when you do a professional job.
You pick the label you want to wear as an engineer, because, your practices define you.
You can call it devops, a mess, whatever you want... but nothing about serverless frameworks is creating that mess. A service is a service, whether it's running Lambda, Spring Boot, or whatever the hottest new framework is... Nothing prevents you from documenting a service. Nothing prevents you from creating basic alarms (it's literally easier for me to create these alarms because of CDK than it is to alarm on the standard service framework I use).
We get it, you don't like serverless. But you're going to rag on OP and question their competence, and then drop some "you pick the label you want to wear"? Come on dude, seriously?
You're strawmanning too here... Can you show me where OP said they are using one tool for all of their tasks? they literally said "for one-off, little processes."
I cant imagine what it's like to be this arrogant.
I’m saying; if that’s your only tool, you’re not doing your job.
“Velocity” is not an excuse for not doing your job to a professional degree. Its a strawman people use to justify doing substandard work (or, often, pressuring engineers who know better, into doing substandard work).
My company's infrastructure was never tended to and was apparently a free for all for a few years. A year or two ago, it was finally decided the free for all needed to end and all the developers saw their SSH root access removed.
My team of 5 SREs is now facing a constant backlash from developers frustrated at the rotten infrastructure, constant bugs, and impossibility to ssh into machines. We're slowly fixing the mess, migrating to Kubernetes and things are starting to get better, but god, how painful it is.
So yea, in its infancy a startup's infrastructure can't be perfect, but every single shortcut you're taking now, you'll pay double down the line. Choose your path wisely.
I cannot tell you the number of times I have implemented "upload your photo and it'll get resized to (profile avatar size from design specs)". It's ridiculous, and it's one of those things that everyone burns time implementing their fun hook into Imagemagick. Now I have one lambda that gets pointed at a new record stream from an S3 bucket, and I'm done.
I cannot tell you the number of times I have implemented "when this user signs up, send them a welcome email." It's one of those things where you construct your email, point it at your MTA, do a ton of configuration, then it may work. Now I have one lambda that gets pointed at a new record stream from Dynamo, which calls SES, and I'm done.
I cannot tell you the number of times I have implemented "clear the Redis cache if a user changes their preferences". You write a clearUserCache hook into your DAO, or you paste it manually into your crud functions, and you always forget something, and six months down the line you start getting bug reports of people's zip code not updating, or something. Now I have one lambda that takes record streams from Dynamo, removes a key from Elasticache, and I'm done.
It's not that you couldn't do this before serverless, of course you could and you still can. It's that it makes that level of code reuse that much simpler. You have all of these helper infrastructure functions that you implement for every single project you work on, and reusing that glue code is so, so much easier in Lambda/GCF/AF/etc.
Email is unfortunately a moving target. What worked before, gradually stops working as the big providers put up increasing obstacles to your own MTA doing a successful delivery.
I've heard Amazon SES also has delivery problems so take with a pinch of salt. But I would hope they generally try to maintain it.
Servers.
That sounds great until you need to add a feature or fix a bug in the reused code. Then you deploy a change to a Lambda function that impacts X other projects immediately, with no chance to test each of them individually.
In a scenario like:
> Now I have one lambda that gets pointed at a new record stream from an S3 bucket, and I'm done.
Ok, so you got AWS set up to fire your Lambda when an object is created in an S3 bucket. You decide you need another "stream", we'll call it, so you start dumping stuff into another prefix. How does one go about testing that the right function is invoked?
A smart person will probably say that they have a dev environment and they manage infrastructure with Terraform. Great! That's probably the best solution there is.
But that still leaves a massive, glaring problem: it's quite difficult to implement any sort of automating testing of this Lambda function setup. In all likelihood, you're probably just pushing a file up to an S3 bucket in dev and watching it run through.
Let's say you made a pass at automated testing, and let's continue with the example of creating resized avatar images. The end product of this Lambda resizing process is probably a different file somewhere in S3. So you fire off the automated test and it fails. How did it fail? Well, if you're lucky, the Lambda function actually had an execution that errored out. Then it's up to you, or your automation, to look up logs in CloudWatch to troubleshoot the failure. What if it didn't error out, and instead just put the file in the wrong place?
This kind of stuff is where Lambda falls over. Running Docker images on EC2 in some fashion puts way more sanity around testing as a whole. You have real Docker artifacts that you ran tests in, not just some zipfile abomination that does nothing to create a good local development environment.
I'm not gonna try and say that it should replace containerized applications (I would choose those 9 times out of 10 given the choice). Unfortunately, there are cases where there isn't a choice or Lambda is still a good choice (like maybe the grandparent's image resize thing, probably depends on a lot of things).
I tend to think of a Lambda as a custom piece of cloud infrastructure. So, in addition to unit tests, I just test them like I would any other Terraform module. I use Terratest to deploy a stack containing the resource under test and a surrounding harness. In this case, maybe the Lambda, a source bucket, a destination bucket, a DLQ, logs, etc. Then execute my test cases, poll for results, do assertions, etc. When it's done, Terratest destroys the stack.
Certainly, I hope you can be the remora to this shark for a long time to come. Just be aware of the benefits and drawbacks of the position you're taking.
The ability to take whatever crazy code with spaghetti dependencies and freeze it into an image and have a cloud provider auto-scale from 0 to ludicrous in seconds is phenomenal ability.
I love cloud functions. They make the perfect Webhooks.
This may take many forms. Pricing models may change. Use cases that gradually see diminishing use may get discontinued (Google Chrome). You might get on some sort of treadmill of having to update details every so often (Facebook API). I can't predict what exactly will happen, but I believe that if your use case doesn't fit "we run a bazillion Lambdas and send tens of thousands of dollars (at least) into Amazon's bank account every month," any service you receive is accidental and contingent.
I use it for similar use-cases GP pointed out: It is a breeze to setup but the best part is there's no devops, no sre (pretty much set-it and forget-it) which is pretty great for something that'd be highly-available yet not be expensive at all even for the smallest of businesses.
Second - with each cloud provider's serverless service, I've noticed that the underlying complexity is shifted into your application and provisioning. They're a great example of leaky abstractions. The name 'serverless' may lure you into thinking that you don't need to think about servers any more, but actually you do have to adjust to the way they've implemented a service that you're using. Think of Lambda cold starts, or have a look at the DynamoDB pricing page. There's still a server running somewhere.
Personal experience: I've found that ECS (Fargate) is a decent step into the serverless realm, if you've been running normal containers on normal EC2s. It does take away the EC2 management aspects, and it's especially useful if you have some applications that are resource hungry, and some are not. It's a cheap way to run ETLs (we use the ECS Operator from Apache Airflow, which itself is on EC2). It doesn't take long to learn, and it's a good place to start. It does come with its own baggage... the 'stack' that you had on a single EC2 is now scattered about; you may need AWS Cloud Map for service discovery, ALBs for networking, SSM/KMS for secrets, Cloudwatch for logging.
That's when you realize, those sticker infested laptops at Devops conferences are actually modern architecture diagrams.
The APIs cost a bit more a month (they’re on quite low cpu/memory) because they’re always on, but it minimises complexity and it’s nice to know that you can move those containers over to another platform if needed.
Actually, re-reading your comment, I'm not sure what you were trying to say. I may be agreeing with you.
---
I think it is possible for serverless to supplant PaaS et al, but I don't know if it's quite mature enough yet.
For stateful systems such as databases, that typically require complications like replication and backup, configuration, testing and maintenance is hard to get right.
But for stateless systems, IMO infrastructure is pretty straightforward.
I'm a fan of OpenFaas: you get the benefits (but not all the drawbacks) of containers and serverless. It's easy to mix and match running 100% locally, or mix-in baked 3rd party components, or running in Kuberenetes or other systems in production. Also, no vendor lock-in.
As the article says, serverless is one of many (many) ways to wrap a quantum of functionality inside an internet-accessible environment. You could have a chunk of python in a serverless setup, a small flask server in a container in k8s, as an endpoint in a monolith, etc. Each of these environments have difference performance, cost and maintenance characteristics, but I think of the star feature of serverless as the light deploy. For companies that don't have a deploy chain, that makes serverless very attractive, but if you've already invested in one it seems less compelling.
I think serverless products have a ton of potential for side projects and hobby projects. It's great that you can just type out python and access it on the internet immediately! That quality just never seemed to be what was holding back SAAS companies.
Not the core of your app, but a tool of great utility. On the other hand, this isn't exactly the Serverless Revolution either.
If you have a server that can handle up to 100 requests at a time, but you're only getting one or two a day, you could probably save money by switching to serverless. On the flip side though, you're also a bit screwed if 1000 requests all come in at once, since, even if you have some autoscaling solution, it probably won't be able to bring new servers up in time. Serverless provides a solution for that case as well, since you have almost unlimited resources.
But yeah, serverless is a tool for solving a certain set of problems. This idea of the "Serverless Revolution" was kind of silly from the start
Not really, since it takes a few seconds to spin up all the serverless instances, so your app response becomes really erratic.
Then you need some magic to deal with database connections from 1000 lambda functions. All of them use their own since they cannot pool.
what platform?
If the next lambda invocation happens within minutes of the previous one ending, you can carry forward a db connection from the older lambda, no magic needed. Just place your db connection object in global scope (nodejs).
Now if you somehow get a spike of requests to some endpoint that goes over your DB connection limit, suddenly all your newly scaled functions fail because they cannot get a connection to the database.
AWS added a service for RDS to deal with this, but it all just feels like a big kludge to deal with a problem that shouldn’t have to exist in the first place.
You will have to analyze your burst rate and the lifetime of each request to figure out the size of your connection pool.
I also am not clear how AWS introduced the problem of not being able to connect to the DB. We've known about the need for connection pools and connection reuse from before AWS was a thing, no?
A process is given a connection from a pool when it requests the connection and holds it until it relinquishes the connection regardless of whether the process was in the logic or data access portions of their lifecycle. It is up to you as the developer to ask for the connection and to return it on an as-needed basis...Which is exactly what you should be doing in a lambda as well.
Just because a lambda container is persistent does not mean the connection given to it is stuck to that lambda even after that lambda returns the connection.
With a connection pool, you don’t need to hold onto a connection longer than necessary, but you are right that you could implement it in such a way that you hold it for life of the request.
How many people have that use-case but not in multiples? How many of those people have such a use-case but would not be better served by platforms such as IFTTT or webhooks on other services?
I guess if we count fargate as serverless, it's much better for this use case.
I'm sure there are other ways of doing it, but the way I do it is using serverless framework (is a nodejs app) together with serverless-python-requirements plugin. The plugin understand poetry (make sure you don't leave requirements.txt because it will use that, you don't need that file if you use poetry).
Anyway it has warts, one big thing is that nodejs developers don't care about norms and standards and constantly reinvent the wheel so for starters you're pretty much forced to create project.json project-lock.json files, which then will create node_modules directory in root of your project. Then you need to start project through npx command (I believe you could still install serverless globally but looks like that's being depreciated).
The node_modules directory will contain over 200MB of javascript code after you install it.
But after all of those things, when you get it to work it is not terrible.
I've created a whole photo sharing website in AWS Lambda, including a complete user accounts system (register, login, forgot password, email verification, profile photos, etc), social aspects, as well as other things like youtube video search and transcoding in the same system.
I also had to create my own Lambda build system because what was out there wasn't what I wanted. I wanted to be able to hit save on a file and have it repackage my Lambda on AWS (using a Lambda to create the lambda) including shared code in 'Layers'. Kind of a "live-reload" for Lambda. It's all working very efficiently.
This was a lot of effort though, but now that I've got a basic system I can extend it to any kind of website.
My photo sharing site for friends costs me about $0.25 every few months, and that is mostly/all the cost of storing gigabytes of photos on S3. No, it doesn't see a lot of traffic, and that's exactly why I chose Lambda to build this site on, because I don't have to run an EC2 instance for my photo sharing site 24/7 if nobody is using it. It's worked out exactly how I wanted, costs me practically nothing to run every year.
And, if I did build a system with a ton of active users, Lambda handles the scaling of that. Another reason I spent the time to create this build system and write all the code to handle user accounts, etc.
This is a bold claim to make without verifying it.
> Another reason I spent the time to create this build system and write all the code to handle user accounts, etc.
But did you also spend the additional time to scale test it and verify it won’t break in confusing and annoying ways when scaling / spiking up? (see the other discussions re: connection pooling above).
The reason a lot of folks are skeptical is that in practice there don’t seem to be a lot of great examples of stateful services using a “pure” serverless lambda pattern but with arbitrary dependencies that actually do scale smoothly and don’t break in annoying ways such that you end up building something more complicated than just using App Engine or Heroku...
I mean, Lambda's ability to scale is a well documented thing. It will happily do more if you ask it to.
The real scaling problem is your wallet, because Lambda is damn expensive once you have any somewhat serious compute requirements.
I'm not sure I understand your point about not needing a deployment chain. I developed a framework (open sourcing it soon) using Jenkins Pipeline DSL for deploying Lambda functions and its very useful and helpful.
Regardless of the target environment the reasons for using a deployment pipeline are still there even if you use a server less platform.
- OS patches and emergent fixes
- Compliance certification
- SSH access control and auditing
- Secret distribution
- Log rotation and storage
- TLS termination
- Configuring and testing auto-scaling policies
- Deployment configuration (rolling deploys, blue/green deployments, connection draining, etc.)Can't wait for codeless, zero code, nocode, whatever name it gets called... something where normal people can build their own ideas without having meaningless discussion about tdd, frameworks and etc, nor the need to maintain codebases, ci/cd and all the jazz
"Low-code" is the buzzword to search for. It's been around a while now.
Of course like all the things the idea has existed a long time, but "low-code" development services seem to be increasingly popular and visible in the last couple of years, and they do seem to be getting more advanced and ergonomic.
Of course, if you already have servers with something like k8s or Swarm running, you might deploy a container there instead.
Cloud Run employs a vendor-neutral container runtime and API (from Knative open source project). It simply accepts any OCI container that can listen on $PORT number.
Similarly, the author talks about serverless not being able to run "entire applications", which again, doesn't apply to Cloud Run. Many people run fully-fledged .NET Core or Java apps on it with a single command to deploy.
Ditto for the author's "Limited Programming Languages" point, Cloud Run is container-based serverless runtime and it runs any language. Furthermore, you don't even have to write Dockerfiles anymore to build containers for plenty of languages thanks to Buildpacks https://github.com/GoogleCloudPlatform/buildpacks.
Furthermore, in my opinion, the serverless revolution is still going on with more services adding edge lambdas support or storage support for edge workers like Cloudflare did with Durable Objects Beta or Workers KV and there's definitely more to come.
Overall I'm not sure what the author was thinking while writing this. It seems they are aware of the intersection of containers & serverless and note that serverless is not just FaaS, but I note an intentional omission of many products that directly respond to his points in the article.
Disclaimer: I work on Cloud Run.
By the way are websockets ever going to be supported on managed cloud run?
Can I just deploy my container to three different places and choose where to point my URL?
I'm right now faced with re-wiring a Python 2.7 App Engine App and feeling the pain of the lock in.
It's just a 12 factor app with a Dockerfile.
What makes Google Cloud Run special between all these is that it's fully managed with an easy auto-scale model and only charges you only for the time requests are being processed.
I think it, and the Knative project more broadly, is the future of serverless.
However, the frequent posts along the lines of "my Lambda function was buggy and now I owe Amazon three billion dollars" definitely kills the romanticism.
You'd be surprised how great it is to pass a parametized query and have it turned into a function quickly.
Doing stupid things on AWS is really easy and sadly billing can be up to 24 hours delayed.
How about we call it auto-snailing. "I wanted to grow fast, but not THAT fast! Slow it down there buddy."
There are lots of posts online that document "talk to amazon and they'll take out the extra charges" but there are even more cases where the people just pay the bill and thus amazon makes money and profit.
Additionally, bill cutting is not that easy. Should they shut down current services? Stop sending in the middle of a newsletter? Delete your S3 Storage? Even then, a large company might accumulate hundreds or thousands of dollars within a second ; even if they decide to cut at all cost, minimal delays might break the limit.
I fully agree that it sucks for experimenting; it is very much the reason I don't have an AWS account. But I can see why they don't have it.
The problem is there's not a clear way to tell the difference between recursive calls and heavy traffic.
AWS components are fundamentally a network of nodes that are sending traffic to other nodes. But there's no unifying language that can model this, so you can't describe what this graph ought to look like. (e.g. you want to be able to say "lambda Foo should be called twice for each event to Bar".)
And that'd be hard to do since most nodes are general purpose computers, like Lambda and EC2, that can send traffic anywhere for arbitrarily complex reasons.
Since you can't describe what it ought to look like, trying to alarm on it becomes a problem of detecting anomalies. You can sort of do it, but AWS doesn't want to do it because they'd be making a promise they couldn't generally keep.
People realising that it's just another tool in the toolbox rather than 1 tool that can replace their entire toolbox.
Much of the cynicism seems to come from that.
I do agree with jason though - open ended serverless things charged per use are no fun. Now you can bankrupt yourself at scale with that bug & the 1000 instances.
Are any of them? Stuff like cloud run is very competitively priced when assuming base case...but by its very nature it can scale up near infinitely. And with it bugs & bills.
The ability to cap things by 2 order of magnitudes would make me sleep much better. (1 would be better).
Currently, FaaS platforms should be viewed like a cron. The majority of the benefit is if the fn is called in the low thousands per day.
Whenever I read these articles, I always wonder what proportion of tech buzzwords from the past decade are just obscure ways of saying "someone else's servers", "automation" and "unnecessarily complicated architecture".
The way people talk about it, I certainly get the impression that a lot of people think that way.
But I agree with you in full otherwise - there is no way that Lambda is going to completely usurp all forms of computing in the cloud.
Things like the Serverless framework[1] help, but ultimately today's serverless architectures feel like they forgot Dijkstra's warning: that the "Go To Statement Considered Harmful." [2]
I'm still big on the approach for the right use cases: variable work loads are a good example.
But ultimately, to be truly useful, we need to see either a) better coupling with non-serverless approaches, or b) a "serverless 2.0" approach that builds in the good parts and fixes the bad parts. (But please don't call it serverless 2.0... those of us who lived through "Web 2.0" will thank you)
[1] https://www.serverless.com/
[2] https://homepages.cwi.nl/~storm/teaching/reader/Dijkstra68.p...
AWS Lambda can run all programming languages via layers.
Cloudflare Workers can run all programming languages that compile to WebAssembly.
"Vendor Lock"
True, but being locked into Kubetnetes isn't a cakewalk either.
"Performance"
Cold-starts aren't a thing for Cloudflare Workers and can be mitigated for AWS Lambda. AppSync and API Gateway don't even have them if you directly integrate with AWS services.
"You Can't Run Entire Applications"
You can and it has been done multiple times. Sure, you shouldn't blindly build everything with Lambda and API-Gateway without measuring. But many services, especially Cloudflare Workers are quite cheap.
Anecdotally, stories of businesses abandoning entire cloud automation projects because they wasted weeks and never had anything to show from it don't seem unusual, so evidently a considerable amount of knowledge and skill is required to get value out of these tools.
I used K8s and serverless tech and while serverless wasn't as simple as some evangelists try to sell it, it was at least a magnitude simpler than K8s.
Even managed K8s, which removed most of the admin plane work with nodes, was still significantly more work to get up and running.
Serverless makes somethings easier and other things harder, so it's advantages over good old fashioned servers are clear cut.
I was quite impressed when I exposed Kinesis via API Gateway, it just worked, although the Kinesis requirement that data be b64 encoded did require some fiddling about with mapping templates.
Sure, a container is better than a serverless function in many ways, but the function isn't what the container is competing with.
https://registry.terraform.io/providers/hashicorp/aws/latest...
The biggest annoyance is unit testing, although it forces certain good habits on us. So for example, since we don't use SAM, we have to do unit testing separately (our code is in Python). But because we can't exactly run a lambda via Python, we have to put most of our code in the layers that our lambdas import.
So we run unit tests on the layers, and the positive is that we can essentially copy and paste the code that our lambdas would run into the unit tests. And then we can keep it DRY because almost all the code is in layers, and that incentivizes us to create generic functions that can be used across all of our lambdas.
For things like unit testing DynamoDB, we use Moto which works exceptionally well.
https://pypi.org/project/moto/
More detail:
Our terraform has infrastructure code for:
API Gateways
Hard-coded API keys we use for testing
DynamoDB tables
IAM roles used by the Lambdas and Step Functions
Lambdas themselves
s3 buckets used for the Lambda zip files and state
Step functions
We did this so we could totally separate our applications and the infrastructure as code stuff since infrastructure changes at a drastically slower rate than applications.
So for example, if we have a service deployed in an EC2 instance, we have to update the instance(s) every month by a certain date. Sometimes, we have to update early if there is a severe security issue. We also have to manage things like security software that has to run on every instance, logging for soc2 compliance, access controls, manage ssh keys, etc. etc.
That's a lot for our engineers to handle every single month. A lot of context switching. It's not so bad if we only have one service to manage. But the actual issue is that we have many services to manage. So every month, our board had about 12 stories related to patching. Patch service X development. Patch service x staging. Patch service x production. Patch service Y development.... and so on.
For serverless, the team responsible for SOC2 compliance gets the reports from AWS without us ever having to do anything. No patching stories, no security software that breaks on occasion, no access management, etc.
All we have to worry about in terms of SOC2 compliance are the basics like principle of least privilege (i.e. lambdas do not have full access to DynamoDB), no public endpoints, we use VPCs blessed by the security team, and that's about it. Things we're used to doing everywhere else and things that we only worry about when writing the terraform.
Now when I add a Lambda, I only worry about writing the Python code for the lambda and copying the terraform code. That's it. There is no patching story for this service.
And yes, of course we're passing the buck. That's kinda what we pay for in the premiums that AWS gets from our massive bill every month. That's kind of the entire point for cloud-based systems, anyway. We're passing the buck on a ton of other things, like not having to manage hard disks, not having to manage network cables, not having to manage a building, not having to manage HVAC, not having to manage backup batteries, etc.
Like CGI or FCGI programs! But with YAML!
FCGI really is a "serverless" environment. Launches an application when called for. Starts more copies of the application if there are enough requests. Shuts down idle copies when not needed. Starts a fresh copy if one crashes. You can even scale with load balancers. Really, that's most of what you need to get work done.
The idea behind serverless includes the fact that you don't pay for what you don't use. So if no requests are served, you owe nothing.
Your FCGI solution shuts down idle copies, but keeps the server running...
FCGI has a server running. And serverless has a server running, as long as its listening for requests.
Essentially it’s a collective pooling which is cheap enough that it can be offered for “free” to the entity creating a serverless workflow.
A better comparison would be "Serverless" is "FCGI in shared environment"... if such a thing exists...
Serverless is part of many different mini-revolutions that have been going on for around the same amount of time, with micro services, rebalancing to have more processing on the client side, managed cloud services within private clouds, and the no-code/low-code movement.
Some applications aren't meant for serverless; they need containers or even on-prem on occasion. But most applications are clients with some basic APIs, and many of those are being thrown up on Netlify and AWS and no one outside of the developers are really noticing.
Can you expand this reference please, Google/Github search didn't give me any obvious results.
Some degree of vendor lock-in is pretty much inevitable no matter what solution you go with.
Limited programming languages may have been an issue a few years ago, but with the support of Java, Go, Python, Javascript to name a few - most of the core user bases are covered.
Additionally, after the introduction of API gateway, most CRUD applications, including those that require async tasks, can be served by the lambda and serverless.
After having used serverless for a while, the biggest turn off for using it (still) seems to be the performance issues. Provisioned concurrency, i.e leaving function containers running, is really just a bandaid and runs counter-intuitive to the original motivation of serverless to begin with.
Another big reason it's stalled is because PaaS has rapidly improved and now deploying a Docker container running whatever you want is just as fast and easy. No need for all the vendor lock-in, frameworks, complex environments and everything else when you can just package up a typical webapp and run it anywhere, even with the same billing (like GCP Cloud Run). In return you get a much better dev environment with all the existing tooling and best practices.
This is not helped by the fact that cloud providers are approaching (and even in some cases surpassing) the complexity and price of running standard compute instances again.
It can be a lot trickier to navigate the nuances of ECS+Lambda than it is to run standard EC2 instances.
Even GCP is not immune, though I find it better in these areas, a common problem I have with cloud run for instance; is container instances which don't have permissions to talk to their linked/associated database, instead you have to instantiate a credentials file programatically inside the container.. which seems superfluous when you're the platform provider and you control both resources.
I've never had to do anything special to access Firestore for instance.
The problem with PAAS and serverless until recently is that they still required a lot of devops activity. Most of that is pure drudgery: setting up networking, dealing with vendor specific weird shit (e.g. amazon's IAM makes everything harder than necessary), micromanaging instance types and trading off overpriced vcpus vs small chunks of memory, etc. Lots of things that you can do wrong, lots of subtlety, lots of poorly documented gotchas, lots of potential bugs, etc. And all of that is needed so you can say "go run this over there" where this is your packaged up docker application and there is months of work by some poor devops person to turn a large amount of poorly integrated tools into a vaguely coherent deployment experience. It's never simple. It's never cheap
I had a decent experience with Cloud Run recently. I was not in a mood or position to reserve 3 months out of my schedule to terraform myself a new server environment (which I actually know how to do). I already had the software and I wanted it running ASAP. So, I felt slightly dirty when I did this but got the job done in a few mouse clicks in their UI. This was shockingly easy after basically spending months piecing together arcane crap in AWS in my previous project.
This got me a cloud run deployment, a build pipeline against my github repository for continuous deployment, and a running service. I tuned the build file afterwards to do more than just docker build & deploy but that was relatively easy (similar to how most CI systems work these days).
They charge for cpu/memory used by requests. The whole thing has been running for a few months now. Less than 4 hours of work to figure it all out from me knowing absolutely nothing about Cloud Run to me having a service up and running. I reconfigured the deployment via the UI a couple of times to fiddle with memory and CPU settings.
I suppose I could sit down and terraform this at some point but I don't feel the need to do that right now. I have more interesting things to do. And technically me spending a day to do that would cost more than running the whole thing has cost so far. That's the point: the cost of devops is out of wack with the running cost for a lot of this stuff.
For the majority of companies out there it's just a tool to overcomplicate your stack and turn it into an engineering playground so you can justify 2-3x the headcount despite no significant productivity increase. But hey, at least your company can now be giving talks about how they solve their (self-inflicted) problems managing all the microservices and throw that buzzword on the careers page.
Prove to me that what you’re trying to do can’t be solved with a Python script, Postgres and a beefy machine. I don’t care about the theoretical benefits of whatever you’re pushing, I want hard numbers. Unfortunately the community has spent quite a lot of its time and effort these last few years making it easy to unnecessarily scale horizontally. I want no part of that.
It's always so strange, being in this space where everyone is dealing with tons of data, and hearing constantly on HN about how "you don't need this level of scale unless you're Google".
That being said, it's always possible to add a bit more complexity if you're not careful.
But nowadays, I just setup my K8 templates with auto-scaling groups on GKE, and it pretty much works just as well as serverless was promised. Compute demand gets seamlessly transformed into compute supply.
I almost never have to think about managing individual hosts. And the cost is cheaper and performance is more predictable than serverless. Plus it's mostly all platform agnostic.
You need to move a few pieces around.
I don't agree that 'once you've gone K8 why bother, just use that' - I don't think K8 is as elegant as the promise of serverless, and it has it's own 'lock in'. Just so happens you may need to have K8s anyhow ... so the pragmatic question then is 'We already have K8s because we have to have it ... so in that context, we can just use it'
If we had a 'severless' version of Docker, i.e. some de-facto standard for it, I think it would obliterate a lot of architectures, just because the promise is powerful.
Developers don't actually want K8s or even true DevOps complexity, we just want a giant computer we can run stuff on and not have to worry.
The reason that the "Serverless Revolution" appears to have stalled is that the customers who have the most to gain from serverless technologies have the least ability to recognize that the technologies they're using are actually serverless. Nor do they care.
In its most reductive form, "serverless" just means "SaaS." And Shopify store owners, for example, don't care about how many servers they run -- they care about how many snowboards they can sell. They could maintain one, twenty, or twenty thousand, or zero -- if the cost to run these things and provide value on top is abstracted away by a few tools it doesn't really matter. So you can use i.e. AWS Lambda to solve their problems, or you could sell ShopifyStoreManagerPlus but they don't care what it's called. They just care that they sold more snowboards.
So "serverless" stalled because the target audience doesn't actually care about the implementation, so the smart companies selling "serverless" solutions all dropped the lingo. AWS Lambda is just another tool in the toolchain; an engine and crankshaft instead of a horseless carriage.
I'll see myself out.
I'm attracted to 'serverless' about as much as I'm attracted to 'version-control-less' or 'testless'
If this was a standard service the standard service would have became overloaded very quickly, allowing the database to remain up but at a degraded performance and the only service that would have been down would have been the one that was getting too many queries. Instead, everything was effected.
I've came to the realisation that if your lambda requires something else to operate, lambda is probably not a good solution.
What's new is that it's become proprietary, instead of being based on open standards. This has made it hard to test serverless apps offline, and made it nearly impossible to move a serverless app from one cloud platform to another.
I do wonder if AWS’s leadership in the area has stalled though.
GCP has innovated with their hybrid Cloud Run to allow running more traditional apps in a serverless way.
Azure and CloudFlare have paved the way to what I believe is the future with stateful functions with Durable Functions and Durable Objects respectively.
Meanwhile Lambda just seems to be gold plating it’s (admittedly great) but stale offering.
AWS has Fargate, run your containers without caring where they run.
In my eyes, the bigger problem with many FaaS products (which most equate with serverless) is the barrier to entry. The setup isn't very user friendly. The code doesn't "just work" like it does locally. Limitations on data size, runtime, etc. cause you to building workarounds in your scripts just to get them running. Not to mention that once you have it all set up, visibility into everything running is a nightmare and only available to the most technical users.
Based on my experience, I'm currently building a [platform](https://www.shipyardapp.com) to try and make serverless data pipelines easier for teams to setup and manage. Would love to hear someone else's perspectives on serverless setups for data management. I know I'm not alone on these existing frustrations.
Batch jobs can't exceed 15 minutes. Memory limit is 3GB. Payload sizes are 256kb max.
If you are used to batch processing large amounts of data, Lambda seems to be the exact opposite of this. I think it is good for "small data, highly concurrent event processing" but this is a very different use case from batch processing data pipelines.
Some of it’s nice. Most of it isn’t. Here are the pros: it’s great for quickly scaling up or down, and we don’t have to worry about physical infrastructure. Here are the cons:
Documentation is inadequate. The docs are often verbose, but outdated or incomplete in critical ways. There also seem to be very few resources/blog posts about how to use big aws features. I suspect there is a secret cabal of infrastructure engineers that hoard this information.
Replication of work is difficult. So much is done through the GUI. This seems really amateurish. How do you quickly capture the state of your entire aws configuration? How do you record and track changes? We have very little insights into our system.
Every microservice is a snowflake. Why do I have to use this apache templating language for api gateway? Weird. Who designed these libraries? Why can’t I rip as much as I want out of a queue? Weird.
The Apache templating language is awful, no question.
- performance: first rule of architecture is that you don't build the entire thing on the needs of high performance...you only address performance as-needed
- vendor lock in: this is the worst reason. Most companies choose cloud vendors and stick with them over many years. I'd love any survey that showed some massive migration between vendors on a regular basis, but it does not exist.
- the monolith: granted, serverless is really best when building something from scratch. Trying to "migrate" to serverless is just a bad idea. Trying to "upgrade" portions is okay, but do it well using the autonomous bubble pattern, making sure you have a solid translation layer
I believe serverless is our best modern architecture, but I also understand that not every company is ready to do the work and/or pay the price of the change in development processes. I'd agree that it's stalled, but moreso out of fear and ignorance than anything else.
Why do you think companies stick with vendors for so long instead of switching?
Isn't this a sign of existing lock-in?
Also, I think lock-in comes in degrees. Some are worse than others. Even if we're slightly locked in right now, we can easily make it worse if we try.
How about simple common sense instead of "fear and ignorance". I have a C++ monolith that serves me just fine already for years. The performance is great and I am able to rent dedicated servers that can serve at least 10 times more requests that I have at a price that is a tiny fraction of AWS for the same work.
Why would I switch and incure more expenses? What's the ROI?
Vendor lock-in is magically valuable, until it is not.
The problem isn’t lock-in per se: the problem is choosing lock-in that will bring surplus extra value for an extended period of time.
https://www.cloudtp.com/doppler/is-vendor-lock-in-keeping-yo...
We’re not buying into IBM mainframes anymore and although there are still customers tied to Oracle databases, vendor lock-in just isn’t that big of a deal anymore.
Is there a way to use one "serverless" framework in one place and then use another in another place?
For example, I love dokku. It is so easy to have the heroku like experience on the cheap.
But I wonder: what happens if I need to scale my application bigger than the dokku server it's on right now? I really like the immediate prototyping possibilities of dokku but the long term plan seems less exact when I read their documentation.
So, is anyone else using dokku locally and Google Cloud Run globally? Or another combination?
This is why I dislike firebase (maybe this has changed). It's incredibly powerful, magic. But, you have to go all in both locally and globally, and there were gaps that caused me a lot of pain.
The tendency would be to use a framework that gives you the best of both worlds, local/dev plus production.
But, why couldn't we use the best solution in one place and the other best solution in another.
Is that a naive question?
As for running actual code, many types of workloads currently do not map to serverless at all and you won't find out until you hit the limits. Anything that needs to read in a lot of data from a file or do a lot of in memory computations are a lot less efficient or impossible in Lambda given the memory constraints. If any single invocation of your lambda needs more than the alloted memory, it'll just fail and you'll either need to implement some heuristics based routing to the different lambdas (same code with the appropriate amount of memory) or you'll just eat cost on every lambda invocation. You'll find you end up doing things like re-writing lambdas from Python into Go to make it easier to fit more in memory. How is that not insane?
Although it's not 100% true anymore, at the time serverless development was often easiest using SAM which ties you directly to AWS CloudFormation which is the biggest mistake any company can make. CloudFormation is easily the worst thing that AWS has ever created. It's fine for infrastructure with lower rates of change but for applications it's an absolute nightmare to the point where it comes up in every sprint retro on some of our teams.
I /won't even go into how crappy testing is. It's a complete afterthought and it shows.
That team has started moving towards ECS/Kubernetes and even running Docker containers directly on EC2 instances in some cases because it's actually saving them time over futzing with Lambda and it's enabling them to do more faster.
The broader idea is a service (or static site even) sending messages to a different service. Given how popular no-code platforms are getting, tiny serverless pieces can help automate operations between these platforms.
The wins also didn't seem great enough to overcome these twin beasts.
Serverless computing is exactly mainframe computing. And by exactly I mean exactly. You see when I was taking CS classes for my degree students bought something called "kilocoreseconds" (kCs) and they are exactly what they sound like, they are a unit measured in the combination of amount of RAM + the amount of compute per unit time your project consumed of the shared computing resource, the mainframe.
You also got a quota of storage with which to hold your projects before you "launched" them on the mainframe with the all powerful "run" command :-). For those of us who liked to explore, it was a very real possibility that you would have to buy more kCs to get through the semester. There was vendor lock in because Doh! you weren't writing operating system code, you were writing application code.
And once you make this connection, you can see all the problems that "serverless" has are all the exact same problems mainframes have for all of the same reasons.
Basically mainframes are "good" when you want to have a small team do all the compute resource management and you have many teams that will want fractional usage of that resource over time. You can developed a closed form solution to the question "what does this compute resource cost to operate per unit time" and you minimize all of the things that affect that cost. Everyone maintaining it are all in one building, all of the equipment is in one room, all of the maintenance is logged and verifiable, etc.
When do mainframes kind of suck? Well when you don't know whether or not you want to use them. Or when you have a different set of OS services you, and only you, need but burden every user of the mainframe.
It is hilarious when you think about it, imagine this giant homogeneous blob of compute, the mainframe, that people disliked because it was "far away" and they didn't have any control, so desktop computers became a thing so that everyone could have their "own" computer, but then networks got bigger so people started virtualizing their desktop compters so IT departments started building bigger and bigger computers to hold many VMs and then because people figured out they could scale by using someone else servers, they started putting their VMs in data centers with a bunch of other company VMs and the data center people optimized by making bigger and bigger data centers, until even the lower cost of a VM was really more than someone wanted to pay for just running a single program so the Data Center types built a layer over all of their machines and services to make it seem like one giant machine. Ending up, nearly exactly back where they started except that now one company owns a mainframe and dozens of companies use it to run their bits of code.
I literally laughed out loud when a low employee number senior engineer at Google proudly announced to the folks that Google was building this new concept of "Data center sized computers." This was like 12 years ago so I'm sure their vision has evolved but it never occurred to them (I checked) that maybe everything old was new again. But I still remember touring computer rooms that took up an entire floor of the building and were basically one computer.
When you get this, you can then go look at where mainframe computers are today. Look at how the IBM Z series can slice and dice itself into logical partitions or LPARs and dynamically become a compute nexus of just the right size for your application, in either bare metal or OS based versions. That is where "serverless" should be looking for inspiration. Modern "web 2" architectures still partition various services into clusters. So you have the EC2 cluster, the S3 cluster, the Elastic Search cluster, Etc. Make the atoms in your world the pivot point and dynamically allocate resource based on your exact needs. Then bill it out in the gigacorepackethours or some other unit of measurement that captures what fraction of the overall datacenter your application consumed and for how long.
Just don't bring back punched cards. They didn't really add value as far as I could tell.
> When to mainframes kind of suck? Well when you don't know whether or not you want to use them. ... Ending up, nearly exactly back where they started except that now one company owns a mainframe and dozens of companies use it to run their bits of code.
I think this reveals that there is an economic distinction between mainframes and serverless, along two fronts: size of commitment and economies of scale. Mainframe time could be extremely precisely subdivided when billing users, but the capital cost of getting to the first kilosecond was extremely high. In serverless settings the capital cost for the first unit of consumption is experienced as zero by everyone except the vendor.
(Yes, IBM runs mainframes-as-a-service, but they don't do it on the cheap).
The second difference is that the hyperscalers enjoy economies of scale and scope that no other organizations can match, including mainframe operations. If IBM is choosing between two different parts for a mainframe, they can pick the more expensive part and foist most of the cost onto their buyers, for whom there is little alternative. Meanwhile, for the hyperscalers, customers are much more price sensitive. More expense falls to their bottom line. But because they operate at such high scale, even tiny improvements are rational to pursue.
Similarly, economies of scope come from the fact that they can pool much larger amounts of variance, which makes their workloads more predictable overall than any smaller organisation can achieve. This is just the central limit theorem mercilessly steamrolling noise into smooth, predictable shapes.
But at least no punched cards.
One of the things that amazed people was the Blekko, the search engine company, operated its own clusters in a data center rather than using AWS. We got a lot of side-eye for that but the numbers were pretty stunning. For our app, the cost of running that data center was about $120K/month, on Amazon (or IBM's cloud) that cost was at least $1.2M/month so 10x higher. What was even more of a pain that getting all the machines in the same racks so that they didn't share "east west" bandwidth with other racks in the data center was a huge ask for these folks, the proposed solution of over provisioning, made it just that much more expensive.
Having said that, there will be someone sprinting breathlessly into the thread to tell me that AWS has lowered prices on X a dozen times and on Y fourteen times and so on. But these decisions are made intermittently, by humans with bonus schemes, and AWS is harvesting the surplus in the meantime.
It feels from the outside as if retail Amazon pays serious, ongoing attention to selling at the minimum possible price, but AWS just doesn't feel the same pressure to do so.
What serverless does is to force scalable architectures.
The price is slim to no control over build and execution environment as a consequence of not taking the underlying architecture into account. For many people this is not an issue. These people will be very content with serverless and should definitely use these types of solutions.
In two jobs now I've been told to do things "serverless first" and both times I've ended up with an ECS solution. My reason always boils down to the other tooling, particularly in regards to HTTP endpoints. AWS API gateway is an absolutely infuriating product.
My current thoughts are if you're working in the AWS ecosystem, with AWS tools, Lambda works fantastically. Using Lambdas to set off alarms and run playbooks has never let me down, and I've done a lot of data processing hanging lambdas off streams and it just works. Really well in fact.
Where it falls over is interacting with the outside world in a sane fashion. HTTP endpoints are a nightmare to set up and deploy in a semi-automated fashion, I've had things that matched what was in the docs that just don't work because of some random gotcha, the undocumented silliness around a "HTTP API Gateway" vs a "REST API Gateway" is another mind bogglingly frustrating thing. Not to mention all the latency issues I've had, which are far too many to blame on a cold boot of the lambda.
Serverless has its niche and it works really well, and real cheaply to boot. But it's not a generalised tool like the loudest evangelists like to pretend it is.
There are certainly use cases for the occasional script run, but that feels like more of job for some cron-like software, and also feels like a far cry from what serverless promised.
The result of it all is cynicism. Serverless takes a lot of work to learn and get good at. The benefit is negligible - no more than deploying a "serverless" thing to an existing web server we use for cross-cuting concerns like exception handling, logging and communication. 10 seconds of automated configuration management, another 10 to deploy. The Change Approval Board meeting to rubber stamp the deployment takes longer.
The breathless hype machine has and always will be just that - hype. Someone's "new, improved, OMG I found this you must use it now" perspective is a pattern that's older than Gartner's hype cycle (which in itself has no basis in fact). It works well in the startup world, but has virtually no place in a business-critical environment, let alone healthcare, defence, telcos or similar.
Little wonder it's adoption has stalled.
Cloud lacks what mobile lacked before the iPhone, a development model. When we finally figure that out, it'll become obvious that table stakes here is an actual full end to end experience with one language and a set of primitives to help scale development for the cloud. That's not FaaS. Its something entirely new. The language will be popular, maybe Go, the focus will be on cloud APIs. Everything else will be become a client. The web era is over, thats not what the cloud is about. Cloud is all about APIs.
IMO what makes something serverless is that you don't need to provision or manage infrastructure, and also no need to worry about scaling as it will go up and down as needed.
Fargate/Cloud Run are probably better serverless examples - you bring a container and config and the cloud provider will handle the rest.
This is why serverless is such a stupid term to begin with, IMO. It's not even an accurate descriptor (the stuff is running on a server _somewhere_), so it really could mean anything.
Something like "cloud functions" to describe Lambda and "managed Kubernetes" for Fargate. Of course, these terms are less sexy, so they'd never catch on.
I like the term "nano services" is it probably the wrong term. But I like it.
We went monolith -> micro services -> micro services + nano services
Everything that adds complexity in a micro service system, adds far more complexity with nano services.
Debugging and tracing is a nightmare. And yes, logging, a lot of logging and instrumentation helps. But if you add too much of it then the whole point of the nano services is lost.
Versioning can get complex, depending on your deploy model. Presumably if you push 100% of your nano services with every deploy you can get around it.
We have about 600 nanoservices running on AWS now across 4 different systems.
My advice, that nobody would ever ask for, keep AWS Lambda for certain special tasks where they are a very good fit.
As always run the numbers, performance, cost to run, cost to maintain. Will the immense ability to scale really help your use case?
In most of our cases, no, they will not.
For some companies it will be a great solution and solve pain-points, save money
And did we mention there's still a server?
Serverless is really not that complicated. But trying to plug a toaster into a super collider just might be.
If you really want serverless, install apache with mod_php and call it a day.
If you already have a server somewhere you could consider spinning up a container instance or kubernetes ( assuming you have at least linux 1 server with excess ram available )
If you deploy your service on kubernetes than you should be able to scale effectively.
Ha. So the name didn't give that away then?
In all seriousness though, I think another avenue for serverless to still explore seriously is declarative application definitions rather than just functions/runtime as a service. Hasura[0] are doing this with GraphQL but I think we'll soon seen (are already seeing?) a surge in these kinds of cloud abstractions that do much more for you than a simple on-demand runtime ever could.
[0] - https://hasura.io/
True, we were only doing a few req/s but Lambda was easy to setup/use and was mostly maintenance free. As is said so often on this site: it's a tool; use it for its intended purpose.
Our development environment was very reasonable: Express or Django locally; aws-lambda->Express/Django on the lambda.
Maybe the revolution has stalled because the emperor in charge has no clothes?
It's a silly term, but successful. If anything, arguing against it made it more successful through mere repetition of the word.
Managing server fleets and keeping them patched is no small feat. How much of the productivity can we gain if we didn't have to do all that with serverless? It's not possible for all use cases today. But can we move from say less than 5% of the use cases today to say, 40-50% of the use cases in another 5 years?
What that leads to is that everyone's code runs on any server. The requires strict inter-server security, authorization and roles.
On a normal server, you can run multiple processes (DB, front-end, back-end, etc) and rely on the fact they're running on the same server to not having to design roles and auth.
In serverless, you need to manage roles and auths for everything.
You see the same issues for data pipes, triggers and queues.
For a few hackathons now my team tries to push a client-side frontend + smart-contracts on a blockchain as a backend solution. This has a UX hit (some downsides of it could be mitigated though) but it made our business logic easier to implement. Depending on the project it has additional benefits to go with this stack. This cannot be applied to every project though and usually the UX hit or the lack of private transactions is enough to favor more traditional stacks.
> One is to optimize your functions for whichever cloud-native language your serverless platform runs on, but this somewhat undermines the claim that these platforms are "agile."
Like this one, don't see how optimizing for a specific language makes something less agile. Then again, "agile" means so many things in the field that I might just be misinterpreting.
I do see a big potential for small and simple projects though. No need to setup instances and worry about networking: just deploy it and go.
I get to deploy stuff "the regular way", that is with a constantly running application, I don't have to re-train my developers, I have extra power in the cluster should I need it, I don't depend on a cloud provider in particular, and my costs are bounded.
Cherry on top: I don't maintain the infrastructure either, there isn't even a SSH daemon running.
It feels like serverless done right. You just deploy your regular Laravel app and don't need any special tooling while developing. Has anyone tried it?
I haven't seen it noted in the discussion below but there's an earlier incarnation of the so-called serverless architecture. Early mainframes could be leased outright or you could buy time on them, in some cases on an ad-hoc basis for one-time batch workloads.
Dare I say that it might even be possible to run twitter at scale with just a simple NoSQL database and some scripts?
That sounds like the opposite of serverless.
Not yet a fan.
That abstraction should allow you to develop serverful, and allow you to go serverless with minimal effort.
For constantly running applications it can be cheaper just to run a server. This fact has been used to avoid migrating to serverless by a lot of orgs i’ve worked with, even when it doesn’t apply to their use case.
At the end of the day it’s either your hard metal server or someone else’s servers (cloud)
It’s then just layers of abstraction to see what is a good fit.
'Serverless' is about billing model therefore purely a marketing concern. It's the very definition of a buzzword, like 'cloud'.
More importantly, app development is way harder since you don't have all OS libraries available to you and cannot easily benefit from previous R&D and workflows.
You basically need to start from scratch or rewrite your app, which simply isn't always an option. There's currently no mature tooling that I'm aware of that would allow you to simply recompile existing apps to turn them into Unikernels (I just discovered Unikraft [1], but I cannot comment on whether it's any good). Edit: [1] http://www.unikraft.org
Debugging is a pain as well compared to more standard development.
Patching, dependencies, configuration drift, security audits/compliance... these are all non-trivial issues requiring non-trivial operations and engineering hours to solve.
It's fine if Lambda is too complex an abstraction for hobby code or a node.js todo-list app, but that doesn't mean it doesn't solve the right problems in the right problem spaces.
The biggest issue we run into is the more servers you have, the harder it is on the database.
We've come full circle.
Netlify, cloudflare, vercel, hasura, heroku, firebase and indeed aws and azure may have something to say about that!
it's true I would need to port the other bits outside the api that use native aws services.
how often do people change their infrastructure once things are working and they move onto building out a platform? I definitely have gone back and forth on the lock-in thing, but I feel like we're always locked into _something_ at the end of the day and it's more important to get something that works and prove it out on the business end.
Not so different and this is from the 70s/80s
I mean, what are we really calling serverless. Many serverless platforms, like fargate allow you to bring a container image, and whatever happens in that container image is non of fargate's business.
I agree with many of the points made though, and admit I am writing this reply mostly because the ckick baity headline got me.
IDK about the others.
pbBouncer[0] in transaction mode tries to solve the problem of limited connections by giving a connection from the pool to a transaction. This has a limitation of not being able to reuse statements; e.g. they have to be deleted before the transaction is released, adding an extra roundtrip for the queries. It also means you must wrap every request into a transaction, if planning to use statements. Otherwise the solution is to sanitize all your values and concatenate a single SQL query. It has the big risks of SQL injections and is against all best practices, but this is what Active Record is doing to solve the issue, and I guess it works for them quite well.
pgBouncer has also the statement mode. Here you can't use transactions and with that, no prepared statements either. So you must then just use the text protocol of the database and concatenate the parameters to your query.
Amazon has the RDS, and I haven't looked into it that much yet. What I kind of think I know about it, is it pins the connections when using statements. Meaning, when pinned, the connection is for one client and cannot be reused until disconnected.
MySQL has the serverless-mysql package[1], which is a client side library killing of stale connections and disconnecting as fast as possible when done. So, basically during the request, it reads the system tables and tries to drop the connections from other sessions. I haven't tried this either, but more I think about it, it really doesn't feel like a good strategy at all. For one, you need a root access to the database for every client, and then you need to modify the state of other connections from the clients. There could be a bug in the client library, that is quite nasty to find and can cause really hard-to-fix issues.
Then the last thing is of course the locking of traditional relational databases. They don't scale to massive concurrency that well; the MVCC chokes when running with many parallel queries[2]. So you need a distributed system, something that can have locking mechanisms that can scale to tens of thousands of sudden connections when the serverless system gets a request peak. Hopefully something with a stateless connection. If somebody knows a system that can do this, I'd like to hear about it.
[0]: https://www.pgbouncer.org/
[1]: https://www.npmjs.com/package/serverless-mysql
[2]: https://15721.courses.cs.cmu.edu/spring2019/papers/02-transa...
> One of the advantages of serverless models is supposed to be that obscure, infrequently used programs can be utilized more cheaply,
No-one: Literally no-one: Infoq: thinks serverless is all about obscure programs.
> Vendor Lock
Well try migrating a massive Java codebase to .NET or a deep Oracle application to Postrges. That's life.
> Functions that have not been run on a particular platform before, or have not been run in while, take some time to initialize.
Boo-hoo mah warm up time. Get over it already.
> you (generally) can't run entire applications on severless systems.
You aren't supposed to - serverless is for developing systems based around microservices, claiming they are intended to replace servers for running COTS products is a disgusting strawman.
> Despite all these complaints, I'm not against serverless solutions per se. I promise.
Promise away. I promise not to care.
> Well try migrating a massive Java codebase to .NET
Java is a language, and an ecosystem of associated technologies. That's not equivalent from switching from one cloud provider to another, or even to your own machines. The equivalent might be, say, maybe switching between JVM implementations.
Today, we're still too focused on what's going on inside the boxes - in the future, more people will realize that the important part that brings value is how the boxes are "plumbed" together. We're already seeing the beginnings of this with serious no-code/low-code platforms, which are ideally suited for these kinds of services environments. (Soon, they won't all be cloud-based, either...)