What a typical serverless architecture looks like in AWS
medium.com
medium.com
I have seen this in practice. A simple CRUD app split across literally 100 of "re-usable" repositories. The business logic is all over the place and impossible to reason about. Especially with step-functions, now the logic is both in the code and on the cloud-level. Each developer gets siloed into their little part and not able to run the whole app locally.
The whole thing could easily fit in a single VM as a Rails or Django application.
The only one that will be happy about this is AWS and contractors because it's a guaranteed lock-in.
It makes sense for the system, and thus for the client.
Just because there are a lot of boxes in the diagram that does not mean the system is complex. A lambda is just a stateless message handler. Half the diagram are lambdas. The rest is just a few persistence services and then stuff related to external interfaces, likr DNS services and API gateways and auth and pubsub.
Any desktop app is way way more complex than this. A dialog box has easily twice the number of handlers.
I also really wish they would not have squatted that term. It should refer to decentralized P2P systems which are truly serverless.
We use GCP and GKS for some things, but we could leave pretty easily. Hardest parts would be DIY Postgres HA and DIY K8S. CockroachDB is almost ready to get rid of the former headache. Haven't looked into the latter headache yet. Of course we really only use K8S for load balancing and HA and there are other options there like Consul and Nomad.
I don't believe that's true at all.
Serverless stuff in gener and function-as-a-service solutions such as the example described in this discussion in particular is in my opinion about placing the needs of the service provider before the needs of the customers.
More specifically, it's about enabling the service provider to waste less computational resources while providing the exact same service.
For example, how many VM instances would you need to put up an API Gateway distributing a couple of HTTP requests to at least one instance of a HTTP server whose only responsibility is to trigger a workflow or update a database? How many instances would you need to keep up to keep a database running that barely has 2 or 3 tables? How many instances would you need to launch to run a batch job that does nothing more than sniff a bunch of files?
Without considering any concern with availability, that's about 3 to 4 instances. All idling, and mostly running support stuff that needs to be there just for your service to be able to handle a request.
This might be your go-to solution for this sort of service, but in the eyes of a service provider that is sinfully wasteful. I mean, half a dozen instances with a utilization rate that barely breaks off 50% just to do the same thing everybody else is already doing?
So, why not cut with all that bullshit and simply tweak a shared API Gateway/message broker/background task/workflow automation/pubsub/database/data store service to do the stuff you need to do?
If you use the communal service and let the service provider manage it with it's dedicated staff, the company doesn't waste half of its computational resources idling by or just duplicating the service everyone is already using.
That's less hardware to power, less hardware to provision, less hardware to maintain, less spiky hardware utilization rates... Less work and less costs.
If you want to discuss lock-in then focus on IAM. Everything else is a way to help the service provider better utilize their current capabilities.
There's nothing about serverless that requires separate repositories or even microservices. I know because I have built a 50,000 line serverless application that is a single repository and deploys functionally as a monolith.
We also don't use step functions heavily, because, like you say, that's basically a cloud-specific DSL that you could write in a real programming language with only slightly worse visibility.
Serverless is, plain and simple, about passing immutable "partial computation" state through events and keeping all mutable long term state in a database.
Make no mistake, serverless cloud services are the greatest building blocks of all time. That doesn't mean you should use all of them, nor use them in the most granular manner.
As always, the architect should make prudent decisions.
You're not really relieving yourself of concerns by using serverless runtimes, just shifting them around. Sure, you don't have to care about managing state at the leaves, but you very definitely do now have to come up with a way to hold state globally with ephemeral/stateless leaves, and it turns out that's... a really hard problem with a really complex solution space.
Not to be overly dismissive or say that one way is better than the other, it definitely depends on the usecase and probably the underlying business model of the usecase. Just to point out that when it comes to engineering abstractions, there really isn't any free lunch. It's all tradeoffs.
I have no idea why some people hear AWS/Azure and they act as if the same architectural principles that have always applied still apply, just on a more granular scale in the case of “Serverless”.
> and they act as if the same architectural principles that have always applied still apply,
They do still apply...
More of a rant about the person who complains about not being able to maintain state at a tier that you never maintain state at in even a traditional architecture.
I’m convinced serverless is the next after containers; but only if it becomes a standard that can encapsulate services in building blocks that you can easily connect with each other and very predictably behavior. We are not there yet, hopefully soon.
This is a standard “onion architecture” with “ports and adapters”.
Sound development principles don’t get thrown out of the window just because you “move to the cloud”.
It’s quite sad how effective cloud marketing penetrated the dev community.
The claim is that getting out of the regular instance-and-EBS-volume architecture avoids cost.
The question being whether "avoids cost" meant Amazon, or your business.
Always assumptions, missing info and costs and poor understanding of it because it’s so damn complicated understanding cost. Consider traffic costing across several regions and you can regularly just piss $20k away from a small architectural decision made wrong.
Estimating costs is Actually so difficult that it has almost entirely replaced the old role of license management under the lie that renting your shit monthly makes that concern go away.
"The difference between theory and practice is greater in practice than in theory."--Herb Sutter in C++ Users Journal some decades ago.
Which causes revenue chasing and profit hunting to “stay alive”.
Not to mention the massive startup credits issued which ensure complete and utter lock-in for the entirety of your business’s existence (successful or not).
This is a terrible practice that is being pushed by the cloud vendors. It’s gotten to the point where now cloud advocates are calling AWS and others an “operating system” your business needs to perform and fulfill its vision.
My company has given a hard push of lowering costs by 20% across all business lines and auditing our cloud usage has been very very shocking.
I am sure as the effects of the pandemic continue, more and more companies are going to realize the same thing.
Just a simple way to provision and snapshot VMs is what 99% of businesses actually need. This is what most businesses have always needed. Many are still using this exact same technology today without any problems.
It's not even exclusive to distributed systems, most computer systems look like this when you actually think about it.
Consider what an application would look like if you mapped out each service and object, it would easily look as complex.
This isn't a new concept, we've had the idea that everything in computing could be represented by little computers communicating with each other since the 60s, Alan Kay called it "object-oriented programming".
The only difference is between the 60s and now, is that instead of mainframes we have platforms, so now we can actually have "little computers communicating with each other".
For example, in an application you can use things like types and functions to make explicit the coupling between different parts. You can iterate and verify changes on the time scale of a few seconds (instead of waiting for a lambda deployment). You can use all the tools for exploring source code we’ve developed as an industry over the decades. And you can atomically change the entire application simply by deploying it.
Of course serverless has its use cases. But evolving a system like this article shows, without most of the tools and guarantees we enjoy when dealing with code, seems like an exercise in pain.
there's certainly _a_ cost; I just haven't seen it be as high as people seem to imagine.
Having been working in one for two years and seen very few "hard" problems come up (and having solved the ones that did), I guess I'm just wondering how it is that our product, with its 5 different GraphQL endpoints, roughly 20 queues, and the equivalent of 100 REST endpoints, has somehow managed to avoid all these supposedly inevitable disasters.
As the team size scales the additional complexity of this architecture let’s teams work More independently.
At least that’s my take on why you’d want to replace a monolith with micro service & serverless.
Whereas shipping as independent services it is harder to cheat because the user profile team doesn’t have access to the billing tables and must go through the billing API.
Of course it is possible for monolith teams to enforce the separation but no team has because it would slow things down in the short term. Then in the long term it’s too complex to unwrap and teams refactor using serverless because that’s what is in style. Of course however nobody is writing about the new problems this approach creates.
I've seen this, a rat's nest of module dependencies, poorly thought-out module/library use, and a lack of ownership and direction around shared bits.
Monorepos also make dangerous things like breaking api changes look easy, or maybe hide them altogether.
Hard problems around ownership, dependency management, and api versions are hard in polyrepos, and I think I'm OK with this. It makes you be more thoughtful.
It is the most magnificent thing I have been a part of so far. It almost feels like we've actually figured this shit out. Looking back at separate repositories from this perspective is like looking back at decrepit playground equipment you used to enjoy as a child.
It does look like what a normal web app would do in say PHP or Django. Is there really an advantage to splitting it into microservices and putting it in AWS Lambda over a load balancer and extra instances when needed? A normal app would have those services split into classes / modules and would run on the same machine (the aysnc tasks being an exception). I imagine my approach would need a lot less ops and coordination between the parts.
And then when you get into the surrounding services and features you really start to feel the benefits. Want to have canary deployments? Want to have a queue of failed events? Want to start piping data into a data lake for future analytics? Need to introduce a decoupling message queue in front of an expensive operation? All of it is almost turnkey.
The advantage to splitting up systems is when multiple teams own part of the service and do not want to be dependent on each other for testing and deploying and signing off on features.
Although given the amount of time I spent this week to be able to send a string from my microservice to a partner teams' microservice and get a different string back, I'm not feeling microservice love today...
“ Q: How can I address or prevent API threats or abuse?
API Gateway supports throttling settings for each method or route in your APIs. You can set a standard rate limit and a burst rate limit per second for each method in your REST APIs and each route in WebSocket APIs. Further, API Gateway automatically protects your backend systems from distributed denial-of-service (DDoS) attacks, whether attacked with counterfeit requests (Layer 7) or SYN floods (Layer 3).”
So I believe serverless/micro services is a way to enforce those boundaries in such a way that a junior developer or schedule pressure won’t allow us to cheat and reach into the internals of another module to get a feature shipped in time. You simply can’t do that when code is running on a different server. So this additional complexity forces separation of concerns that can’t be cheated by schedule pressure or shortcuts.
Of course the downside to this is we have to now manage a different kind of complexity.
As developers I wish we pushed back and simply built our monoliths as a set of independent npm packages, jars, or whatever your language supports. Doing this enforces clear separation of interfaces / dependency management. If we did that then the complexity of debugging and supporting the architecture presented, which strikes me as several orders of magnitude more complex, would be avoided and instead we simply manage at our code level.
If you do test, you really can only test each unit in isolation - testing the kind of setup in the article is just too hard and the tooling for doing automated tests on that kind of thing is not very mature. There is stuff like localstack/samstack/etc but I haven't seen anyone really use it as part of a workable test strategy for serverless AWS.
Consider what happens outside of serverlessland: when people have 15+ microservices each of which needs a different docker container and different data store. Do they tend to do whole-system testing of it? No, in my experience people tend mumble something about "contracts" and test individual bits in isolation.
This is on top of the open secret which is: it's very, very hard to predict what the ongoing costs of a serverless solution will be. Most of the services are priced by usage but the units are small (seconds, minutes, etc) and the sums are too (eg $0.000452 per hour of xyz). Post-project bill shock is a considerable problem.
Source: worked at an AWS partner consultancy on many cloud "architectures" similar to the article.
- I had lots of lambdas written in TS. Things were going out of hands. So I refactored them into three lambdas only. A large one that has a switch statement to execute the correct function, and two other ones for some quite long-running processes.
- The big lambda was executed too many times. Moving it to an actual server would cost me way less. So I did. Plus I could add more security features that would need AWS API Gateway if I chose to stick with lambdas, which is too expensive.
- Initially I used Amplify. And while recently its bundle size has been reduced substantially, it's still big (see https://bundlephobia.com/result?p=aws-amplify@3.0.12). In my case, I only needed the Auth package, which is 50kb (https://bundlephobia.com/result?p=@aws-amplify/auth@3.2.7). However, my options were quite limited and I found myself changing things around to make my app compatible with the way AWS Cognito worked[1] (used by Amplify Auth). I ended up removing it and handling authentication on my own.
A better stack for a serverless app would be Vercel for deploying the app (it's cheap, it has SSL, and it has automatic deployments) and either lambdas or a server, depending on your needs, for handling backend stuff.
Also, use something like aws4fetch instead of the actual SDK if you can. It will save your visitors some kilobytes :)
---
[1]: Email/password auth with Cognito is not the best thing out there. So I wanted to add Google Sign In. At the same time, I wanted to use DynamoDB's row-level access (similar to Firebase's write rules, but more limited). My plan didn't go too well, as I needed the user's Id (provided by Google) when writing to DDB. But Cognito wasn't returning the Id... I spent a couple of days on it, then I figured that I'm just wasting my time.
The number of OSs you patch or upgrade over the course of running this for a few years is... zero. And the time and effort to scale up, run failover tests on your servers, migrate to new hardware ..also zero.
But the servers are just a small part of it. The application services around it make it possible to build enterprise grade scalable services with low ops overhead really fast.
There is so much custom code to essentially fling data between services. I imagine only a small percent of the code and dev time is spent on business logic. These Rube Goldberg machines are my least favourite part of AWS, especially when dealing with at-least-once delivery.
What's wrong with it?
If the appeal is about free usage under 40k monthly active users, you probably don't need an external complex managed auth solution in the first place.
Visit the amplify-js and amplify-cli github repos and search for cognito issues.
Try using the cognito console, it is a litter box of warning messages and exceptions. Try connecting Cognito and Pinpoint if you are in Europe, 2 years and it still doesn't work, but no indication.
See its cloudformation configuration and how many field changes cause a Replace. Yes, replace Cognito and goodbye users.
Until recently user names were case sensitive, and they didn't have basic account enumeration protections.
Its UX is exactly how DynamoDB can do damage to a product.
I understand the wish to keep everything in AWS.
Especially after I read many times thaz Auth0 is pretty expensive. Is this true?
I would have preferred to keep everything in AWS just to make it easier to use one set of credentials (IAM) to manage permissions.
That's when I figured that I'm not gaining anything by using Cognito.
What does your dev/sandbox environment look like? Are you using localstack?
Thanks for your support in the middle of this.
To answer your question: our dev environment is iso to staging and production with AWS accounts for each developer. The cost is close to 0. Deploying and testing a code change takes seconds. Deploying and testing a config change however is still a little bit longer for sure...
There is a TON Of stuff here that is not related to your application business logic, but still has to exist SOMEWHERE for a complete production application.
Let's go through what's actually here that your single VM/Rails application does not/cannot encapsulate:
* CDN
* Domain name and certificate management
* OAuth/identity handling
* File upload
* Workflow orchestration
* Facebook messenger/email integration
And for those complaining that this is untestable, it's...exactly the opposite. ALL Of this is infrastructure as code declarable in a configuration file. Which means you can spin up a beta environment as quickly as production, or hell every developer can spin up this entire environment in their account (FOR FREE i might add - because that's the whole value proposition with serverless which is that you don't pay for something if you don't use it)
What's inside these Lambda function is still standard code you can apply good software engineering and SOLID principles and unit test thoroughly and in isolation.
And if you think workflows, and asynchronous queues, and event buses are overkill - well maybe they are for your use case. If you have 1 API invocation per minute you don't need ANY of this.
But if you're handling any serious scale, this will scale elastically with minimal operational intervention.
Remember - your company is not just your devs but your ops people too. And this makes their life way easier.
I understand the frustration with the 'serverless' buzzword, but let's not miss the forest for the trees. This shit is REVOLUTIONARY. And all people can do is try to one-up how much they could replace it with a simple shell script. The lack of humility in this crowd is mindboggling.
Look I agree - at both extremes (very little usage, A LOT of usage) serverless architectures like this are either overkill or to expensive. But there is avery fat middle (I'd venture >90%) of software projects out there for whom this would be both the easiest and the cheapest architecture (in terms of total cost of ownership) to build and maintain.
Does it look like an unmaintanable untestable mess?
Serverless covers 2 scenarios very well: when you have volume so low, that you don't want to pay for reserved capacity: if nobody is using the stuff, you pay next to nothing.
And also if there is a lot of load, it's very very easy to scale it to cover very high demand.
Don't get me wrong, I absolutely love AWS SAM, but I believe testing serverless applications locally is still an unsolved problem.
That's not to say I disagree with your comment though...
I'm a solution architect in an environment where 20 teams of five developers each are deploying API- and web-oriented code daily, and have been for a year, on a 100 % serverless architecture. The teams are able to deliver MVPs within days, from scratch. We don't pay for any allocated capacity, only usage. Aside from bonuses such as being able to spin up a complete set of services as a temporary environment, in minutes and for free, the production environment is also incredibly cheap.
I can't say whether this stack would have been as effective if we weren't at the size of benefitting from a microservice way of working, with independent team responsibilities and explicitly defined ownerships. But what I can say is that in my 12+ years in IT, I've never seen developers be this productive. Especially considering the complete development lifecycle.
I'm sure plenty of this would have been possible using a traditional architecture, but having 500 lambdas in production and knowing that you can go home and not get a phone call... pretty nice. The reason being that they are small and thus easily testable and securable, in addition to the obvious (fully managed, auto-scaling etc).
We've had to solve issues of, course. Some silly, some unfortunate. But I wanted to provide a counterweight to the apparently popular opinion that serverless is complicated. To me, that's like looking at a car and complaining that it's more complicated than a train. Both have their use cases.
If you find yourself writing many Lambdas just to do simple transforms and piping around data, it's always a good idea to check if the two services can directly talk to each other first.
The docs for most services are really bad. I wasn't fond of learning VTL for AGW either :/
Cost was a big concern for us. Strangely enough, the AWS experts insisted it would be impossible to estimate the cost of the serverless architecture until we deployed it at scale and tried it out. Is this still the case, or has it become easier to estimate costs during development?
In our case, costs did go up significantly after the serverless rewrite, but they also added additional functionality and complexity that made a 1:1 cost comparison impossible.
Serverless was interesting, but if I was involved in another backend project I would want to understand the cost better before diving into the deep end with all of these intertwined services.
I was less involved with that team over time. I got the impression that serverless was quick and easy for the simple server-side operations that could be mapped easily to AWS' serverless building blocks. Of course, this is where they started and showed initial promise.
It broke down as they started taking more complicated server-side functions and forcing them into serverless style. It felt like a square peg / round hole situation that ballooned in complexity just to make it serverless.
If I did it again, I'd have the teams start with the most complicated server-side functions instead of picking the low hanging fruit first.
And, of course, be open to using regular old servers where it made sense to do so. AWS Serverless provided a lot of internal political ammunition because the cloud team could show up in a meeting with complicated diagrams exactly like what you see in this article, whereas previously we just showed single blocks for servers and another block for databases. The complexity did a great job of convincing execs that the AWS experts knew what they were doing, but then of course they were committed to going with the serverless options.
I think they probably considered you a "sucker". Think you're ever going to be able to migrate that app to anything other than AWS?
Obviously, no, we did not expect to migrate our AWS Serverless backend to something other than AWS.
Avoiding vendor lock in makes sense in few, select circumstances.
Avoiding vendor lock in for the sake of avoid vendor lock in leads to over engineered services that take longer to deliver because people are going out of their way to avoid using vendor-specific tools.
In all of my time as an engineering manager, the number of times that avoiding vendor lock in has solved more problems than it caused is still zero.
Use the best tools available at your disposal to get the job done. Cross the vendor-changing bridge if (and only if) it becomes a requirement.
I have seen two businesses fail under the weight of AWS billing and be unable to get out from underneath it before they ran out of funding. It’s definitely not something to look at lightly.
How many companies fail because they run out funding - period?
Congratulations, you found two examples out of the millions of companies that exist in this world, both with astronomically greater funding and scale than anyone you or I are likely to work on. Yet people continue to obsess about these unicorns and assume the choices they've had to make are clearly the right choices for everyone.
The greatest problem with modern devops is peoples delusional ideas over the scale upon which they're really operating and where they're really going to start feeling the pinch. I've come across way too many tinpot "we're netflix!" setups that crumble under their own unmanageability as soon as they have budget and staff taken away from them for a couple of quarters.
The reason for managed cloud is the same reason for using any other vendor. To let you focus on your core competencies that give you a competitive advantage. Do you also think companies shouldn’t use Github, Microsoft, Workday, Oracle, SalesForce, Atlassian, etc.?
Do you know how much infrastructure you can buy for the fully allocated code of one employee?
Maybe they know something that you don’t know...
Seeing that even Amazon admits that only 5% of Enterprise workloads are on any cloud provider, who is arguing that it is always the right choice?
I command a higher than (local) market salary for an individual contributor in no small part because of my expertise when it comes to AWS, but I’m the first person to tell someone who asks me for advice when it doesn’t make sense to bother with the complexity and cost of a cloud provider and just use a colo, VPS, or just use AWS Lightsail (AWS’s answer to companies like Linode).
$10k does not get you very far with AWS.
How long have you stuck around to find out? Never had to undertake a massive task to migrate an existing system that's closely bound to its legacy platform? Because that's the other side of the coin.
I also highly suspect that number is not zero, you're just not counting the obvious decisions people make every day that avoid lock-in because they seem like common sense. But I suspect you're not building everything using ColdFusion, though, or doing your version control through Perforce...
> Use the best tools available at your disposal to get the job done.
Behind this phrase lies the fallacy that there is any such thing as "the best tool" and the implication that anyone doing anything different is a clear idiot. In reality, everyone uses "the best tool", it's just they have different metrics by which they measure it.
I used AWS for a recent personal project. Lambda was not the right approach for me and primarily due to cost. I had an idea of how long a request takes to process (ball park estimate) and the expected throughout (requests per second). From that I knew that at minimum, I would be paying $x a month, and as the execution time of my code goes up or I get a spike in my traffic volume, it would only increase from there.
If you’re at scale, AWS Lambda blows up real quick.
If you use the standard APIGW solution where it routes to your various lambdas, you would have to rewrite code.
Lock in or no lock in, there’s also a cost of writing the deployment scripts and setting up your development vs. production Lambda stack. And if you are using Lambda, you’re going to end up making various decisions to keep your costs low. For example, I was working with a JVM stock. Spring is too heavy, and even Guice has some execution cost. So on Lambda, I would use Dagger... a decision made purely because I’m operating on Lambda.
Building an architecture on Lambda requires certain decisions to be made up front. Saying that you can just lift and shift later on is very simplistic and will be costly depending on what you’ve already done on Lambda.
I would always say estimate your costs first and think about the long term picture about your request rates, patterns (spikes), and growth rates.
Changing your deployment to use Docker/Fargate is creating your Docker container by copying from the same output directory to your container in your Docker file, pushing your container to ECR, and running a separate CloudFormation deploy to deploy to Docker.
I’ve deployed to both simultaneously.
There are no decisions to be made up front. You create your standard Node/Express (for example) service and add four lines of code. Lambda uses one entry point and Docker or your VM uses the other.
It’s simplistic because it is simple. Java is never a good choice for Lambda.
What about writing CF to begin with for all your Lambda functions and setting up all their permissions to work alongside whatever other resources you have?
>Changing your deployment to use Docker/Fargate is creating your Docker container by copying from the same output directory to your container in your Docker file, pushing your container to ECR, and running a separate CloudFormation deploy to deploy to Docker.
Nobody said anything about Docker or Fargate. This has nothing to do with the topic. My problems had no need for using Docker. Do you randomly pick a tool kit and just hope it works out? I suspect you're not really working at scale and costs aren't a concern.
>I’ve deployed to both simultaneously.
Great job. We are all proud of you I guess, but this is not how engineering works, especially when you're building at scale.
>It’s simplistic because it is simple. Java is never a good choice for Lambda.
It's actually not simplistic. Software engineering is about being able to make tradeoffs. Not everything is a CRUD/web application. It's funny you keep going back to Node/Express as your examples. Have you looked into Express and what dependencies it has, or do you just randomly pick the new framework of the week? I don't mean this as an attack, but I'm more than slightly annoyed that you keep referring to Node and Express, when that has nothing to do with the problems I'm trying to solve.
The funny thing is you've written out all these lengthy posts but never stopped to ask about whatever constraints the system has. You started with the solution and are basically now looking at one of my constraints (JVM ecosystem) and are saying it's not a good choice.
In the real world, we often don't have the ability to randomly swap out a language. There's enough properly tested code that already exists or a massive code base that entire teams are supporting. They're not going to drop that just so they can go pick up a shiny new framework or product offering. LOL.
There are no “all of your lambda functions” with the proxy integration. You write your standard Node/Express, C#/WebAPI, Python/Flask API as you normally would and add two or three lines of code that translates the lambda event to the form your framework expects. I’ve seen similar proxy handlers for PHP, and Ruby. I am sure there is one for Java.
As far as setting up your permissions, you would have to create the same roles regardless and attach them to your EC2 instance or Fargate definition.
But as far the CF template. Here you go:
https://github.com/awslabs/aws-serverless-express/blob/maste...
Just make sure your build artifacts end up in the path specified by the CodeUri and change your runtime to match your language. I’ve used this same template for JS, C#, and Python.
Nobody said anything about Docker or Fargate. This has nothing to do with the topic. My problems had no need for using Docker. Do you randomly pick a tool kit and just hope it works out? I suspect you're not really working at scale and costs aren't a concern.
You’re using your standard framework that you would usually use. You can use whatever deployment pipeline just as easily to deploy to EC2 with no code changes.
Great job. We are all proud of you I guess, but this is not how engineering works, especially when you're building at scale.
Actually it is. Since you are using a standard framework choose whatever deployment target you want....
Have you looked into Express and what dependencies it has, or do you just randomly pick the new framework of the week? I don't mean this as an attack, but I'm more than slightly annoyed that you keep referring to Node and Express, when that has nothing to do with the problems I'm trying to solve.
I also mentioned C# and Python. But as far as Java/Spring. Here you go.
https://github.com/awslabs/aws-serverless-java-container/wik...
The funny thing is you've written out all these lengthy posts but never stopped to ask about whatever constraints the system has. You started with the solution and are basically now looking at one of my constraints (JVM ecosystem) and are saying it's not a good choice.
I’ve also just posted a Java solution. The constraints of Lambda are well known and separate from Java - 15 minute runtime, limited CPU/Memory options and a 512MB limit of local TMP storage.
In the real world, we often don't have the ability to randomly swap out a language. There's enough properly tested code that already exists or a massive code base that entire teams are supporting. They're not going to drop that just so they can go pick up a shiny new framework or product offering. LOL.
My suggestion didn’t require “switching out the language”. Proxy integration works with every supported language - including Java. It also works with languages like PHP that are not supported via custom runtimes and third party open source solutions.
Feel free to keep writing more paragraphs. You’re trying to solve problems that don’t exist.
It seems like you start from solutions and hope that the problem will fit. Lambda seems to be your hammer and you write as if everything is a nail.
Good luck!
So lambda seems to be the solution and my “hammer” even though I mentioned both EC2 and ECS? Is there another method of running custom software on AWS that I’m not aware of besides those three - Docker (ECS/EKS), Lambda, and EC2?
I've re-platformed so many projects these past few years because of these issues.
I wish more would take that to heart instead of marketing for Amazon.
Lambda -- Serverless in general -- is fantastic when your application has any amount of downtime (on the order of 10 seconds). If there is even the briefest moment your system can be turned off, you'll save dramatically over a traditional server model.
I have seen more then one project fail when they use lamdas to serve apis. This is because of the cold start problem, but also because lambda does have scaling issues unless you work around them. By scaling I mean greater then 1000 tps.
All of the services performed within lambdas SLA but failed to meet the requirements that the project had.
The solution was always to wrap whatever function it called in a traditional app and deploy it using a contanerized solution where response times dramatically improved and services became more reliable.
The idea behind lambda is good, but it can't beat more traditional stacks at the moment. One could argue that most projects don't have these requirements, and I would agree, but the marketing behind lambda doesn't make that clear.
I personally don't understand who the target audience is, as small apps/companies very rarely (if ever) need to scale and large enterprise companies are probably better off hiring a sys admin and getting their own dedicated servers, which would lead to a lot more flexibility, performance and reduced costs.
Or was your comment sarcastic and I didn't get it? Whenever I tried to run the same gulp tasks on a different machine, it almost never worked first-try.
Moreover the cloud providers should make it harder for an uncontrolled "block of architecture" to accidentally spend too much money. The focus seems to be on "always available", but depending on how fast it's spending your money, it might be better if it crashed.
Expanding further, perhaps each "block of architecture" should have a separate LLC dedicated to it to control billing liability. Incorporate your "lambda fanout" to keep it from bankrupting your "certificate manager" when it goes haywire, and let it go out of business separately.
What’s harder is when your app hits an Amazon limit and you have to figure out if it’s hard or soft, how quickly your TAM can do something about it, etc.
The over use of microservices and cloud computing is terrible.
I've used Lambda and I fear maintenance burdens due to the underlying environments migrating; they'll keep the runtime for a running function for you but if it's old I don't think you can update that without having to upgrade runtimes.
Also you get billed in 50ms increments at minimum, and so beyond that there's little incentive to make a function faster / a lot of incentive to stuff more logic into a function and make it less composable.
I think I'd rather define an AMI w/ Firecracker using Packer, and stand up a box with Terraform, and write logs to CloudWatch, and have a much more predictable billing cycle / greater control.
AWS supports runtimes as long as there are upstream support security updates available [1]. If there are no security updates anymore, it's a good idea anyway to update to a newer version. If course you don't have to do that if you host it on your own, but it's a good idea nevertheless.
Also mind that most of the runtimes AWS deprecated already are old NodeJS versions. That's apparently because NodeJS only offers 30 months of support for their LTS versions [2]. So choosing a language which offers longer support (for example Python which offers 5 years of support [3]), might be a better choice if you don't want to update your software regularly.
Also you could still provide your own runtime, which wouldn't be deprecated at all via custom runtimes [4].
> Also you get billed in 50ms increments at minimum, and so beyond that there's little incentive to make a function faster / a lot of incentive to stuff more logic into a function and make it less composable.
You get billed in 100ms increments [5]. While I'd like to see smaller increments as well, I encountered more money thrown out of the window by overprovisioning memory (and therefore CPU) for AWS Lambda functions, than you could possibly save by optimizing for billable duration.
[1]: https://docs.aws.amazon.com/lambda/latest/dg/runtime-support...
[2]: https://nodejs.org/en/about/releases/
[3]: https://devguide.python.org/#status-of-python-branches
[4]: https://docs.aws.amazon.com/lambda/latest/dg/runtimes-custom...
I do wish there were more examples of Lambda functions, like a marketplace of some sort. I'm using NodeJS only because certain tutorials offer examples in node.js (like static site authentication): https://douglasduhaime.com/posts/s3-lambda-auth.html
I actually logged into CloudWatch just to see that point on 50ms vs. 100ms and yeah you're right, it's 100 not 50. I'm pretty new to Lambda, and wonder whether the billing insights dashboard will give me the breakdown b/w duration and memory.
Thanks!
The company grew and the product saw heavier use. The original monolithic server became slow and problematic. Hosting cost became very high. Separate teams ran into conflicts while working on the server.
So we have incrementally switched to an FaaS-centric architecture, similar to the one in the article. The new design makes it easier to get excellent stability, performance, cost and observability. Now that the development team is a little bigger, FaaS/microservice design makes it easier for the teams to work without blocking each other. I’m sure that same benefits could be achieved in a traditional monolithic server with careful design and control of the development process, but it seems easier and quicker to accomplish with FaaS.
Indeed, a simple REST API sometimes requires dozens of handlers. One might worry that this is complicated because each handler is somehow running in a separate container... but that’s really the cloud provider’s problem. Execution is very finely segmented, while logical division of the code-base is generally free to follow natural domain divisions. It’s pure benefit to us: endpoints have excellent isolation. We have excellent control for behavior in overload scenarios.
FaaS handlers tend to be simpler than each action in the former monolithic server. With simplicity, it’s easier for us to ensure handlers follow a common set of best practices. With simplicity, consistency, and fine-grained isolation it is easier to reason about the performance and capacity of the system. While there are scenarios where FaaS is prohibitively expensive, our use stays within its sweet spot and is significantly less expensive than running the previous monolithic server.
A downside of FaaS is that it quickly became impossible to run the entire system on a developer laptop. We really need to run things on the cloud infrastructure for any kind of integration test. So far, it has been hard for us to provide developers or test infrastructure with independent Integration test sandboxes. Vendor lock-in is a problem: it will be hard for us to change cloud providers if ours becomes problematic.
This can easily be 1-2 very simple services.
I'm starting to accrue a list of projects that started out looking like that, with product owners swearing blind that was the sum total of all they could want, but ended up becoming much more complex. Then you get into the pain of sequential lambda cold starts, etc and before you know it everyone thinks you're in too deep to turn around and the pain just won't go away.
Just one more AWS service and all our problems will be solved...
Sure, if you want to inflate developer costs. Isn't a huge portion of the cloud argument that people are expensive, infra is cheap so who cares if you spend a lot on AWS as long as you have fewer sysadmins? If you suddenly care how much you're spending on infrastructure but, apparently, don't care how much you spend on developers to work within such a convoluted system, why not exit the cloud at that point?
At this point, it might be feasible to reclaim "serverless" as the domain of p2p systems, as some others here have suggested. The non-p2p folks have clearly given up on the "it's too complex to reason about" line of argument, so now might be the time. :)
Also, this "simple is better" theory has been proven to be true again and again in a vast range of disciplines and use cases.
I can give you just an opposite anecdote - and from personal experience - not from what I found on the internet.
We have a lot of microservices that are used by our relatively low volume website, but also used by our sporadic large batch jobs that ingest files from our clients and our APIs are sold to our large B2B clients for their high traffic websites and mobile apps.
When we get a new client wanting to use one of our microservices, usage can spike noticeably. We have deployed our APIs to both Lambda for our batch jobs so we can “scale down to 0” and scale up like crazy but latency is not a concern and we host them individually on Fargate (Serverless Docker). Would you suggest that it would be architecturally better to have a monolith where we couldn’t scale and release changes granularly per API?
I've seen it time and time again. Anecdotal of course, but we're all just shooting off anecdotes on HN anyways.
But how pray tail did I “construct a strawman with a one word reply” - “details?”
And then when I asked for personal experience a random article on the internet was found. At least I was able to speak from personal experience. I can also go into great detail about best practices as far as deployments, logging, security, troubleshooting, testing (functional and scaling) and tracing without citing random articles....
Serverless isn’t going anywhere until we get something like kubernetes for functions.
What appears to happen is the "get going coding" part is dramatically accelerated, that saves on the order of days of waiting; a few weeks/months if new tech stacks the organization hasn't worked with before are in play. But man, once in production, I'm still puzzling over how to manage the operational costs.
There is a huge disconnect in the business community in their perception of cloud that I'm collaborating with. They're taking the delivered "get coding fast" benefits which do save time and money. They see we're only paying for infrastructure the moment devs start hammering fingers to keyboards so to speak, and not waiting around for teams to stand up that infrastructure and paying for all of that while it's not delivering benefit. They take that experience and are applying the same expectations to the operational, production support side.
I don't know how others are doing it delivering on the operational and production support benefits, but nowhere near the same scale of savings are happening so far for teams I work with. What I'm experiencing is while there are marginal savings on the nodes of services (the actual service itself), we're adding a lot to staff payroll at high-end pay scales and spending a lot of their time on all the connections and complexity between them; between troubleshooting and adding new capabilities (more troubleshooting), it is eating up a lot of expensive staff time.
I'm pushing for greater automation through DevOps, but that's encountering lots of resistance now with the economic outlook because DevOps generates a net increase in absolute payroll even with a marginal decrease in operational staffing once the automation is burned-in. In fact, because automation takes away the mundane operational aspects, we end up having to cost-justify far more capable and expensive operational staff that can handle the corner-cases that emerge as business-as-usual activity, after automation ate away all the usual cases. While headcount goes down, total payroll goes up, and budget-watchers resist the idea that they're saving overall by delivering more stable, consistent, quality service with more feature points per dollar, and if we continue to carry on the old way our costs would be easily triple to order of magnitude more. This is a marketing/education perception problem to work upon nearly daily.
I love working with the cloud tech stack, but at the end of the day I have to show my business stakeholders it really did save net budget. For judiciously-targeted and considered projects where I've had a lot of input tactically it has worked really well for them and me. For strategic-level "Cloud All The Things!!@!%!" initiatives, the cost savings get way murkier to discern; this is the perception problem on steroids. I'd love links to detailed readings from people who have been in the trenches of successful strategic-scale "we went to the cloud whole hog" efforts, describing the traps and pitfalls to avoid and how they showed a net benefit.