Serverless Best Practices
medium.com
medium.com
- It has autoscaling built in now but even that leads a lot to be desired, autoscaling is slow and your lambda functions will error out for significant amount of time(I have seen it happen upto 5 minutes) for 5X spikes until it scales up the IOPS. This fills up the queues quite fast and can lead to patterns where your spike increases because of failure retries. You can control this somewhat by limiting lambda invocations though. - Limited number of scale downs allowed. - Scale down doesn't completely reduce the throughput limits. I have seen situations where the write load went to 0 on certain times but scale down reduced it to something like 400 IOPS which can be a huge cost drain.
I am still looking for a DB which has done this well. Essentially what I want is compute layer separated from the storage layer in the database well enough so that both can be scaled independently and the compute layer can react to quick changes in load patterns. Does anyone here have any recommendations?
As long as there are machines/containers to be started, there will always be some latency, though we can expect it to improve in the next years.
AWS now has the Aurora Serverless, though I don't have any real-life experience with it's latency yet (VPC requirement is a downer for now). https://docs.aws.amazon.com/AmazonRDS/latest/AuroraUserGuide...
Yeah, I meant that only. DynamoDB and Aurora Serverless are examples where they are decoupled and can be scaled independently. So, I wanted to know about more products like them but ones that scale with load really well.
> AWS now has the Aurora Serverless, though I don't have any real-life experience with it's latency yet (VPC requirement is a downer for now).
I have tried Aurora Serverless and I like it but it also plagued by the provisioning problems for now. The lack of options to customize database parameters was kind of a downer. Also for some weird reason they only support MySQL in t type instances.
> (VPC requirement is a downer for now)
BTW, you have an use case where you need to use DB from outside VPC and you can't do peering?
For example in traditional MySQL setups to scale up the compute capacity your server has to be unavailable for some time.
We should ask for the same thing with rdbms's.
Lambda in VPC have unacceptable cold-start for interactive applications. This is partly why DynamoDB use is so popular with Lambda, it does not have to be in VPC.
Maybe I'm just using it wrong. I wish Dynamo came close to its apparent potential.
Maybe, but you might also be using it for the wrong thing. Any business logic that requires table scans (which are perfectly fine on RDBMS), is going to suck in any document database. The way you win at Dynamo is by making sure that one user operation is as close to one Dynamo call as you can make it. If your workload can't be designed to suit that paradigm, then Dynamo will only bring you pain and suffering.
That said though, migrations are terrible in Dynamo. Altering a table will often involve creating a new one, backfilling it, and then cutting over. Which you have to implement your own logic for. Also, think the GSI limit is 5? In practice it's actually 4, because if you want to recreate a GSI, you need to create a new one, and then drop the old one, so you have to keep one spare to accomodate that process.
> The biggest point to make here is that serverless architecture may well require you to rethink your data layer. That’s not the fault of serverless.
Well, it is the fault of serverless. It’s a shortcoming. Own it. The trade off may we be worthwhile, I just can’t tell yet. I’m trying to wrap my mind around serverless as regards database-backed apps, but obviously it would be preferable to be able to use the RDBMS of my choice.
1. Aurora Serverless RDBMS (In pilot)
2. Lambda function property caching (for which we cache connections)Also, can you elaborate/ point me to info on 2?
Background info: A lambda that spins up stays live for a while, and can continue to handle requests. The function handler is called again, so anything scoped to it, will disappear with the function, but anything scoped outside of the function will persist.
Ergo, move your connection creation outside of your handler.
Realistically you may need some abstraction to properly close and reopen the connection in the event of failure (so you can't have a 'bad' lambda, where something has happened to the connection and you can't fix it without deploying code again, to trigger a lambda refresh), but for the most part it just works.
If you absolutely must hit an RDBMS, be quick about it to free up the connection, and definitely consider if the latency matters that much; a few dozen ms more to initialize and close the connection is going to be a LOT cleaner.
But that said, how badly it penalizes you for scaling connections up depends on the DB.
If there are other parts of your application which are consuming the same data but in different patterns, it might not be necessary that systems like DynamoDB work in those cases. You might need to make several secondary indexes in those cases and all of them require provisioning of their own. The other alternative being setting up some ETL layer which transforms the data and sends to appropriate sinks, from where they can be consumed; but for small services that adds more moving parts and cost to the system.
Crappy application-servers pump queries like nobodies business, so you need to shave every byte of them you can.
Beyond that, MySQL for example pre-allocates a fixed amount of memory for read, join and sort buffers per connection to efficiently handle different types of queries, though the buffer sizes are all tunable to optimize performance. There is also some OS-scheduler overhead involved in creating and running a separate thread per connection. A thread-cache optimizes the case where lots of short-lived connections are being created. (I dont have direct experience with PostgreSQL but I read that it forks a separate process per connection, which would presumably incur a significantly higher overhead than threads.)
Beyond this, any more specific discussion requires quantifying connection 'weight' to dispel/confirm your superstition that 'all the predominant RDBMSes are heavyweight'. In an apples-to-apples benchmark, it's quite possible that the connection-weight of some RDBMSes might actually be lighter than other database engines. I dont know myself, but would be interested if anyone has any such data to share on this.
The problem really boils down to the tight coupling between the HTTP layer and the rest of the stuff all the way down the stack. If you assume connections are always long-lived and stateful, you can make performance and memory optimizations by reserving memory buffers and threads for the sole use of a single connection, and using threadlocals for storing session data.
But if each connection gets its own thread (or thread pool), idle connections lock up resources and can prevent new connections from being opened, because all of the threads are currently allocated to existing connections. And setting up every new connection is a bit expensive and wasteful if it's only going to be used for one request.
I think the industry trend toward transitory connections, and even transitory application instances a la Kubernetes pods, is a good thing for a lot of reasons. If nothing else, the knowledge that connections can be short-lived leads to more fault-tolerant software because the business logic layer cannot depend on assumptions about the HTTP layer. The big downside is that a lot of older stuff isn't really suitable for that kind of world, serverless or not, and it can be really hard to refactor.
Maybe DynamoDB works better with this model, because it is build up without the idea of maintaining connections?
Well, I don't know much about Oracle and the other players, but AWS at least tries to own it with Aurora Serverless.
E.g., the sections could become:
Function Don’t Compose Well
Functions Don’t Connect to Data Stores Well
Otherwise Unnessary Queues are Needed
Based on this article, it makes it sound like current serverless isn’t a very widely useful tool.
But, I think it’s used a lot more than necessary for end user facing APIs they need fast responses.
If you really insist on calling your entire stack serverless, do what we do and run these servers as ECS tasks via Fargate. That is plausibly serverless, albeit long-running. You get all the perks of a serverless environment with all the perks of something like beanstalk (without having to patch a server!). The drawback in this scenario is that running ECS tasks via Fargate don't provide as much flexibility in cost.
We are successfully running Lambda functions connecting to AWS Aurora MySQL using IAM authentication. It has some quirks but after figuring it out, works well.
Just run a connection pool with only 1 connection and expect some additional latency on cold start, though it's negligible compared to loading libraries, runtime, etc.
> Just run a connection pool
You just described the high-level architecture of what I believe is the majority of web application servers and the applications they run: an RPC layer over database accesses that run on persistent connections.
> Each function should do only one thing. The problem with one/a few functions running your entire app, is that when you scale you end up scaling your entire application.
Sure, but is that a problem in general? Same thing can be said about a monolith vs microservices, and there is always a trade-off. A function with a larger code-base makes the code easier to navigate; may be easier to deploy. First reason about / compare cold-start times before making the decision to break the function.
> Functions don’t call other functions
Is the cost of calling the other function negligible (e.g. the functions are called rarely)? If yes, feel free to do it.
> Use as few libraries in your functions as possible (preferably zero)
Security risk is another discussion, not related to serverless. Worry instead about the cold-start time. If it's short enough for your use case - do use libraries. Should be no problem for Go, Node is usually OK. Java - native compilation is slowly approaching.
> Avoid using connection based services e.g. RDBMS
If adding another library is not a problem in your case, then why not? It's not too difficult to create a pool with a single connection. As for security, there are some new options like AWS Aurora IAM authentication. No VPC needed.
I think this point may might also be about reducing complexity, not just cost. Pushing to a queue instead of directly calling another function decouples the two functions and might make it easier to reason about how logic flows.
The question is more one of errors, I think.
You pump your results in a queue before you let another function do the next step, if one function fails it's not the whole process that is lost.
Step Functions, for example, lets you chain functions and does retries for you if one function failed.
This one will get me into the most trouble. A lot of web application
people will jump on the “but RDBMS are what we know”
bandwagon.
It’s not about RDBMS. It’s about the connections. Serverless.
works best with services rather than connections.
Guess what a service call will use to do its RPC...? a connection. This advice makes no sense. Maybe you don't care about the response to your service? That's different from saying "don't use connections". (I doubt most serverless RPC calls are over UDP like protocol)If you don't intend to scale beyond what a server can offer (which is a lot!), you probably shouldn't be "serverless" in the first place.
Use RDS for DB servers and just use Elastic Beanstalk. It takes most of the complexity out of deploying load balanced web servers.
Serverless for app servers is sometimes fine if you don’t need real time responsiveness like working with queues.
However, 99% of the work that I've done involves users hitting buttons and us responding to them synchronously. In these scenarios, I simply can't figure out how queues (and chains of serverless functions as advocated by this blog) are supposed to work (if they are at all). There seem to be many ways to solve this when the queues are all flowing freely, but as soon as there's any sort of pressure on the system these things all look to fall down.
Looking at the amazon booking flow as an example -- it appears that they always show a "your order has been placed" page with a big green banner synchronously at the end of the cart flow. Some time later the user may then receive an email saying their payment method was declined. This certainly works, but a) it's horrible UX and b) it only works at the final stage of the process.
I see queues (and serverless) advocated as good architectural decisions, but every time they come up in a lecture/blog they're given in toy or data-sciency sort of examples. Is it possible to use these patterns in a sensible way where users are actually involved? (the blog mentions CQRS, but that seems... not a perfect solution)
We're building a user-facing app from scratch - the plan is to start with big coarse grained services, and just use REST for service to service communication when a user is waiting on the line.
But of course favor breaking up the services so there's no communication required is the first preference, and queue-based where it makes sense is the second preference.
Serverless is a meaningless term. The real name is platform-as-a-service, which is decades old. Running individual functions was just taking that to an extreme but any complex app will have multiple functions working together so it's right back to the same thing.
Many "serverless" environments are in fact converting to running an arbitrary docker container for as long as it needs to run, effectively becoming the next-generation of PaaS where you container can be whatever size you need, while still not worrying about infrastructure and individual servers.
I'm mystified by the word 'serverless'. Why are we using this word? There are clearly, _quite clearly_ servers involved. The wacky bugs involved in these 'server-less' services... are going to come down to what's happening on the servers involved.
EDIT: my question seems to have offended. Unfortunate, and unintentional!
It doesn't seem like a technical term at all.
Another commenter mentioned that it applies only to the price list, and that makes sense to me.
Serverless... pricing models. The pricing models are just silent.
I don't. Divorced from the context of a very specific vendor offering, serverless could mean almost anything. There are many applications, services, and platforms out there that don't require customers to manage their own servers.
Now if we're talking about Amazon or Google's specific things that have the marketing term 'serverless' applied... then many (but not everyone) can start getting specific (like in this posted article).
EDIT: I think another interesting question might be why so many people react defensively when this misnomer is criticized
There are many words in all languages that when viewed from certain perspectives don't seem accurate. People drive on parkways and park in driveways. Going around in circles like this amounts to trolling.
'trolling' makes it sound like critics are acting in bad faith. : (
In the context of AWS...
App/web servers
EC2 — I don’t have to manage the underlying hardware but I still have to either overprovision to deal with spikes, setup autoscaling based on metrics that I define, write scripts to make sure that new instances in an autoscalimg group is up to date, deal with OS patches.
Serverless - lambda - I give AWS a zip file and do a little configuration. It figures out how to autoscale.
Build servers
Same issues as above with EC2 and my build environment has to have everything every developer needs to build and one configuration can interfere with another developers build.
Serverless CodeBuild - I create a Docker image of my build environment or use one of the prebuilt ones and tell CodeBuild to spin it up on demand. If 10 developers need to build ten different repos at once, No problem. When no one is building, no charges.
Databases
EC2 - all of the problems about the EC2 and I have to manage OS patches, database patches, replication, and backups
RDS - removes most of those issues.
Serverless Aurora/DynamoDB. You don’t have to worry about “instance sizes”.
Docker.
ECS vs. Fargate - same as above With lambda.
I'm curious what size packages people are loading in lambdas currently. My largest lambda is a 4.7MB zip.
If anyone has some anecdotal info on runtime + package size I'm super curious to hear.
> So if you have to use an RDBMS, but put a service that handles connection pooling in the middle, maybe an auto scaling container of some description simply to handle that would be great.
Not sure I understand this. If I put another service in between me and the database don't I just end up having to open connections to that service?
I have an Aurora instance my lambda talks to, and it creates a new connection every time, and that's a bummer so I'm curious about how I could optimize that patch.
Really, it's all a sliding scale. Buying into any cloud, even just at the VM level, means you've accepted your automation tools are going to be platform specific, and require changing recipes/configs/etc to go elsewhere. If you want to leverage anything beyond that, such as object storage (S3 in AWS, which is super common), then your code has to become aware of IAM and S3 endpoints. Your code now has to change if you want to deploy it elsewhere.
I would contend your definition needs to be changed if the logical application of it basically makes everything more complex than 'someone else's server' to be "a hack"
[1] https://maas.io/
[1] https://www.ovh.com/world/public-cloud/storage/object-storag...
I like having out of the box authentication with Cognito, out of the box DBs with RDS and Dynamo, out of the box object storage with S3, Serverless as an option with the API Gateway and Lambda...and I'll worry about vendor lockin when Amazon ups the price of an AWS service.
But I understand your point about just connecting a bunch of APIs and being done with it, if you personally don't own the product, or if you are expecting to sell it soon.
A nitpick: I don't host my own object storage, I just pick S3 alternatives that provide standard APIs to access them.
But I’ve seen so many people get themselves into trouble by using dynamodb as a silver bullet and not thinking heavily about the design of their data. If you try and use RDBMS patterns (PK and FK relationships) there’s a whole heap of pain due to atomicity only being on one record in one table. Some times you can’t escape these patterns as well as it is the best fit for data or you require strong consistency of your model (e.g. financial).
They released dynamodb transactions (for Java only) but from what I’ve seen that just increases the complexity of solutions, you end up with more reads/writes per call and a heap of other stuff surrounding your data. Serverless story currently for stuff like this really has drawbacks