AWS cancels serverless Postgres service that scales to zero
datanami.com
datanami.com
It's possible and here is the recipe: https://twitter.com/nikitabase/status/1725394843285766463
In the future we are considering separating the buffer pool from compute and turning off compute doesn't drop the buffer cache. It's a research project but PolarDB did that for AliCloud.
Though computers are pretty fast these days (and RDMA exists), so if the pieces are in the same data centre and using good interconnects it's probably not going to make much of a practical difference.
One area where Aurora is superior to non-Aurora in AWS is replication lag. When running non-Aurora MySQL servers, we had replica lag sometimes reach an hour making read replicas useless. With Aurora provisioned, replica lag is always under 30ms.
If your app is database heavy and you end-up using database pretty much 24/7, don't fall for serverless hype and go with Aurora Provisioned.
> Some unsuitable workloads for DynamoDB include:
> Services that require ad hoc query access. Though it’s possible to use external relational frameworks to implement entity relationships across DynamoDB tables, these are generally cumbersome.
> Online analytical processing (OLAP)/data warehouse implementations. These types of applications generally require distribution and the joining of fact and dimension tables that inherently provide a normalized (relational) view of your data.
> Binary large object (BLOB) storage. DynamoDB can store binary items up to 400 KB, but DynamoDB is not generally suited to storing documents or images. A better architectural pattern for this implementation is to store pointers to Amazon S3 objects in a DynamoDB table.
For each of these workloads you switch the query pattern to accommodate for DynamoDB and provide a solution for your workload and/or app. That is actually the secret of NoSQL. You do the work upfront.
So your reading of the recommendations is correct on the surface but incorrect on the fundamental usage aspect of it. Lets have a look at each:
> Services that require ad hoc query access. Though it’s possible to use external relational frameworks to implement entity relationships across DynamoDB tables, these are generally cumbersome.
A: Dont do this. Don't do ad hoc query access. Define a series of access patterns and lay out your data to support them. You dont get ad hoc query access for Netflix, Amazon or your Airline Travel website...
> Online analytical processing (OLAP)/data warehouse implementations. These types of applications generally require distribution and the joining of fact and dimension tables that inherently provide a normalized (relational) view of your data.
These are not most SQL Patterns. This is OLAP and was always or should be done with MPP systems, not with your relational database like your SQLServer or your Oracle. So this is going on an edge...
> Binary large object (BLOB) storage. DynamoDB can store binary items up to 400 KB, but DynamoDB is not generally suited to storing documents or images. A better architectural pattern for this implementation is to store pointers to Amazon S3 objects in a DynamoDB table.
Your Relational Database will not store BLOBs more efficiently than S3 anyway....
"When is "ACID" ACID? Rarely." - http://www.bailis.org/blog/when-is-acid-acid-rarely/
So your post is both irrelevant as a response to mine about consistency and more generally irrelevant to the entire discussion here.
From the article: "The textbook definition of ACID Isolation is serializability (e.g., Architecture of a Database System, Section 6.2), which states that the outcome of executing a set of transactions should be equivalent to some serial execution of those transactions. This means that each transaction gets to operate on the database as if it were running by itself, which ensures database correctness, or consistency."
" This means that each transaction gets to operate on the database as if it were running by itself, which ensures database correctness, or consistency. A database with serializability (“I” in ACID), provides arbitrary read/write transactions and guarantees consistency (“C” in ACID), or correctness, of the database. Without serializability, ACID, particularly consistency, is generally not guaranteed"
I am of course ignoring the Consistency you are certainly aware that exist in DynamoDB with Strong Consistency and DynamoDB Transactions.
Moving consistency management outside of DynamoDB as you propose doesn't circumvent the CAP theorem's limitations.
How would you do it? Assume we want proper pagination, and not rewrite the app for cursor based "Load more" style pagination. Why? Because the React Admin provider API insists. https://github.com/marmelab/react-admin/issues/1510
You are not supposed to do the same query patterns. My argument is that you can substitute your relational database, by changing the app and the layout of the data to match the proper patterns for DynamoDB.
"Migrating to DynamoDB from a relational database" - https://docs.aws.amazon.com/amazondynamodb/latest/developerg...
""How to model one-to-many relationships in DynamoDB" - https://www.alexdebrie.com/posts/dynamodb-one-to-many/
You can go further by using the streams feature to dump your data into an analytical database for your querying needs.
It gets worse than that though, as Dynamo's usage on large tables depends on its hottest shard, and other than setting primary keys, you have no control over how it shards. The story of people moving to dynamo and touting its advantages, just to move out a year later because the shape of their data makes dynamo prohibitively expensive as the data grows is pretty common. I once worked at a place where they stored historical data in dynamo, leading to a substandard hash key. Most of the time the db was idle, but when it wasn't, the sharding scheme made them pay a good 500x capacity than it was actually in use, because while a shard was red hot, others were completely idle. The monthly price for this relatively unimportant feature ended up being higher than the office's rent, in San Francisco.
If your table's key is UUIDs, and the chances of querying one record or another is almost perfectly flat, then sure, dynamo away! But if you walk away from it very best use case, and you suddenly start having any amount of data... dynamo can become really, really expensive. I'd not say you should never use it, but it's a really scary first place to go, precisely if you care about costs.
Are there any truly serverless SQL databases out there?
sqlite ;)
We also tried serverless Redis. We have a tiny machine at %10 CPU utilization, but sustained queries. It costs 20x more. We were very afraid of the result we just ran it for an hour to get an idea.
Serverless, pay per use sounds only good if you have very few or unusual traffic.
We just run a monolithic but stateless Spring Boot Application and a dedicated RDS instance with reader in stand by. Only serverless dependency is SQS because it works nice enough. This simple infra has allowed us to iterate and pivot our startup very fast as needed. We still need more customers though.
And most of those cases can be handled by a $5 a month instance. Thus, usually, the most serverless can save you is $5 a month.
I suppose the real appeal of serveless is you don't have to worry about maintaining the server, doing updates, etc.
Scaling to zero is tricky. In order to make it work you need to
1. Separate storage and compute which aurora has done
2. Have a proxy between the user and an actual VM in which you host Postgres
3. Spin up this VM after authenticating a connection, stand up Postgres inside the VM and run the query
You need to do all that very quickly- ideally in a couple hundred ms. This can only be accomplished with having a warm pool of VM or having extremely fast microVMs like firecracker.
I believe AWS aurora just uses ec2 and that’s why scale to 0 is turned off
I had settled on "scale to zero" as my chosen definition for it, because it was the only definition that truly made sense to me and that I very strongly valued.
Apparently AWS don't think it means that.
I guess it just means "you'll never have to SSH in and upgrade anything"?
The other thing that surprises me here is that one of AWS's biggest selling points is how rarely they break existing apps by turning off services developers depend on.
In theory it provides better incentives for both parties:
Customers don't have to pick instance size, etc. to try and accomodate their workload. They can buy throughput, etc. in fractions of an actual instance. They can look at their quota usage to determine whether to "scale up" or "scale down".
Service providers are incentivized to optimize their software. If you sell instances + a service on top, there's a perverse incentive not to make your code too fast or people will buy fewer instances. In serverless, providers are incentivized to reduce their costs and optimize the entire stack.
Scale to zero isn't necessarily a property of serverless - a service provider can agree to keep your data on ice for free, but that's a business decision.
You manage purely logical containers like databases and tables, the serverless service takes care of the logical to physical mapping. This lets them get better density, manage patching scaling, and a ton of other things
To me that is just "managed". I agree, serverless means scale to 0 (compute and costs). I'm so glad I left AWS Aurora Serverless after how much I disliked v1 and never considered v2. If I was still stuck on them my costs would double or more for my dev/qa environments (which I had scale to 0).
We have "managed" for that.
But if you manage to get a meaning for "serverless", I would love to know it too.
Somebody answered the GP with "it's SaaS, not IaaS". The more I think about it, the more I think this actually represent what people mean when they say it.
It seems in many cases RDS Postgres is cheaper and even when using Aurora, using fixed-size instance classes instead of "serverless" autoscaling classes (with ACUs) is a lot cheaper. Fixed Aurora on-demand instances are roughly 17% the cost of ACUs. Fixed instances are usually >= 30% cheaper on top of that if you reserve.
Aurora heavily favors performance and availability over cost.
Like clock work, at 3pm PS most weekdays there was a spike in traffic as our customers became active on our system.
I would love to see the real numbers behind the costs. Is time of day dependency the real price driver here? Are some customers getting more value than others out of this system?
what knowledge gap do i have about the collective womanhood that I don't understand why a service which is predominantly used by women gets traffic spikes at a specific time of day?
Is 3PM PST when most women sync their in-built RTC?
That starts 6pm ET, so post work and kids, it's "free time". By 6pm PT(9pm ET) we would be hitting a peak, for the next hour or two and then winding down till 11. It's the prime time for "adult working women" as a demographic.
On the contrary, vCPU a completely cloud-centric product oriented concept, intended to abstract away what CPU you're actually running on, and how fast. These abstracts make orchestration and availability easier, but I think the key is a stable, public internet access. If you get a static IP, I don't see what wonders this does for "microscale". In fact, you probably have 90% unused "microcapacity" on all the personal hardware that's often idling, that is a lot less micro than any cloud provider will give you for free.
They claim 99.99% availability or whatever the figure is, that it’s self healing, allows up to 2 replicas to be down for write availability and 3 replicas to be down for read availability. They promise “millisecond consistency”, whatever that even means.
Some of the blogs mention that they use distributed state to extract slivers of consistency across the nodes, which is just scary stuff to hear.
I’m seeing more and more usage of Aurora due to the magical component it offers, but I can’t wrap my head around what it’s actually supposed to be, and what the failure model is, like I can with RDS PG and MySQL.
CP means you can’t be available, meaning your system will stop replying to queries.
AP means you’re always available, but you could receive stale data when querying.
It’s irrelevant whether it’s a single binary running on EC2 or a billion services working in unison. It’s a property of the system.
https://cloud.google.com/blog/products/databases/inside-clou...
They do actually provide quite a lot of detail about Aurora storage this works. This 2019 reinvent talk gets pretty deep into the weeds.
https://www.youtube.com/watch?v=uaQEGLKtw54
https://d1.awsstatic.com/events/reinvent/2019/REPEAT_Amazon_...
This year I started to run into some issue with PS mainly around their plans changing (went from pay for reads/writes/storage to pay for compute/storage). Yes, yes, I know they still offer the $30/mo plan but it's billed as "Read/write-based billing for lower-traffic applications" and they dropped all mentions of auto-scaling. That coupled with them sleeping your non-prod DB branches (no auto-wakeup, you had to use the API or console) even after saying that was a feature of the original $30 plan rubbed me the wrong way. Eventually the costs (for what I was getting) were way too out of whack. My app is single-tenant (love it or hate it, it's what it is) so for each customer I was paying $30/mo even though this is event-based software (like in-person, physical events that happen once a year) so for most the year the DB sat there and did nothing.
Given all that I looked into Neon [1] (which I had heard of here on HN, but PS support suggested them, kudos to them for recommending a competitor, I always liked their support/staff) and while going from MySQL to Postgres wasn't painless it was way easier than I had anticipated. It was one of the few times Prisma "just worked", I don't think I'd use it again though, that DB engine is so heavy especially in a lambda. I just switched over fully last week to Neon and things seem to have gone smoothly. I can now run multiple databases on the same shared compute and it scales to 0. In fact it's scale up time is absurdly fast, the DB will "wake up" on it's own when you connect to it and unlike AWS Aurora Serverless v1 it comes up in seconds instead of 30-60+ so you don't even have to account for it. With AWS I had to have something poll the backend waiting to see if the DB was awake yet, to fire off my requests, if it was asleep. With Neon I don't even consider it, the first requests just take an extra second or two if that.
I don't have any ill will towards PlanetScale and I quite enjoyed their product for almost the whole time I used it. Also their support is very responsive and I loved the branching/merging features (I'll miss those but zero-downtime migrations aren't required for my use-case, just nice to have). In fact if I had written my app to be multi-tenant then I'd probably still be on them since I could just scale up to one of their higher plans. It does seem like Neon is significantly (for me/my workload) cheaper for more compute, I had queries taking _forever_ on PS that come back in a second or less on Neon all while paying less.
All that said, I _highly_ recommend checking out Neon if you need "serverless" hosting for Postgres that scales to 0.