Amazon Aurora Limitless Database
aws.amazon.com
aws.amazon.com
For the other 10%, we take them aside and politely explain that they almost certainly have an unusable staging environment per the scope of our B2B project.
Testing in production is a wonderful path if you are comfortable talking to business people and making lots of compromises.
Make sure to prune old data from your tables. This one got to this limit because it eventually got too large that queries to delete old data would time out... so it just kept growing.
> Sharded tables – These tables are distributed across multiple shards. Data is split among the shards based on the values of designated columns in the table, called shard keys.
It sounds like this is very much managed CitusDB on top of Aurora, but without any details about the implementation its impossible to know if its just repackaged Citus, or some new novel technology built specifically for Aurora.
[1] https://aws.amazon.com/about-aws/whats-new/2023/11/amazon-ti... [2] https://assets.amazon.science/dc/2b/4ef2b89649f9a393d37d3e04...
So are you using vector clocks in the backend to handle transaction ordering and maintain consensus?
Super exciting announcement, and I am really looking forwards to learning more!
edit: More details make it seem like this _isn't_ Postgres on the frontend, its just mostly postgres compatible, which would explain how they got away from the xid based MVCC. So it seems like this is an entirely new distributed database, rather then a modification to existing postgres.
given the mentions of "log as a database", I believe that depending on the transaction the storage responds differently. like how mysql mvcc uses undolog to essentially rollback database internally so that transaction sees data consistency. they could be doing something similar i.e regardless of whatever postgres uses, if they can get a reference of transaction and its start time then they can use the custom storage engine to rollback the log and respond in that way
From what I hear from folks using things like Vitess, if you're used to a monolithic SQL database there's often some things to learn to mentally model your query cost well after you move to a sharded world, and understanding more up front can save heartburn later. Writing up those details well is a good thing that AWS could do.
1. Scales to zero, no cost when not using
2. Allows for SQL over API like aurora v1. This point is important, it allows for faster access from other serverless technologies, i.e. lambda
only case I could think of is that businesses that want "scale to zero" have very low total expenditure on the database. is this the case for you as well?
because serverless without going to zero still solves a major problem for a lot of companies with some decent scale, where there is some decent traffic and its too hard to implement database autoscaling and most businesses have unequal traffic (daytime and nighttime etc,...)
The point is: You don't want to need to switch technology at a later stage, you'd like to build with the right technology from the get-go.
I almost always go for DynamoDB first, since for small projects it's essentially free, and for huge projects I do not have to worry about the ops overhead that follows normal SQL databases.
A lot of people also have traffic patterns that go:
- Very high volume during the day - Almost no or little volume during the night when your users are sleeping
Aurora Serverless v2 gets us closer to this for SQL, but does not scale to 0 so it's not too nice for the initial phases of projects.
would be the non-editorialized submission, and would have advised potential readers that one cannot currently play with it without a ton of hoopjumpery
one will also want to watch out for this, buried 7 paragraphs in:
> The preview runs in a new Aurora PostgreSQL cluster with version 15 in the AWS US East (Ohio), US East (N. Virginia), US West (Oregon), Asia Pacific (Tokyo), and Europe (Ireland) Regions.
https://www.lastweekinaws.com/blog/us-west-1-the-flagship-aw...
As if data egress charges aren't gouging customers everywhere!
We are using the ~equivalent SQL Server architecture in Azure right now (Hyperscale tier) and really enjoy the experience. Not having to worry about specific machines and being able to treat it the same way as an express instance (from an app perspective) are extremely compelling selling points.
The only downside I saw in the Microsoft offering was the fact that there was a throttle for transaction log writes (100 megabytes/s) which is required to maintain resilience across the cluster. You don't want your replicas to wind up 3 weeks behind, etc. My hope is this limit will only increase with time, but even so it is quite workable (for us) today.
I may spin up an instance of Aurora (when available) to take a look at how it behaves around our worst-case serialized txn paths. Our use of SQL Server is very "light" in terms of its weird claws/hooks, so we could pretty easily shift to another RDBMS vendor in the future, assuming we can sell our clients on the 3rd party mix.
Maybe they improved since. Let me know how you find it:)
But they must be confident that that whatever the real limit is, is just not realistic for any use case.
Maybe it's because most projects will never need it, but when you eventually do it feels like you're just left to yourself
Are we giving away something to get multiple writers to same table? Or does the routing to the correct shard somehow solve for that?
Limitless horizontal scaling is really cool and all, but does anyone not running a Fortune-500 Tech company actually need this kind of firepower?
This being part of native aws fully postgres compatible is exciting though.
I wish AWS brings columnstore or custom storage engine features.
But I can see why AWS wouldn’t want to bring columnstore tables - it would directly compete with redshift.