Amazon Aurora Postgres: First Thoughts
linkedin.com
linkedin.com
If I understand this correctly, they're using the same block store backing the primary database for the replicas. If that's the case, then wouldn't an issue with the blocking store hose both simultaneously? Or are they just jump starting the replicas copy of the data from the primary's block store until it's fully replicated?
> And from a management point of view, Aurora makes the database administrator's job far simpler since you no longer have to closely monitor your tablespaces and expand your block storage as needed (and reallocate tables and indexes across multiple tablespaces using pg_repack) in order to handle growing your dataset.
This is indeed very cool. Scaling from 1 MB, to 1 GB, to 1 TB, and beyond with no manual interaction is truly amazing. The storage pricing is a net win for Aurora as you only pay for what you use at the same price as gp2 EBS volumes ($.10/GB/month) whereas the latter has to be preallocated.
""" These are each replicated 6 ways into Protection Groups (PGs) so that each PG consists of six 10 GB segments, organized across three AZs, with two segments in each AZ. A storage volume is a concatenated set of PGs, physically implemented using a large fleet of storage nodes that are provisioned as virtual hosts with attached SSDs using Amazon Elastic Compute Cloud (EC2). """
The storage nodes are using local SSD instance storage, not EBS.
That seems like a problem that's easy to overcome. Allow people to configure the local storage size and you have a reasonably fair solution.
Still has its limitations for really large datasets, but that's also one you have with vanilla postgres.
That said, building a large, scalable Database still requires understanding the workloads and testing them out on production sized datasets. No easy button yet!
This article is definitely an interesting one as it highlights some of those trade-offs that they now need to account for.
Under classic RDS, when your application makes a SQL connection (the data plane) it's talking to a more or less stock Postgres instance, the same as you would have if you ran it locally.
Aurora, on the other hand, is involved in both the control plane and data plane. Your SQL connection is to a Postgres instance that's been forked/modified to work within Aurora.
Not sure the type of data he's storing, but 2 billion seems a bit much for 1 or a few tables. Hopefully he's thought of sharding those into multiple tables by date or key.