I ask this question every year and postgresql have not deliver this. If there is any, there are hardly any documentation on it.
See also, ye manual: http://www.postgresql.org/docs/9.4/interactive/high-availabi...
Postgres-XL looks great for scale out, but you need 4 independent types of servers. Even with all those moving parts, it doesn't provide availability. If you want fail-over, you need pacemaker for the data nodes with traditional sync replication, and something like VRRP for your balancer, and something else to failover the coordinator. Several of these pieces can be tricky to set up in a cloud provider.
BDR looks nice, but it looks like there could be lots of gotchas for consistency in there. Maybe it is a magic bullet though... I don't know much about it yet.
Contrast with something like rethinkdb, mysql-galera, cassandra, etc, you start up enough nodes for quorum, tell them about each other, and you're pretty much done. The clients can handle the balancing, or you can use a pooler/balancer.
In my perfect world, I'd install postgresql-awesome-cluster-edition on 3 nodes, add the 3 IPs (or turn on multicast discovery, if my env can support it), and away we go for read scalability and availability. I do this today for mysql-galera, and other than the fact it's mysql, it's awesome. For writes, if you add 4 or more nodes, there should be some sort of shard system like XL has.
That said, postgresql is still clearly the best SQL and even noSQL single node server out there, it's a really great piece of software.
And there are plenty of options for dealing with the issues presented in those series.
I do think (for once) PostgreSQL is addressing it's core weakness and by version 10 will likely have horizontal scalability locked down. The new API is a really positive step.
This is a cost-effective way to be always maxed out (expanded as you say) with fixed price, this can scale well beyond the needs of almost every business. And , as I said, if you happen to be the next Facebook, Uber, Airbnb or whatever, you will acquire the know-how to scale.
If you don't need transactional semantics and you do need globally distributed multi-master key/value storage, I would not switch from Cassandra to PG, you already have a good solution for that.
If you don't need transactional semantics and you also do not need globally distributed multi-master key/value storage, use whatever you want, it doesn't really matter.
Cassandra is one of the better ones out there, but you have to deal with its data model and weird consistency promises (which however weird you think they are, are weirder)
The correct way to cluster also changes dramatically depending on your use case. Sure there are things like RAC that promise to make it just work, but those don't scale more than a few nodes.
Mongo is kind of the worst in this - it clusters in one weird way, has bad tooling, and subtly destroys your data at scale.
The general philosophy with postgres is to do it right, or not do it. There are ways to do specific kinds of clustering, but all of them (just like mongo, oracle, etc) have a lot of nuances to them.
If you have a natural shard key, use a bunch of schemas and table inheritance, and eat the downtime during re-shards. Check out citus as well. They have their issues, but they can help you hook up what you need.
I take it you haven't actually used Cassandra much in the last few years. It's data model is almost identical to a typical relational one and it's consistency promises are quite clear:
http://www.datastax.com/documentation/cql/3.0/cql/aboutCQL.h...
And I've scaled Cassandra clusters from 1 to 100 nodes in hours with no issues. It really is quite simple. Likewise have had no issues with MongoDB replica sets. It is definitely not "really, really hard".
> postgres is to do it right, or not do it
What a pathetic cop out. PostgreSQL has been around for decades they've had plenty of time to have a proven, stable solution implemented.
The lack of vector clocks in Cassandra can lead to some very non-intuitive (possible wrong) behavior - check out their counter implementation for some rage on that. It's pretty well made though, and I think C*, Hbase and Postgres all have great uses (along with Redis, and a lot of others)
Mongo tends to get things subtly wrong in ways that corrupt data, or that don't scale, and it gives up both A and C.
I'm genuinely interested. How does Mongo corrupts data? Thanks.
Mongos does weird magic as well when it gets confused, and will confirm writes to the wrong shards during "interesting" situations.
Is this behaviour what you are referring to?
> A rollback reverts write operations on a former primary when the member rejoins its replica set after a failover. A rollback is necessary only if the primary had accepted write operations that the secondaries had not successfully replicated before the primary stepped down. When the primary rejoins the set as a secondary, it reverts, or “rolls back,” its write operations to maintain database consistency with the other members.
They are working on improving FDW API to help make foreign tables available as inheritance children. I think It's a step forward...
[1] https://github.com/citusdata/pg_shard [2] http://www.citusdata.com/
Disclaimer: I work for Citus Data.