CRDB: Pros: natively distributed so HA and scalability are built-in, simple deployment and configuration, can run on Kubernetes for automated HA operations. Scaling tables is automatic and single-key and small-range OLTP performance is very good. Supports JSON and most data types for compatibility with most things that use Postgres.
Cons: Still maturing and has bugs like `select unnest(some_array_col)` not working. Obviously cant run any Postgres extensions so SQL w/JSON is all you get. Performance on large scans is slow, they're working on this but the distributed consensus required for queries means they will never match the latency of a single-node Postgres. Advanced queries are either very slow or unsupported or every slow.
CITUS: Pros: pure Postgres including extensions so you have access to advanced functionality. If you use shard key for queries, lack of distributed consensus gives low-latency performance just like single-node, but distributed transactions are still possible. Citus scales queries across all CPUs (on nodes holding the accessed data) so greatly improves query performance.
Cons: only distributes data in "distributed" tables (sharded) or "reference" tables (full replicas on all nodes). All other data just sits on single master node. HA uses Postgres streaming replication, requiring an inefficient 2x increase in costs, and is not seamless with failovers. Generally requires much more maintenance because it is still Postgres. Sharding does not accept multiple columns. No columnstores so large scans can still be slow, but they have ZFS in beta.
--
Summary:
CRDB for simpler OLTP with very low ops overhead and great availability, scaling, and durability.
Citus for advanced OLTP or OLAP, low-latency sharded access, and full access to all Postgres features.
TiDB is another competitor but mysql dialect and still early, missing lots of features.
For pure data-warehousing, we used MemSQL which is incredibly fast but can be expensive.
SQL Server is a great all around database if you want in-memory tables, columnstores, native graph queries, full-text search, and very high performance and can live with a single-node design (with optional HA cluster).
Being able to just drop the DB into a k8s cluster and not worry too much about failover gotchas and leader election has a lot of value. As does being able to throw more nodes on for more performance. Complicated OLAP queries aren't in scope.
Performance is fine, why do you say it's scary? OLAP will just be slow, but it's also distributed and unoptimized. Highly concurrent OLTP can't really get slow unless you're trying to stretch the cluster over multiple geo regions.
Also, CockroachDB is not yet suitable for:
Heavy analytics / OLAP"
I think Citus is actually very well suited for Heavy analytics / OLAP
And yes, cockroach isn’t Postgres but it has SQL, versus something like Mongo or Cassandra.
Citus is best used when transactions don't cross shard boundaries, in which case they execute on a single node and give you the low-latency to match.
The part about 2PC still holds.
Spanner uses Paxos (consensus) for replication within a key range (shard), but two-phase commit across shards: https://ai.google/research/pubs/pub39966
Citus relies on PostgreSQL's streaming replication, which gives higher throughput than Paxos, but Paxos has better availability characteristics. On the other hand, Paxos with leader leases as used by Spanner is similar to streaming replication both in terms of performance characteristics and short downtime during failover.
1. Analytical, but less data warehousing and more of a HTAP (hybrid transactional/analytical processing). In this case you're often ingesting a lot of data, often times sensor or log data from many endpoints, and then providing analytics across that data. The analytics needs to be up to date within minutes, and responsiveness of reports within seconds. You can see how Algolia (which powers the search for HN) uses Citus for this in their blog post - https://blog.algolia.com/building-real-time-analytics-apis/
2. Transactional. For a couple of years now Citus has had full transactional support when targeting a single node. Single node transactions can actually cover a breadth of use cases because it can span across tables as long as tables are co-located within the same node. We often see this is the case for multi-tenant/SaaS applications. In recent releases we also added support for distributed transactions. These transactions do have a higher overhead, but can often be hard to detangle from an existing application, thus us building support for distributed deadlock detection then adding distributed transactions.
Generally we're continuing to improve and support both of those use cases and have our usage base actually pretty evenly split between the two.