Landlord: Per-Tenant Stats in Postgres with Citus
citusdata.com
citusdata.com
Once again I implore you to consider adding a "starter" or "hobbyist" tier for $20/month or even $5/month. If you want to really be competitive, your product should be free to use until you have more than 200mb of data stored in your database, at which point usage-based pricing kicks in based off of the number of rows used.
We understand some of the costs of switching and have actively been working to build tools for you to plan to use something like Citus, but then migrate when the time comes. Libraries like activerecord-multi-tenant or django-multi-tenant ensure you're ready for Citus and work perfectly well on a single node Postgres database.
Then when the time comes, which could be at $100 of RDS or $500 of RDS or larger, we have Citus warp [1] that allows you in a fully online way to replicate from RDS directly into a Citus cluster. Using it we've had customers with > 4 TB of data cutover with less than a minute of impact to their database.
I know that's not as simple as starting on Citus and continuing to scale, so hopefully we can address this better in the future, but for now it's the best answer we have at the moment.
[1]. https://www.citusdata.com/blog/2017/12/08/citus-warp-pain-fr...
For us the bigger issue is the steep jump in price for prod level plans and the lack of GCP hosting.
Also just a side note : The community edition is lacking some very important features like shard rebalance which is available in Enterprise edition to be fair.
It's not, you are just probably not old enough for the ones that are legal.
Source: Bought a house before, neighbors moved in with 4 kids who ran around screaming loudly and ignored that my yard wasn't theirs to play in. They also had 2 rather large labs that barked loudly day and night. Working from home with that sort of constant noise pollution is super irritating.
Also, CockroachDB is not yet suitable for:
Heavy analytics / OLAP"
I think Citus is actually very well suited for Heavy analytics / OLAP
And yes, cockroach isn’t Postgres but it has SQL, versus something like Mongo or Cassandra.
Citus is best used when transactions don't cross shard boundaries, in which case they execute on a single node and give you the low-latency to match.
The part about 2PC still holds.
Spanner uses Paxos (consensus) for replication within a key range (shard), but two-phase commit across shards: https://ai.google/research/pubs/pub39966
Citus relies on PostgreSQL's streaming replication, which gives higher throughput than Paxos, but Paxos has better availability characteristics. On the other hand, Paxos with leader leases as used by Spanner is similar to streaming replication both in terms of performance characteristics and short downtime during failover.
1. Analytical, but less data warehousing and more of a HTAP (hybrid transactional/analytical processing). In this case you're often ingesting a lot of data, often times sensor or log data from many endpoints, and then providing analytics across that data. The analytics needs to be up to date within minutes, and responsiveness of reports within seconds. You can see how Algolia (which powers the search for HN) uses Citus for this in their blog post - https://blog.algolia.com/building-real-time-analytics-apis/
2. Transactional. For a couple of years now Citus has had full transactional support when targeting a single node. Single node transactions can actually cover a breadth of use cases because it can span across tables as long as tables are co-located within the same node. We often see this is the case for multi-tenant/SaaS applications. In recent releases we also added support for distributed transactions. These transactions do have a higher overhead, but can often be hard to detangle from an existing application, thus us building support for distributed deadlock detection then adding distributed transactions.
Generally we're continuing to improve and support both of those use cases and have our usage base actually pretty evenly split between the two.
CRDB: Pros: natively distributed so HA and scalability are built-in, simple deployment and configuration, can run on Kubernetes for automated HA operations. Scaling tables is automatic and single-key and small-range OLTP performance is very good. Supports JSON and most data types for compatibility with most things that use Postgres.
Cons: Still maturing and has bugs like `select unnest(some_array_col)` not working. Obviously cant run any Postgres extensions so SQL w/JSON is all you get. Performance on large scans is slow, they're working on this but the distributed consensus required for queries means they will never match the latency of a single-node Postgres. Advanced queries are either very slow or unsupported or every slow.
CITUS: Pros: pure Postgres including extensions so you have access to advanced functionality. If you use shard key for queries, lack of distributed consensus gives low-latency performance just like single-node, but distributed transactions are still possible. Citus scales queries across all CPUs (on nodes holding the accessed data) so greatly improves query performance.
Cons: only distributes data in "distributed" tables (sharded) or "reference" tables (full replicas on all nodes). All other data just sits on single master node. HA uses Postgres streaming replication, requiring an inefficient 2x increase in costs, and is not seamless with failovers. Generally requires much more maintenance because it is still Postgres. Sharding does not accept multiple columns. No columnstores so large scans can still be slow, but they have ZFS in beta.
--
Summary:
CRDB for simpler OLTP with very low ops overhead and great availability, scaling, and durability.
Citus for advanced OLTP or OLAP, low-latency sharded access, and full access to all Postgres features.
TiDB is another competitor but mysql dialect and still early, missing lots of features.
For pure data-warehousing, we used MemSQL which is incredibly fast but can be expensive.
SQL Server is a great all around database if you want in-memory tables, columnstores, native graph queries, full-text search, and very high performance and can live with a single-node design (with optional HA cluster).
Being able to just drop the DB into a k8s cluster and not worry too much about failover gotchas and leader election has a lot of value. As does being able to throw more nodes on for more performance. Complicated OLAP queries aren't in scope.
Performance is fine, why do you say it's scary? OLAP will just be slow, but it's also distributed and unoptimized. Highly concurrent OLTP can't really get slow unless you're trying to stretch the cluster over multiple geo regions.
I guess the broader question would be: Why use GCP? I like them well enough, but I find it easier to prototype on AWS. Why would a developer develop the MVP on GCP vs AWS and then go to scale there? (I realized that K8S has made this, recently, a moot point, and it comes down to cost ... but beside that is there anything?)
EDIT: I realize that as a neo-Luddite, I just might be used to AWS, and would be open (mostly) to arguments to move to GCP if it is better (I'm not not talking about $/Gbps)
EDIT2: As a (mostly) fan of Google - I probably would have gotten myself acclimatized to their platform if it was any good 10 years ago, but it wasn't, so I didn't, and I guess I'm asking: is there a good reason to learn how to switch?
GCP has better performance (especially raw storage, compute and network), cheaper and simpler pricing, and more consistent experience. The raw primitives to build with are better designed and integrated, and scale without any tuning.
We're not building MVPs (which are just as easy) but have a working product and only use GKE and VMs. If you need all the managed services then AWS is the right choice, but K8S has made that a non-issue for us and allows us to use VMs (which are the most reliable part of any cloud), along with spot/preemptible pricing and better colocation for efficiency.
You say that you find that prototyping is as easy with GCP as with AWS? Do you believe it is because you are familiar with GCP? Or do you believe that they are truly at par with prototyping productivity?
(And yes - we've gone far afar of the upstream topic - but I do appreciate your feedback)
Look into AppEngine which is a few CLI commands for deploying. Same with Firebase which is widely considered a great mobile platform with everything you need. You can also use the Serverless.com framework to build apps on Cloud Functions, and they recently announced running any container in a serverless fashion (similar to Azure's ACI).
If you still want to run servers, then they have the best VM platform around.
As far as I can tell, it would be a great move for Citus because right now they are competing against Aurora PG within AWS. On GCP and Digital Ocean they have no equivalent competitor in the datacenter.
With the rest of a company infrastructure on GCP, you’re not likely to make a move just for Citus sake either.