Making PostgreSQL Scale Hadoop-style: Benchmark Numbers
citusdata.com
citusdata.com
But lack of sharing is what we get when major open-source projects do not choose the GPL.
https://github.com/citusdata/cstore_fdw/blob/master/LICENSE
They're really nice talented people, and a great example of a company giving back to the community.
"How much you got?"
You pay a lot, you get a lotta scaling.
I hear you, and we are working on fixing that even as it might take some time. The challenge for us is that for an enterprise, alternatives could cost literally in the millions (see Oracle pricing at $100k's for just a single 8-core commodity machine). For start-ups, we have offered Citus for prices lower than $5k per node in the past, and we provide an entirely free community version as well.
Essentially, our take is to not have pricing be what stops you from using CitusDB. And if you are an enterprise, the value you get from using Citus should far exceed that you'd get from any other alternative out there.
It's just that "please call for pricing" means a negotiation with a sales guy, which many people find uncomfortable, unless they work in corporate purchasing.
On that other hand, I guess they could just still be working out their pricing ;)
https://www.monetdb.org/content/citusdb-postgresql-column-st...
This benchmark confuses CitusDB with PostgreSQL + cstore_fdw extension. CitusDB scales out PostgreSQL to multiple machines, and cstore is a columnar store for PostgreSQL. The author has a clarification posted at the end.
For single node Postgres + cstore numbers on TPC-H, we found that a few simple changes notably help. 1/ Analyze on foreign tables + increasing work_mem helps join queries by 2-4x, and 2/ Using the double precision instead of the numeric type increases aggregate function performance by 6x.
Lastly, we agree that vectorized execution can result in notable performance wins! See https://github.com/citusdata/postgres_vectorization_test for some initial work. We hope to incorporate some of MonetDB's vectorized execution features in cstore_fdw in the future.
The Citus vs Hadoop comparison feels a little apples vs oranges as presented.
I worked a bit with Netezza appliances which use an older version of postgres which can spread queries across a Bladecenter ... I wonder how this compares.
The downside of the Netezza (beside the huge cost) is that it is not expandable at all - to get more Netezza you need to buy another multirack system.
Also there is a bottleneck getting data in and out as there are individual host servers that you launch jobs through (ibm x3650s if I remember correctly).
Hadoop does a significantly better job than something like Netezza in those 2 areas.
I guess the head to head comparison would be Citus vs Impala/ Hbase? That is probably where a 'massively parallel' postgres setup that can scale horizontally would out perform its hadoop counterpart.
Free hardware?
I think showing Q2 and Q11 numbers would've been great, because for something like Tez, this is how those plans look in Hive (before the cost-based optimizer work)
http://people.apache.org/~gopalv/tpch-plans/q2_minimum_cost_...
http://people.apache.org/~gopalv/tpch-plans/q11_important_st...
Postgres's query planner should shine for those.
I tried google'n for tutorials but there are none.
There are no books on clustering or sharding postgresql too? At least I haven't found any.
To add sharding on top of that is a similar tutorial, but even more complicated.
I'm genuinely asking. I did not use it yet, but when I was researching database design subject that was my assumption we would use if we need to scale horizontally.