PostgreSQL Parallel Aggregate
blog.2ndquadrant.com
blog.2ndquadrant.com
Postgres is so amazing I still find it hard to believe it's free. I abused it many times (as a graph database; as a real-time financial data analytics engine with thousands of new ticks coming in each second; as a document storage with several TBs of data per node), and amazingly, it just worked. Magic.
If in doubt, choose Postgres.
While I was listening this presentation I've thought that it's quite weird functionality and intersects a lot with transaction visibility rules that MVCC already has. So it is interesting what is your use case?
Though I stored document relationships and social graphs, not semantic graphs, as they describe in this paper. I have little experience with semantic graphs — perhaps, they have different requirements.
Joins are good. :)
PostgreSQL is only free if your time has no value..
(Paraphrased. In jest..)
I worked with a certain other database from a company starting with "O" and ending with "racle", and support/documentation quality was abysmal. Our support contract was quite basic, but still, it expected to be better than free alternatives.
It was very NOT.
But in comparison to an enterprise solution like DB2 or Oracle? Not even close...
If you are trying to pull nosql into the argument, that is not even apples to oranges, plus with nosql, you get to retool your entire architecture.
But parallelism allows using more resources (particularly CPU cores) for a query, while the overall efficiency (the total number of instructions, CPU time etc.) remains about the same. It's just that it's split over multiple cores.
So if all CPU cores are already saturated on your machine (e.g. because you're already running multiple queries), it's not going to improve things.
I can say that I use Postgres and have not been disappointed with performance and it costs me $0. I encourage you to try it out. If you're a heavy user of SSMS, then you will have to change your flow a bit to accommodate psql. However, once you get a handle of psql, it's actually a pleasure to use.
- disks (SSDs) are very fast now, so cores are saturated more easily (when queries actually process the data instead of just reading it)
- multiple parallel (random) reads will likely be faster on HDD and SDD to some extent (esp. on larger RAID setups)
- the best optimization is still lots of RAM and people have that these days, so 100% CPU utilization during queries happens more often than not (the benchmarking setup seems suitable for more than 1 billion rows...).
Let's say you've got a web application with some cookie-cutter OLTP queries (which are well suited to serial plans ... no parallel aggregation benefit here). Your CPUs are running at 50% then some big OLAP query is going to come in and get maybe ... 13x faster.
Which is a nice performance improvement, but not 25x.
These numbers are a good validation of the work done in PostgreSQL, but you can't look at it and look at your X core machine, and say "my workload is going to be 25x faster!"
in my case, the OLTP queries run against a dedicated slave that's mostly busy doing nothing, waiting for the next OLTP query to happen. So for me this will be very, very nice :-)
If you know the number of cores you get that the efficiency is 80% (w.r.t. linear scaling) which is very good.
Edit the title of the HN post has changed since I made this comment... probably by being merged with an earlier submission of the same URL. Now I understand why HN readers are misinterpreting my comment. I did not intend to be negative.
How do you figure? As an 'EC2 person' I found that EC2 has made arbitrary specs utterly mundane and that is kind of the whole awesome point of EC2.
Need 200GB of RAM for the afternoon to process some data? No problem, just press a button and a minute later you have 200GB of RAM for the afternoon to process some data. Then just turn it off when you are done and forget about it.