The founders both have strong academic and commercial experience and the advisor is Prof. Joseph Hellerstein (he has both an excellent academic reputation and good experience in industry/tech transfer).
Note that TPCH is a "decision support benchmark". There are other technologies for helping postgres with these workloads as well: a column-store approach like https://github.com/citusdata/cstore_fdw, or the recent work on parallel sequential scans in Postgres core (http://rhaas.blogspot.com/2015/03/parallel-sequential-scan-f...), etc.
> "In an attempt to avoid losing my entire life to this website, I no longer comment as often as I used to, and, when I do, unless it is related to a topic (iOS jailbreaking) where it is part of my "job" to respond, I make it something of a policy to not look at things people say in response until at least a month later."
Given that their query engine uses an LLVM JIT compared to the standard PostgreSQL interpreter, these results are very reasonable. This is similar to comparing an interpreter vs a JIT for a programming language.
Even with the best intentions, benchmarking a database is notoriously difficult. I work on Presto, a distributed analytic SQL database, and we are often asked how it compares to other systems, or how it will perform for a given workload. The answer is always "try it and see", both because how it performs for your configuration and workload is the only thing that matters, and because there are so many variables that factor into the result.
Peter Boncz (creator of VectorWise, HyPer) says it best here: http://www.tpc.org/tpctc/tpctc2013/slides_and_papers/005.pdf
Remember also that the postgres planner and execution machinery is ANCIENT. Plan scoring still assumes Disk I/O = 10 MemoryAccess = 100CPU instructions. Modern CPUs are SIMD and Vector processing, but PG planning/exec machinery isn't vector aware. There's a 200X difference between L1 cache and main memory access that the PG planner can't see. And what about SSDs and NVMe vs disk? Spark gains its speed through caching intermediate result sets along with the ability to reconstruct them. But postgres doesn't cache intermediate results or query plans last I checked. Plus a million other things.
Postgres is wonderful, but not because it has natively fast analytics. 8X is nothing. 180X is better, but not as fast as Vertica, Aster, Greenplum, et al. I say kudos to Citus and Vitesse for leading the way.
I think we've already seen plenty of evidence that for OLAP type queries, there is a TON of room available for improvement vs. off the shelf PostgreSQL.