Processing 6 billion records in 4 seconds on a home PC
github.com
github.com
In other words, it's not actually really looking at a full 6 billion records for that query. More accurate would be the next query discussed, "Query 1", which takes 72 seconds to look over a much more significant portion of that 6 billion records.
It's still a pretty impressive set of numbers (as one would expect from GPU SIMD processing), but it irks me when short descriptions bend the facts to try to sound more significant. (Not to mention anything about the disk time subtraction.)
[0] http://www54.sap.com/content/dam/site/sapcom/global/usa/en_u...
Can anyone explain why that's a valid benchmark for him to use? Certainly the Hadoop version had significant disk access time?
If there is back and forth talk between the GPU and CPU, things will slow down considerably.
Relatively cheap compared to having to build clusters, perhaps, but $1000 isn't cheap for a desktop computing GPU.
A mid-high tier card (GeForce GTX 770) is closer to $400. A mid-range gaming card (GTX 760) is closer to $260.
Those finding the topic link interesting may also be interested in this CUDA radix sorting article[0] from 2010, as it featured "one billion 32-bit keys sorted per second."
[0] https://code.google.com/p/back40computing/wiki/RadixSorting
Of course for gaming the Titan is not considered cheap at all.
The problem with GTX770 / GTX760 is that their double-precision performance is gimped by NVidia. AMD cards are better at that price range for general purpose compute (which is why AMD cards are almost always used for BTC mining: http://www.extremetech.com/computing/153467-amd-destroys-nvi...)
Anyway, sorting and searching are all memory-constrained problems. GPUs have significantly faster RAM with significantly more bandwidth than CPUs. Hopefully DDR4 fixes the problem... but that isn't going to come for another year or two.
So, in other words, i subtracted time that both actually have to spend, for no good reason.
In order to further make results look better, I only subtracted it from my database, instead of running tests myself and subtracting it from both.
Look, disk time counts, whether it's hadoop loading it into the memory of a given machine, or you reading it from disk and transferring it into a GPU piece by piece.
Your "hopefully it will be irrelevant" is, well, crazy. I work for an employer with plenty of in-memory systems (very large ones in fact), and it certainly doesn't discount disk time. In fact, it matters a lot!
Of the 6 design considerations I listed, none of them are really addressed here. If you outgrow a single GPU then you have a huge performance penalty growing (that's a vertical growth). If you want to make your own operations (very common), then this would be impractical.
It's a nice idea but it'd be better to compare against things like Memsql and the like, where they have been designed from first principles for fast SQL processing. I'd recommend just dropping any Hadoop/HBase comparisons and compare within the same class, Hive is embarrassingly slow even in the class it's in (compare it to Google's Dremel/F1 or Apache Impala).
Your other considerations are still valid though. Although the point was to show the inefficiency of Hadoop/MapReduce when it comes to relational operations.
1. The TPC-H benchmark is measured in price-for-performance ($/QphH, or dollars per queries-per-hour). At 4 seconds for Q6, he's getting ~900 queries per hour. The cost of his rig is probably ~$2k, so he's under $2 per QphH. The top TPC-H scores are around $.10, but <$10 is pretty good for a first go.
2. The standard knock against GPU processing is the time it takes to load GPU memory. GPU processing may be blazing once data is in memory. But there was an MIT paper last year claiming you couldn't load the GPU fast enough to keep up. Evidently, he's keeping up.
With regard to comparing his performance to hadoop/hive - yeah it's apples and oranges, but he's in good company. Hadapt, Hortonworks Stinger, Cloudera Impala, Spark/Shark and others all rate themselves on how many times faster they are than Hive.
And frankly, I don't buy the whole "the point of MR is for huge, horizontally scaling networks" If you factor out Yahoo!, Facebook, Amazon, LinkedIn and a few others, the largest remaining hadoop clusters are all WELL south of 1000 nodes. And most run on homogenous high-end hardware.