1.1 Billion Taxi Rides with MapD & 8 Nvidia Pascal Titan Xs
tech.marksblogg.com
tech.marksblogg.com
I was working on a similar concept (GPU-accelerated Hadoop MapReduce) at the time, and everyone I talked to in the Valley (and believe me, I knocked on a lot of doors) was perfectly happy with their Hadoop jobs running overnight or in an hour instead of instantaneously. It also amazed me how many people were using Hadoop only for the distributed storage.
Maybe data scientists love taking coffee breaks...
In the end, I scrapped the project because nobody wanted it -- everybody said they were only doing I/O bound work, which my solution didn't really accelerate.
Interestingly enough, I talked to people from Fusion.io (data center PCI-E SSDs) who mentioned that everybody told them that they didn't need Fusion because they were only doing compute bound work.
A YC company, BigCalc, was trying to do a similar thing with accelerating R (by rewriting library routines in highly-optimized pure C), targeting hedge funds, but suffered a similar fate as my map reduce project.
Shockingly, people just don't need speed.
Like your hedge fund example: why pay for accelerated R, when you can spend just a little more money on hiring a low-level C/assembly programmer instead and get 50x more speedup?
It's not just startups, big companies suffer the exact same problem. Look at Portland Group's OpenACC initiative, for instance, which intends to make GPU acceleration as easy to use as OpenMP, with pragmas etc. But everyone interested in GPUs just go and run straight CUDA or OpenCL code instead, there are essentially no "casual GPU acceleration users". Also see Cray Chapel.
Edit: spelling
I agree that while some users don't need this level of speed, many do. Hadoop use cases are often very different than those that demand the speed of in-memory processing for real time answers. Think about all the money going into high-performance solutions like Spark, HANA, Vertica, Redshift, Exadata, MemSQL, etc. And many use cases demand even more speed, which is where GPU acceleration comes in.
We just launched this year but already are being used in telco, hedge fund, social media, adtech, defense and retail deployments, with many more on the way. Most of our customers have been inbound, motivated by a common pain point of not having real-time visibility into their datasets. For example, Hadoop would not be acceptable for our use cases at Verizon, where troubleshooting needs to occur in real-time and not with lags measured in hours.
I'm a firm believer in horses for courses. Not everyone needs the fastest tech, but for those that do, it makes all the difference.
In my anecdotal sample, most of the users of hadoop/bigquery/etc seem to use it less because of raw query speed but more because the datasets are simply to large for any classical solution (i.e. much larger than fits into RAM on a single machine)
This is big enough to handle datasets of over a hundred billion rows, which is not big data by everyone's standard but big enough that its almost intractable to handle in real-time with other solutions, particularly if the use cases requires a lot of scans (witness Mark's other benchmarks).
Also, querying a static dataset that you can fit completely into RAM is not exactly "intractable to handle in real-time" without a GPU. You can do that just fine using the CPU right now.
Case in point: Even MySQL can handle 256gb of rows in memory with ease if you give it that much ram... Essentially, any single-machine database (gpu or not) can only process the same datasets that traditional single-machine databases (like MySQL) can already handle, albeit maybe a bit faster.
I think a good example for a "newsql" usecase that does _not_ fit into classical solutions like MySQL or Postgres would be web tracking (clickstream) data: Even a medium sized web property (say alexa top 100 in germany or france) will generate dozens of gigabytes of tracking data per hour. over the period of months, this adds up to hundreds of terrabytes of data. There is no way to load that much data in one piece into any of the classical databases.
EDIT: Maybe I should add a disclaimer: I'm the founder of another open-source database product that could be considered a competitor to MapD.
The scan performance of MySQL on that much data is going to be dramatically slower. MySQL may be good at indexed accesses but is simply not build for analytics workloads like those tested in the blog post. Not sure exactly how MySQL benches against Postgres, but the latter takes minutes over the same queries. http://tech.marksblogg.com/billion-nyc-taxi-rides-postgresql...
Unless I misread the post it's just comparing apples and oranges. If anything, the two benchmarks show that one piece of hardware (RAM) is orders of magnitude faster than the other one (SSDs). My previous comment was specifically about MySQL running on a machine where the whole dataset fits into memory.
Of course, I'm sure mapd is faster than mysql/postgres for some usecases. But the benchmark doesn't prove that in a fair comparison.
EDIT: Maybe I should add a disclaimer: I'm the founder of another open-source database product that could be considered a competitor to MapD.
Fair enough, but even if the data is in RAM a CPU solution will still be much slower. See this benchmark running the same queries on a 7-node Redshift cluster. http://tech.marksblogg.com/billion-nyc-taxi-rides-redshift-l... And MySQL for all its strengths is not in an analytics database and for these types of queries will be much slower than Redshift.
Note I said rack not single server.
Instead it reads like a marketing blog from Nvidia with some fairly meaningless benchmarks in fractions of a second.
http://toddwschneider.com/posts/analyzing-1-1-billion-nyc-ta...