BlazingDB uses GPUs to manipulate huge databases in no time
techcrunch.com
techcrunch.com
The weakness of GPU databases is that while they have fantastic internal bandwidth, their network to the rest of the hardware in a server system is over PCIe, which generally isn't going to be as good as what a CPU has and databases tend to be bandwidth bound. This is a real bottleneck and trying to work around it makes the entire software stack clunky.
I once asked the designer of a good GPU database what the "no bullshit" performance numbers were relative to CPU. He told me GPU was about 4x the throughput of CPU, which is very good, but after you added in the extra hardware costs, power costs, engineering complexity etc, the overall economics were about the same to a first approximation. And so he advised me to not waste my time considering GPU architectures for databases outside of some exotic, narrow use cases where the economics still made sense. Which was completely sensible.
For data intensive processing, like databases, you don't want your engine to live on a coprocessor. No matter how attractive the coprocessor, the connectivity to the rest of the system extracts a big enough price that it is rarely worth it.
select id, name, age, avg(income) from people group by gender
In this case only the income and gender columns would actually be sent to the gpu and they would do so in a compressed fashion to increase the "effective" bandwidth of data over PCIE. Even more interesting is that id, name, and age, would be pulled from our horizontal store instead of our compressed columnar store in order to minimize the number of iops necessary to fill the result set.
Outside of CPU >> GPU, I'm not sure what other data movement you could be talking about. A SAS HBA or Ethernet NIC or Infiniband HBA are almost always going to be operating over the same PCIe bus the GPU uses. In the rare instances they're built onto the CPU, the "network link" is likely going to still be significantly slower than the fastest PCIe slot.
GPU's may not be great for every unit of work a RDBMS has to perform, but given their ability to rapidly compute hashes it could help a lot with joins (as evidenced by PGStrom).
And a big on the PowerPC roadmap http://www.nextplatform.com/2016/04/07/ibm-unfolds-power-chi... "With NVLink, multiple GPUs can be linked by 20 GB/sec links (bi-directional at that speed) to each other or to the Power8 processor so they can share data more rapidly than is possible over PCI-Express 3.0 peripheral links. (Those PCI-Express links top out at 16 GB/sec and, unlike NVLink, they cannot be aggregated to boost the bandwidth between two devices.)"
NVLink has 80GB/s [1]. DDR4 quad channel (Xeon Servers) has ~120GB/s [2]. So no this rationalization doesn't fall apart. Furthermore in the event NVLink gets faster then RAM, then you'll still be bottlenecked by RAM access, as you'll buffer here.
This of course is ignoring weird systems where you attempt to maintain ACID coherence of tables between GPU, CPU, and Disk memory. But then GPU memory size become inherently limiting as even the biggest max out at ~32GB.
[1] https://en.wikipedia.org/wiki/NVLink
[2] http://www.corsair.com/en-us/blog/2014/september/ddr3_vs_ddr... (2channel -> 4channel x2)
Furthermore, all that memory bandwidth is calculated against all the cores. So you have to be VERY careful in usage patterns (it doesn't work like a giant CPU). Not to mention how much it costs involved per GB or TB!
I have some experience in this area, and where GPU's & database really shine is building (and especially) re-building indexes.
Running a database on GPUs isn't going to replace all DBs overnight. But, the hardware does have it's uses; just like improved SIMD on CPU's and NVMe storage, plus developments in networking - Basically look where Intel is going since we're coming to the limits of silicon transistors..
Though, I would like to see Xeon Knights Landing compared to GPUs for db uses.
Well yes and no. Both IBM and Oracle have gotten impressive performance using database-specific co-processors, but these are not GPGPUs, they are dedicated hardware that sits on the storage path. Baidu are reinventing that wheel with FPGAs too.
How relevant is that when you're looking at multi-TB data sets that don't fit into computer RAM? Sure, the RAM <---> CPU bandwidth may be very wide, but the SSD connects to the computer over the same PCIe bus.
And also: when did you have this conversation? GPU performance has changed very much year by year, so what wouldn't have been worth it 2 years ago might be a huge gain now.
It's not about GPU performance, it's about the latency and bandwidth of getting that dat to the GPU. If once you ship data to the GPU, you reuse it many times for many calculations, that cost is amortized and it doesn't matter as much. But if you ship data to the GPU and use it once, then that cost will probably not be amortized. I think of databases tending to fit in the latter category.
[0] - http://www.nersc.gov/users/application-performance/measuring...
Or the ability to not be ridiculously wasteful of existing resources (also exists).
Your naysaying doesn't make you smart. Your naysaying makes you cut off from learning a different, better way of doing things.
Big joins are our best use case. Joins are hard for many databases to optimize when they have not seen them before or are not "expecting" them like when you let amazon know how to partition your different tables onto the same physical machines so that redshift can return your query in a reasonble amount of time. But most SQL operations can be accelerated by the use of GPU's. Order by (holy smokes it helps), arithmetic or date transformations (20-30x for comparable cpu code), predicates, group by. All of these operations are happening over vectors of data. SIMD rock out when it comes to running these kinds of loads. The only use cases that we actually think are very poorly suited to gpus thus far (and this is a nut someone will one day probably crack) is wild card string searches. Some of our competitors handle this by caching all the data in GPU RAM but we consider that to be "cheating" since you would never be able to justify the pcie transfer to do wild card string searches.
Why not put GPUs on your analytics machines? Or a cluster of them with SPARK. Or heck, distribute the spark cluster on top of your database.
Save the GPUs for training neural nets and physics simulations. When the fundamental hardware capabilities of GPUs make them worth investing in for database workloads, I'll change my opinion. Until then, any budget for expensive DB hardware is much better spent on NVMe storage.
There have been some very sweet algorithms for doing B+ tree searches for Itanium and for AVX and those work just as well, if not better on GPU.
With some thinking, branch operations can be converted into math and applied to masses of data without checking for branch conditions. This wastes some work but is still faster than branching.
In particular I mean sorting non-English text which any serious database needs to be able to do. Subtracting characters doesn't work when you're dealing language specific Unicode collation [0].
What you can do is decide how your comparison should be collated and preprocess your strings into a sort key form that is strictly big-endian binary (first byte has highest weight). Or little-endian I suppose, whichever works best for your hardware.
Your sort key doesn't have to be the actual string.
select * from table where column1 = "some awesome text here"
In this case we actually do a comparison between hashes of the data. Hashing is cheap, fast, and makes comparisons on the GPU's be a breeze. So long story short, we have no data dependent branching. We do this by never using certain statements inside of kernel code.
the use of "if" is expressly forbidden at blazing for any gpu code and it's use is punished viciously (said individual usually has to be the one that captures meaningful input from one the 80 log files of our 80 gpu cluster ).
Editing to mention the way you can encode a long string down to a 1 byte is by doing a dictionary compression and then bit packing the keys. On gpu the way you do this is getting the max key (min key is always 0) and then you can store this data in 1 byte (if max(key) < 255).
Decompression for processing: We can roundtrip decompress 8 byte integers RRleDeltaRle 4x on an AWS g2.2xlarge faster when we use the GPU. This includes sending the data TO the GPU and brining it back. Our decompression segments on CPU were set up so that every thread was processing a segment to be decompressed so we were using every avaiable thread at about 100%.
Sorting data: Here the difference can be startling. On an aws g2.2xlarge we are able to sort orders of magnitude faster than you can on GPU. Checkout thrust to run some exmaples http://docs.nvidia.com/cuda/thrust/#axzz4K7CRY352
A few modifications there can let you run this with both and NVIDIA backend and one that runs on CPU threads. It will run orders of magnitude faster on GPU than CPU on a small gpu instance on amazon. Even a laptop gpu would still outperform the cpu sorting capacity by at least an order of magnitude.
Why not?
It's certainly easier to run parallel Atari sims on CPU, because CPU programs are typically written as single-threaded or with parameterized number of threads.
Running parallel Atari on GPU is completely possible, with either running an Atari game on the each of ~30 SMs or each of the 32 * n_SMs ~= 1000 warps. However, because GPU code is typically written and delivered as kernels which utilize the full GPU, this type of embarrassing parallelism over SMs or warps typically can't be gained from using an existing library.
How many fewer cores, and how much less GPU RAM, and how much slower were GPUs in general when this was tried? They're changing quite rapidly—significantly faster than any other component in a modern PC. Any attempts more than 2 or 3 years ago aren't very relevant.
For graphics this usually doesn't matter as a minor color or vertex deviation isn't noticeable, but for compute it can be devastating.
We do cryptocurrency mining on an industrial scale and constantly see single bit errors from hardware that is brand-new without modifications.
That's very surprising and interesting. How do you detect these single bit errors?
Dont the GPU specs allow for a certain lossiness in the math? Or like, at least they don't conform to IEEE 794 float specs w/ regard to order of operations, precision, degradation, etc.
So like, do a shitload of math ops in a glsl shader with a deterministic outcome, render the result to a texture, take the texture back and make sure the RGBA values match bit for bit with the numbers you expected?
Or to detect single bit errors in the gpu's local memory or caching just attach textures, read data into them, read it back, render back, etc etc.
Detection of false negatives: Compare solution distribution and frequency to expected models; switch to debug kernels if outside tolerance.
However, this works because the mining problem space is stateless and follows strict mathematically predictable models.
A DB is stateful and the answers generally can't be verified without consulting a secondary copy, which is why I'm super curious how they would engineer correctness and reliability in a cost-effective way using GPUs.
I'm assuming you are using GeForce cards and not Tesla cards which have an ECC memory protection mode?
I've tried to collect some statistics on GPU memory errors rates but have found them to be normally extremely rare. The only time I've reproducibly seen them is due to faulty hardware, where the errors become highly reproducible and the GPU needs replacement. The other theoretical cause of bit flips is supposed to be random errors due to cosmic radiation but I've never been able to observe that using memory testing software (though I did only run the experiments in AWS).
Could it be that you have faulty or low grade GPUs? I assume these are all low-cost OEM parts, given your application? Or maybe there's something odd about your data center environment?
Regarding the GPU database application, I think the answer is to just use the Tesla grade GPU with ECC memory enabled.
AMD's hardware specifications are more open too which lets one build your own shader compilers and get direct access to the iron.
SHA-256d (Bitcoin) and Scrypt (Litecoin) have ASICs, X11 (Dash - formerly Darkcoin) has FPGAs.
I cannot find any reference to handling of soft errors in their material. One rather banal approach is to do everything twice and check the results; effective, but I'm not sure that they do this. They may be simply putting their heads in the sand.
The only thing that matters for them here is the aggregate, cross-sectional bandwidth to your data’s working set in memory. For databases, especially for the approach that many GPU databases take (light on indexes since GPUs aren’t great at data-dependent memory movement or pointer chasing, just brute-force scan much of the data), the working set size is something that will only fit in main memory.
Instead of using 8 GPUs with a peak global, cross-sectional memory b/w of 8 * 320 = 2560 GB/sec and brute-force scans, you can parallelize across ~40 CPU nodes each with ~60 GB/sec b/w to main memory. The cross-sectional bandwidth will be about the same, and the cost to split and join the results of the query is likely small in comparison to actually doing the work, assuming the intermediate results are reasonably small. You can use a broadcast and reduction tree; the added latency of the broadcast and reduction tree's depth likely won't add much, since there isn’t much data to broadcast in a query, and the data returned by each machine for the reduction is hopefully (!) tiny in proportion to the actual data scanned.
If you want to consider indices on data, then maybe the heads of that can remain resident in a CPU’s cache, and will make the individual CPU scans even faster. The GPU caches are tiny and mainly serve to patch up strided loads and other bad uses of memory.
Whether or not it’s worth it one way or another depends upon how large your database size is, the relative cost of GPUs versus CPU nodes to get the memory you want and the cross-sectional b/w you need, perf/W and other issues.
You’re probably nowhere near close to arithmetic throughput bounds on GPUs or CPUs since these workloads have very low op / byte loaded ratios compared to typical HPC workloads, so that aspect of GPUs doesn’t matter. If you’re doing expensive pre- or post-processing on GPUs as well, then that may push the balance more towards GPUs.
which are why we actually prefer to have only 1 GPU per server when we are making our own boxes. We find that our most optimal running environment is when we have smaller instances with only 1 gpu. This is due to the fact that two smaller rigs with a gpu each will benefit from increased CPU RAM throughput (basically double 1 rig) and you have 2x the PCIE bandwidth since you are splitting it across two machines. While we don't have the enviable op / byte loaded that some machine learning tool sets might use we are able to greatly enhance these throughputs by using compression and doing things like running multiple arithmetic operations in one kernel call.
I vividly remember the Intel 8087, a math co-processor to the Intel 8086 that came out in 1980-1981. All the floating point arithmetic was offloaded to it.
It ended up disappearing as a separate chip with the Intel 80486 in the late eighties.
The good old days indeed!