Alenka: GPU database engine
github.com
github.com
It reminds me of the big hoopla for GPU h264 encoders. When they came out everyone realized the quality was worse and not much faster.
Some things don't lend themselves to parallel processing, notably anything linear like transactions.
I mean yeah the GPU can sort a hundred billion items a second but how often do you really need to sort that many items using a database? In 99.9% of uses you have indexing or limits on the number of results.
Just saying, this program looks more like a stream processing platform with a SQL-like frontend than a full database
However, there are so many types of databases around. Lambda architectures are all the rage now - you keep one database for your transactionals, and another for analytics. Analytics are huge, in the multi-billions of dollars every year and they've become one of the most important parts of steering a business and deciding on new strategy. Larger businesses don't just 'go for it' anymore, they analyze, and inspect, and dig deep into their historical data to find out if something is worth doing.
GPUs tend to lend themselves well to analytics, contrary to transactions. Specifically, columnar databases. When the columns are all of the same data type, and the data locality is high, GPUs perform /very/ well.
Regarding your sorting point you may not really want to sort everything, you got that bit right. But what if you want to perform a `JOIN` on a bunch of data?
It makes more sense to sort it first, because the JOIN would be much faster - matching keys would be much easier.
Now, if you were performing really fast SORT on a GPU, you're saving precious processing time.
If you're running analytical workloads on big data sets, you're typically I/O bound to start with. It seems like managing moving little pieces of it back and forth to the GPU to compute is going to be a big PITA, add lots of little latencies, and gain you absolutely nothing. What am I missing there?
2. What if you only push indexes or similar up to the GPU, like an AB-tree index? You're keeping all of the 'heavy' stuff down, and only uploading a representation of it, to be later replaced with the actual data.
3. Think compression/decompression done on the GPU directly.
What about parallel queries over a Restriction-Union normalized data model?
I think the benefits would be similar in nature to columnar stores as you note:
> GPUs tend to lend themselves well to analytics, contrary to transactions. Specifically, columnar databases. When the columns are all of the same data type, and the data locality is high, GPUs perform /very/ well.
Data warehousing is a $30B/year business and it often boils down to a price/performance game. Any technology (including but not limited to GPUs) that can change the equation by at least an order-of-magnitude will be disruptive in that space.
...
> The precise goal of databases like Redshift, ...
Huh? It was my understanding that Redshift was a fork of PostgreSQL and is therefore not only transactional but also relational.
Simple Google search confirms transactions in Redshift: http://docs.aws.amazon.com/redshift/latest/dg/r_BEGIN.html
I am also confused by the implication that OLTP and OLAP are mutually exclusive. It is my understanding that they are not...
Just because a database can support transactions does not mean that it is geared for transactional workloads, and hence would likely not be called transactions.
OLTP+OLAP is called HTAP. Its an active area of research but to my knowledge there is still no silver bullet there - systems that do it often store data in both row and columnar format to ensure they are optimized for both, with associated overhead.
To that end, I always thought of OLTP vs. OLAP model and engine optimizations to be more or less orthogonal to whether or not a DB is transactional. I would even suggest that the inclusion of transactions in Redshift as justification for my seemingly unconventional view:
> Some PostgreSQL features that are suited to smaller-scale OLTP processing, such as secondary indexes and efficient single-row data manipulation operations, have been omitted to improve performance.[2]
But who knows, maybe I am barking up the wrong tree so-to-speak.
[1] https://en.wikipedia.org/wiki/Database_transaction#Transacti...
[2] http://docs.aws.amazon.com/redshift/latest/dg/c_redshift-and...
No, it is not transactional.
Yes, it is relational.
Are there any GPU (shader) encoders?
There are dedicated fixed-function hardware encoders (VCE and NVENC) on graphics cards, and they're very popular, because they're the best way to record game footage if you don't have dedicated capture hardware. You don't want x264/5 hogging the CPU when you're playing games!
Not everybody can afford that setup, and I think even an 8-core will struggle with 4K…
The target for most streamers is Twitch 1080p H.264, because that's the platform where the money is made. 4K is not needed or useful, and H.265 won't help simply because twitch won't stream it.
It seems to me like there are mainly two kinds of cost-limited database workloads - complicated operations on datasets that fit in RAM, or simple operations on datasets that need to be distributed. For the former you're better off writing custom software, and for the latter you're I/O rather than CPU limited... Maybe SSDs or low-power-high-memory ARM servers could change that equation
And today's NVENC h264 encoding quality is approaching x264 levels of quality, until you get to the lower end of the bit rate spectrum. x264 really shines at bpp values of 0.05 and lower, a feat NVENV has yet to achieve.
I can't speak for the h264 situation, as I wasn't involved at the time, but with the newer h265 hardware, it is obnoxiously faster with minimal (in terms of end-user, not distributor) quality loss.
I remember when h264 CPU accelerators hit the scene. It was a game changer.
1. GPU RAM storage isn't as fine grained as CPU virtual memory. While GPU's have virtual memory they don't have the same degree of Copy On Write hardware utilities. This make rolling back snapshots and transactions within memory very difficult.
2. GPU's don't have dedicated non-volatile storage (well some do, but it is used more as a cache, and is treated as volatile). So for loading data you:
SSD -> RAM -> CPU *copy* CPU -> RAM-> GPU
This isn't really a question of technological maturity it is more a question of PCIe doesn't allow an NVMe SSD to talk directly to a PCIe GPU. Nor do modern OS's have any model how to do this, nor do GPU's support file systems.3. SQL based data querying is parallel friendly. At it's core SQL is 99% Filter/Map operations. The original goal of SQL was to be bottle necked by HDD access times so the vast majority of the query work is done in O(n). With a
SSD -> RAM -> CPU *copy* CPU -> RAM-> GPU
You are already paying that O(n) load+process price at copy time (minus branches).So the savings are only present if the data can persist on the GPU. Ultimately having multiple CPU threads to do the Map/Filter do the same job
There isn't just one problem:
1. GPU's don't support features to make holding data in RAM useful
2. PCIe doesn't support features to make loading data into the GPU fast.
3. The very design of SQL makes copying data into the GPU moot. Spreading a Map/Filter over 10-20 threads isn't rocket science.
Re CPU vs GPU performance you are neglecting the fact that GPU RAM can be an order-of-magnitude faster than CPU RAM, not to mention the fact that complex queries can often become compute bound (geospatial queries are a prominent example).
Why couldn't CPU RAM be used as well in a GPU database system?
There is no market demand for this currently. GPU's don't need advanced MMU's because nobody wants them. The same is true for PCIe data loading. CPU vs GPU performance you are neglecting the fact that GPU RAM can be an order-of-magnitude faster than CPU RAM,
No I'm not. If your data set can remain in GPU ram long term yes there is a big performance speed up. The problem here (as I outlined before)1. GPU MMU's are less advanced, so they don't handle transactions well (see above comments).
2. As you can't perform transactions efficiently, you need read only data. Which is possible but many databases aren't read only, or databases NOT being written too are rare.
3. Ensuring your full dataset can fit in GPU memory (typically <20GB) is a damning limitation. Streaming data into the GPU is rather sloppy and limits compute though-put.
There's a well-known way to increase network throughput and reduce latency by running a whole dedicated IP stack + hardware driver in the user space of the process that needs it. This removes the bottleneck of the OS kernel.
I wonder if there's a way to remove the bottleneck in the GPU case by something similar, by dedicating a piece of hardware to the SSD interconnect without having a CPU as an intermediary. AFAICT you still cannot DMA from a disk directly to GPU RAM, even though you have enough PCIe lanes coming to the GPU. Is this true? If so, can it change in a way that is still compatible with traditional, "normal" operation of the PC architecture?
Still, it is pretty cool technology and I think something that will be integrated with a "real" database near you sometime soon.
But there really isn't a brightline between what's a table, and what's a filesystem, and what's a database. You can put a blob in a database or BigTable, and MongoDB is really close to being a flat file conceptually. You can have a filesystem or database that is content-accessible. You can have a filesystem that is atomic and supports rollbacks and can store relational data like symlinks. A virtual filesystem like LVM can support schema-like volumes on top of it.
At the end of the day it's all just technology that lets me abstract my writes so I can deal with a more simplistic model backed by certain guarantees about behavior. I want to write a program that does XYZ, not write a filesystem/database driver. From there it's all just various tradeoffs.
https://aws.amazon.com/marketplace/pp/B01M0ZY2OV?qid=1484735...
So, in a nutshell, is it possible to run CUDA on a non-nVidia GPU nowadays (AMD or else)?
CUDA → HIP translator. HIP is an abstraction over CUDA and AMD's HCC: https://github.com/RadeonOpenCompute/hcc/wiki
There's also this: https://github.com/vtsynergy/CU2CL but I know nothing about it.
The exception is littered throughout bison generated code: "This special exception was added by the Free Software Foundation..."