You can try it out yourself here https://colab.research.google.com/drive/1r7S15Ie33yRw8cmET7_...
Or use dockerhub https://hub.docker.com/r/blazingdb/blazingsql/
The benefits are.
Greatly increased processing capacities. We can just perform orders of magnitudes more instructions per second than a cpu with the gpus we are using.
Decompression and parsing of formats like CSV and parquet happens in the GPU orders of magnitude faster than the best cpu alternatives.
You can take the output of your queries and provide it to machine learning jobs with zero copy ipc and get the results back the same way. We are all about interoperability with the rapidsai eco system.
// sorry if this is a stupid question.
Or do you mean being able to read a database's file format natively? If this is the case there are many reasons. 1. There are many poorly/non documented formats 2. Even if you decide to read some other DB's format natively, those formats change over time 3. Little control of how and where the data is laid out
Is it distributed? How do I set it up in a distributed mode? Does it support nested parquet (something that even spark itself struggles to support inside SQL).
Right now we use k8s on Google K8s Engine(GKE) to deploy in distributed mode.
We don't supported nested at present, there are Rapids teams looking into this.
In summary, you get snappy, interactive query speeds on large data sets. I've ran that locally and the results are pretty amazing compared to Postgres or even Tableau in-memory.
I'm personally more excited about GPUs in stream processing; its just quite a natural fit: https://github.com/rapidsai/cudf
https://fastdata.io/plasma-engine/
* It's not open-source and I work there.
We've also optimized our storage formats and multithreaded our disk reads, such that we can easily hit many gigabytes per second on flash storage. Plus, new persistent memory technologies like Intel Optane will enable even more instant reads from "cold" storage.
People have been building columnar databases to do analytics quickly. GPUs (with CUDA) can run analytics operations (think join, group by, math, sorting) on columnar data in a much more efficient manner. They're designed for operations on vectors, which columns are.
We've been doing this ourselves too with SQream DB: https://sqream.com. It's an enterprise data warehouse with GPU acceleration. We use CUDA exclusively too.
Also from 4 years ago