cuDF – GPU DataFrame Library
github.com
github.com
Unless cudf has implemented some clever dask+cudf kind of situation which can intelligently push data in/out of GPU as required?
What you can do on even consumer GPU's is mind blowing.
Open, mmap and then import the host pointer to GPU API (Cuda or Vulkan). Paging should work as expected, but "stupid" access patterns hurt throughput.
This requires a GPU, a CPU and a kernel driver that supports it. E.g. my laptop Intel GPU has the host pointer import Vulkan API but last I checked it does not support pointers imported from mmap.
https://developer.apple.com/metal/jax/
And MLX
[1]: https://github.com/google/jax/issues/16321 [2]: https://github.com/google/jax/issues/17490
See discussion: https://news.ycombinator.com/item?id=39930846
I've been watching cuda since it's introduction and Polars since I had an intern porting our Pandas code there a couple years ago but I had no idea Polars would go this far, this fast!
They made GPU processing at scale accessible to everyone, I have been a long term user of Rapids and found that even as a data engineer I can do things on an old consumer GPU that would otherwise require a 20+ node cluster to do in the same time.
The ability to run coffee 100-1000x faster with this is just icing on the cake.
(I've run through this with most of my Pandas training material and it just works with no code changes.)
I like pandas, and python.
We have like 12 different types of it in the wild. I think it's time we came up with a 1 or 2 GPU HW standards similar to how we have for CPUs.
I get it: some of these are legacy, others are hand optimized python since default pandas is so slow. But I'm hoping that, over time, we'll improve the runtime of the other stages of analysis too.
HoloViews hvPlot Datashader Plotly Bokeh Seaborn Panel PyDeck cuxfilter node RAPIDS