Limited by VRAM is a huge constraint for me. Even if it is slower, being able to load 100GB+ into RAM without any batching headaches is worth a lot.
Unless cudf has implemented some clever dask+cudf kind of situation which can intelligently push data in/out of GPU as required?