Full-scale file system acceleration on GPU [pdf]
dl.gi.de
dl.gi.de
edit: blind as a bat, says so right in the paper of course:
PMem is mapped directly to the GPU, and NVMe memory is accessed via Peer to Peer-DMA (P2PDMA)
[1]: https://nvmexpress.org/wp-content/uploads/Enabling-the-NVMe-...
[2]: https://lwn.net/Articles/767281/
[3]: https://www.nvmexpress.org/wp-content/uploads/NVMe_Over_Fabr...
Once you got that then the CPU is just the orchesterator, and wouldn't necessarily need to be so beefy.
The PS5 and Xbox both have GPU-access of NVMe Flash.
------
So you are right. But what you are talking about happened like 5 years ago.
EDIT: https://devblogs.microsoft.com/directx/directstorage-develop...
Looks like 3 years ago for Win10. But I feel like I heard it sooner than that as NVidia or AMD specific API calls.
Also, Optane was like $4 per GB, so a moderately-sized drive, like 256GB, is already above $1000.
But this work used the Optane DC Persistent Memory DIMMs that only work with certain Intel server CPUs. I'm not sure what the typical price people actually paid for those was, but it probably was not actually more expensive than DRAM.
You don't want this kind of thing happening when it is running a filesystem.
Maybe someone could run workloads across CUDA and ZLUDA (Nvidia, and other hardware), but really we just might need more reliability to efficiently and reliability run a file system across disparate GPU hardware.
It would be interesting to know if this approach could optimize the performance of training and inference for large models.
>GpuRamDrive
>Create a virtual drive backed by GPU RAM.
https://github.com/prsyahmi/GpuRamDrive
Fork with AMD support:
https://github.com/brzz/GpuRamDrive/
Fork that has fixes and support for other cards and additional features:
seq1m 2205 2190 q1t1
rndq32 41.31 38.77
rnd q1t1 34.70 32.80
To be honest i didn't know what to expect, aside for a very high reading and writing speed. I was a bit disappointed in seeing random reading and writing were so slow, the only use i could think about would be having photosets or things like that over there, and then saving the session on ssd when closing the program, but it is easily solved by using a newer nvme ssd
1) How does this work differ from Mark Silberstein's GPUfs from 2014 [1]?
2) Does this work assume the storage device is only accessed by the GPU? Otherwise, how do you guarantee consistency when multiple processes can map, read and write the same files? You mention POSIX. POSIX has MAP_SHARED. How is this situation handled?
3) Related to (2), on the device level, how do you sync CPU (on an SMP, multiple cores) and GPU accesses?
Just quoting the paper:
>Using GPUfs, Silberstein et al . [ 24] demonstrate that offering a library interface to CPU FS eases access to storage for GPU programmers, but GPUfs only calls a CPU-side file system. GPU4FS offers a similar interface to GPUfs, but runs the file system on the GPU.
In this case, it is indeed novel to run the logic of the filesystem on the GPU itself. It's definitely worth the investigation!
Making the requests asynchronous and issuing lots of requests in parallel is what makes it possible to get good performance out of flash-based storage; P2P DMA would be a relatively minor optimization on top of that. DirectStorage isn't the only way to asynchronously issue batches of storage requests; Windows has long had IOCP and more recently cloned io_uring from Linux.
DirectStorage 1.1 introduced an optional feature for GPU decompression, so that data which is stored on disk in a (the) supported compressed format can be streamed to the GPU and decompressed there instead of needing a round-trip through the CPU and its RAM for decompression. This could help make the P2P DMA option more widely usable by reducing the cases which need to fall back to the CPU, but decompressing on the GPU is nothing that applications couldn't already implement for themselves; DirectStorage just provides a convenient standardized API for this so that GPU vendors can provide a well-optimized decompression implementation. When P2P DMA isn't available, you can still get some computation offloaded from the CPU to the GPU after the compressed data makes a trip through the CPU's RAM.
(Note: official docs about DirectStorage don't really say anything about P2P DMA, but it's clearly being designed to allow for it in the future.)
The GPU4FS described here is a project to implement the filesystem entirely on the GPU: the code to eg. walk the directory hierarchy and locate what address actually holds the file contents is not on the CPU but on the GPU. This approach means the application running on the GPU needs exclusive ownership of the device holding the filesystem. For now, they're using persistent memory as the backing store, but in the future they could implement NVMe and have storage requests originate from the GPU and be delivered directly to the SSD with no CPU or OS involvement.
(I worked on a FUSE filesystem that had these issues.)
I think the main benefit here is not having to do memory copies through the CPU, which frees up memory bandwidth for other things.
You can improve file-open overhead in conventional filesystems, too. Including the FUSE one I was working on.
are shaders turing complete ? ;)
Issuing individual truncates of 1B files can be just as much of a CPU problem then an IO one for example.