Is this true for dgpus? How does this work?
Is this true for dgpus? How does this work?
On the device (dGPU here), it is possible to route memory accesses to part of the internal address space to the PCIe controller. In turn, the PCIe controller can translate such received memory access into a PCIe request (read or write), in the different PCIe address space, with some address translation.
This PCIe request goes to the PCIe host (CPU in a dGPU scenario). Here too the host PCIe controller can map the PCIe request, using using a PCIe address space address, into the host address space. And this can go to the host memory (after IOMMU filtering and address translation usually). And all this back for the return trip to the device in case of a read.
So latency would be rather high, but technically possible. In most application such transfers are offloaded to a DMA in the PCIe controller doing a copy between PCIe and local address spaces, but a processing core can certainly do a direct access without DMA if all the address mappings are suitably configured.
If I map some host memory to the GPU… I get worse latency and worse bandwidth. Most likely not a win.
GPUs have been able to access "host" memory for a long time now, with a few restrictions: you have to setup the GPU mappings first and pin the pages in memory.
I say in theory and used an asterisk because I think it's generally the case that the driver could lie and just maintain an illusion for you by flushing a staging buffer at the 'right time'. But in practice my understanding is that the memory writes will go straight over the PCIe bus to the GPU and into its memory, perhaps with a bit of write-caching/write-combining locally on the CPU. It would be wise to make sure you never read from that mapped memory :)
OpenGL drivers have a habit to try to second-guess the application (though this depends on the driver, eg. Nvidia guesses a lot, Mesa not so much), but passing/not-passing GL_CLIENT_STORAGE_BIT to glBufferStorage should decide whether buffer should reside in CPU/GPU side memory.
In D3D11 directly mapped GPU memory is known as D3D11_USAGE_DYNAMIC.
Seems to be that there are other ways now, to achieve this, on lower levels in the hardware, 'transparent' to the layers above. 3D-Vcache stacked on top of the die, and https://www.techarp.com/computer/amd-infinity-cache-explaine... comes to mind.