It allows the CPU (or any other device in the PCI bus) to write/read data to/from the device when it’s convenient for the sender, without having to interrupt whatever the receiving device is currently doing.
You still need to coordinate with the device to tell it that you’ve written to its memory, or read from it. But that’s a pretty cheap operation.
The alternative is the both CPU and GPU would have to stop what their doing and manage the data copy while doing nothing else.
So it’s basically the difference between sending someone an email vs giving them a call.
With DMA your emailing a big document, the later calling to make sure they received it. Without DMA it would be like calling them and reading the document out over the phone. One is clearly better for everyone’s productivity.
I could still read this two ways however: one where the memory is on the peripheral, and one where the memory is main memory, where the peripheral is copying to/from using DMA. Which one is it?
You send the device a circular list of descriptors (pointers) to a region of main memory.
In order to send data to the device, you write your network packet to the memory region associated with the pointer of the current ‘head’ of the descriptor list.
So far, you have a ring of pointers, one of those pointers points to a location you just wrote to in ram.
You then tell the device that the head of the list has changed (as you just wrote some data to the region that the head of the list is pointing to - so it can consume that pointer), the device then goes ahead and copies the data from ram into an internal buffer on the card. Once the data is consumed, the tail pointer of the ring buffer is updated to indicate that the card is finished with that memory region.
> __padding 45 minutes ago [dead] [–]
> Typically with devices like network cards (that also operate over PCI-E) You send the device a circular list of descriptors (pointers) to a region of main memory. In order to send data to the device, you write your network packet to the memory region associated with the pointer of the current ‘head’ of the descriptor list. So far, you have a ring of pointers, one of those pointers points to a location you just wrote to in ram. You then tell the device that the head of the list has changed (as you just wrote some data to the region that the head of the list is pointing to - so it can consume that pointer), the device then goes ahead and copies the data from ram into an internal buffer on the card. Once the data is consumed, the tail pointer of the ring buffer is updated to indicate that the card is finished with that memory region.
But equally a peripheral can expose its own memory and ask the host to write into it.
Cheap devices tend to do the former because it avoids the need to have expensive memory built in. They can just “borrow” system memory. More expensive, performance optimised, devices tend to do the latter.
It’s also worth mentioning that DMA tends to work between every device attached to the PCIe bus. So Microsoft’s DirectStorage API seems to be using this feature, by having the GPU directly read data from an SSD, without the data ever touching the CPU or main memory.
As others have said, you'd normally configure the MMIO space for uncached access, or you'd need to be careful to force the memory ordering you need. The device specific interfacing requirements would be the guide there. Devices can indicate if their MMIO ranges are prefetchable or not, which should indicate if stray reads would cause side effects or not.
One bonus of MMIO is DMA could interface with other devices, whereas I don't think devices are allowed to drive the I/O bus like that.
It's interesting to compare to embedded processors without a memory management unit, like this STM32 reference manual, see p.68 and following:
https://www.st.com/resource/en/reference_manual/dm00124865-s...
Everything looks like a memory address. Note that it's not actually memory, the processor just diverts requests for that address to the peripheral instead of memory. But on that little ARM processor, if you want to write to actual RAM, that's memory addresses 0x2001 0000 to 0x2001 BFFF. Data in the onboard Flash memory is at 0x0020 0000 - 0x002F FFFF. If you want to talk to something on a serial port, write to registers from 0x4001 1000 - 0x4001 13FF. If you want to show something on an attached LCD, or pull a buffer from the USB or Ethernet peripherals, or work with GPIO, or do anything at all, really, it's at some memory offset. This chip has some DMA, you can set it up to automatically push from one peripheral memory space to actual RAM or vice versa. But everything happens at a region of memory.
This is perhaps the DMA scenario you describe in the last two lines, my thinking is that it would make sense to do this all the time, at least when transfers are large.