If you can map all VRAM into host linear/physical address space, you can just use write combining page table entries to it and get efficient buffered write back caching for the CPU writes that created the very bytes you think of DMA-ing in the first place.
Write combined on cpu will batch a few bytes here and there but you’ll never hit full 4k transactions without dma. I brought up a few new gpus on pcie cards and spent many moons staring at interposer dumps.
It’s ok for small things but even once you get into the command buffer range it’s slow slow slow without dma.