If you can map all VRAM into host linear/physical address space, you can just use write combining page table entries to it and get efficient buffered write back caching for the CPU writes that created the very bytes you think of DMA-ing in the first place.