Understanding x86_64 Paging
zolutal.github.io
zolutal.github.io
Page Cache Disabled (PCD) – pages descendant of this PGD entry should not enter the CPU’s cache hierarchy, sometimes also called the ‘Uncacheable’ (UC) bit.
I guess this is some optimisation where it's decided the page isn't worth taking up valuble CPU cache state. And now I'm wondering what this algorithm is that decides it..
For example GPU drivers and texture copying - even on integrated GPUs that have cache coherency with the CPU (so don't "need" to manually flush it), there's no point filling the CPU cache with data you're only touching once with effectively a memcpy() on the CPU, you'll just evict more useful data you'll probably need to readback after anyway.
For anything that isn't MMIO, you would prefer using ordinary WB caching and non-temporal stores to avoid populating the cache instead of UC mode.
> Writes to MMIO space allow the CPU to continue before the transaction reaches the PCI device. HW weenies call this “Write Posting” because the write completion is “posted” to the CPU before the transaction has reached its destination.
I like this series as an intro: https://xillybus.com/tutorials/pci-express-tlp-pcie-primer-t...
Not even newer, but instead it's a pretty common feature for GPUs for the past couple decades or so.
DMA blocks can only do so much when the (often older but still well used) APIs don't map well to the synchronization required - either allowing the user to immediately free and/or reuse the buffer used to pass in data (so likely requires a synchronous CPU copy to a staging area before the API function returns), or direct memory mapping of resources. Putting either of those in the cache is often wasteful, as it's unlikely for any line to be re-used before being flushed anyway, then passed over to the GPU DMA block to do whatever asynchronously.
And there's also non- device driver use cases - I've seen image processing libraries intentionally skip the cache if they know they're not going to be touching the data again for some time, and the data set itself is large enough. I assume other users exist, I just see those as I work on GPUs, and images are a big source of large data sets.
WC allows these to avoid clobbering the cache while having at least some chance of using the memory/pcie bus effectively.
RISC-V is quite boring, majorly opting for well-proven approaches. And a lot of its technical value comes from doing just that.
Yet x86-64 and ARM would have a hard time justifying their complexity, when RISC-V achieves the same with simplicity, and matches or beats them on objective metrics[0].