Apple Patent Shows GPU Dynamic Caching Has Been in Development for Years
tomshardware.com
tomshardware.com
Isn't this not too different from what Intel DVMT was doing twenty years ago?
https://en.wikipedia.org/wiki/Dynamic_video_memory_technolog...
Then again, this is the same Apple that took "integrated graphics" / "UMA" as a marketing point.
In any case, in Intel's DVMT you have a dedicated memory space for graphics. It is a dynamic amount, but it is still dedicated for graphics. You need to copy data from system memory over to graphics memory for it to be handled.
In UMA, no data copying is needed. There's 1 memory space
https://www.intel.com/content/dam/develop/external/us/en/doc...
https://en.wikipedia.org/wiki/Heterogeneous_System_Architect...
Yes, they do, it's mentioned as some major M1 revolutionary technology all the time. It's utterly absurd, but that's the power of Apple's marketing.
If it OOMs, then one valid assumption is that it OOMs because it has less memory available than there is for a GPU on M1 but we can't know that if you fail to mention the specs.
No you don't. You can map any page of physical RAM into the GPU address space.
This has been the case since at least the Intel i810. Here's a relevant part from the i815 documentation:
But aren't caches already a dynamic copy of main memory? Don't GPUs already prefetch buffers as needed? What is actually new here?
I know this might seem like a lifetime for SW people but <3 years from patent to customer for a GPU HW innovation isn't really that long? It's pretty typical, really.
However, I'm always suspicious of how innovative patents actually are so it wouldn't surprise me to find out it's just barely different than what CUDA or unified memory on game consoles can do.
There's the linked patent itself [0] - but that specifically seems to refer to using the MMU as the part that's doing this dynamic allocation and translation - but the majority of those resources above are "before" address translation, with the MMU often being more at the L2 cache level rather than embedded within the shader clusters themselves. Maybe this isn't the "normal" MMU and instead a simpler address translator specifically for those resources that has been added of for this? Or is it something like the parameter buffer (The block of memory used to store the intermediate data between the tiling & rasterization state, and the pixel shaders). That has been "dynamically allocated" in a similar way since Apple were just taking PowerVR cores directly.
And then what is the cost of this re-allocation? If there's a cost to performance to allocating new pages, or (like worse) running out of spare pages, it might mean the (graphics API) user still has to be aware of their resource usage in scheduling shaders, so less of the promise of "Just throw things at the hardware and it'll do things optimally" than people might hope.
It seems to me that every explanation of dynamic caching in terms of memory is "wrong" - as seen here and in several articles written by folks more familiar with PC hardware.
I think where some folks might get it wrong is thinking of Apple silicon as being like PC hardware, where VRAM and RAM are distinct pools of memory and using the CPU to move data between them (or other devices) is very inefficient. PCI bus attached devices having separate pools of memory has given rise to a plethora of technologies to allow directly read from RAM (DMA), or to enable a GPU to read from NVMe (DirectStorage on Windows), and so on.
The patent from the article seems to describe unified memory, part of the M1's architecture, not whatever "Dynamic Caching" is; but I'll admit Apple makes it a bit hard to understand what exactly is the case.
There's only one person I'd trust to describe what this feature actually is - I'll wait for Asahi Lina to break down what dynamic caching is and whether this is a hardware or OS feature.
[1] https://devblogs.microsoft.com/directx/hardware-accelerated-...
The patent basically describes using a page table to dynamically map pages of RAM between the GPU and CPU, something that Intel's integrated GPUs have been doing for 20 years.
Almost EVERY patent is basically something that has been done for many years, plus some improvement. Improvements are usually inconsequential, but sometimes they are very valuable.
It's possible that the latest gen of RT cores are a slightly larger portion of a TPC (but we don't know), but on the other hand, TPCs are a smaller portion of a die now as TPC arrays do not include L2, and L2 got many times bigger. So chances are, RT over die area is still somewhere on the order of 5%.
And in Nvidia designs as far as I understand RT cores sit idle in non-RT workloads, yes.
If anything, Tensor cores are almost twice as large as RT cores, and tensor cores sit idle whenever you don't use DLSS (or professional ML compute).
See https://www.reddit.com/r/hardware/comments/baajes/rtx_adds_1...