GPU Caching Compared Among AMD, Intel UHD, Apple M1
chipsandcheese.com
chipsandcheese.com
What do you mean by this? They do support integers from 8 bits and up and natively support 16-bit floats. Are you referring to something else, like 8-bit floats?
But other GPUs support doing operations on limited-precision data types faster. And Nvidia has dedicated matrix multiplication units that can perform very wide limited precision operations per cycle (Apple has similar units but they are part of the CPU clusters).
Apple since A15/M2 offers SIMD matrix multiplication intrinsic (very similar to VK_NV_cooperative_matrix). But the performance is limited by the fact that each SIMD only offers 32 ALUs. If they add the ability to reconfigure these as 64 FP16 ALUs (or 128 FP8 ALUs) and then maybe even doubled the ALUs like Nvidia/AMD recently did with their architectures, they could achieve much higher matmul performance for ML.
I have very limited knowledge of these things but I did compare matmul with PyTorch using GPU vs not and there is a dramatic improvement. So even if its not fully optimised yet, it's still a huge bonus to have this available. If it could be improved another 4-6x that would be stupendous.
(For context: am observing an 8000x8000 matrix takes ~1s to multiply with CPU and 5ms to multiply with GPU).
Is this true for dgpus? How does this work?
On the device (dGPU here), it is possible to route memory accesses to part of the internal address space to the PCIe controller. In turn, the PCIe controller can translate such received memory access into a PCIe request (read or write), in the different PCIe address space, with some address translation.
This PCIe request goes to the PCIe host (CPU in a dGPU scenario). Here too the host PCIe controller can map the PCIe request, using using a PCIe address space address, into the host address space. And this can go to the host memory (after IOMMU filtering and address translation usually). And all this back for the return trip to the device in case of a read.
So latency would be rather high, but technically possible. In most application such transfers are offloaded to a DMA in the PCIe controller doing a copy between PCIe and local address spaces, but a processing core can certainly do a direct access without DMA if all the address mappings are suitably configured.
If I map some host memory to the GPU… I get worse latency and worse bandwidth. Most likely not a win.
GPUs have been able to access "host" memory for a long time now, with a few restrictions: you have to setup the GPU mappings first and pin the pages in memory.
I say in theory and used an asterisk because I think it's generally the case that the driver could lie and just maintain an illusion for you by flushing a staging buffer at the 'right time'. But in practice my understanding is that the memory writes will go straight over the PCIe bus to the GPU and into its memory, perhaps with a bit of write-caching/write-combining locally on the CPU. It would be wise to make sure you never read from that mapped memory :)
OpenGL drivers have a habit to try to second-guess the application (though this depends on the driver, eg. Nvidia guesses a lot, Mesa not so much), but passing/not-passing GL_CLIENT_STORAGE_BIT to glBufferStorage should decide whether buffer should reside in CPU/GPU side memory.
In D3D11 directly mapped GPU memory is known as D3D11_USAGE_DYNAMIC.
Seems to be that there are other ways now, to achieve this, on lower levels in the hardware, 'transparent' to the layers above. 3D-Vcache stacked on top of the die, and https://www.techarp.com/computer/amd-infinity-cache-explaine... comes to mind.
Later
>Intel: 700, AMD 1400, Apple: 2100
I wouldn’t call 2x and 3x “similar”.
Also I don’t see why author thinks desktop chips with integrated graphics are meant to be paired with a discreet GPU. Surely the opposite is true. I got a faster CPU by not getting one with integrated graphics.
Finally, doesn’t the fact that apple has a fundamentally different rendering pipeline relevant?
On a tangential note, it's great having an iGPU even if you are almost never going to use it. If your discrete GPU borks, you have a fallback ready and waiting. If you do use it alongside a discrete GPU, you can offload certain lower priority tasks like video encoding/decoding to it.
Now AMD's desktop chiplet-based CPUs have a tiny GPU in the IO/memory controller die, ill-suited to anything more advanced than everyday web browsing.
also not meant for anything but office use, debugging ease of live and maybe offloading some dedicated GPU task to the included decoder in the future.
there is still a good chance we will see some APUs soon like a 7700G it will be interesting to see if they will be Zen 4.
Is it still all that fundamentally different? All of the RDNA parts are tile-based renderers (I think even the Vega series GCN parts made that switch?)
Apple (inherited from PowerVR) adds another twist on top: the rasterised pixel are not shaded immediately but instead collected in a buffer. Once all fragments in a tile are rasterised you basically have an array with visible triangle information for each pixel. Pixel shading is then simply a compute pass over this array. This can be more efficient as you only need to shade visible pixels, and it might utilise the SIMD hardware better (as you are shading 32x32 blocks containing multiple triangles at once rather than shading triangles separately), plus it radically simplifies dealing with pixels (there are never any data races for a given pixel, pixel data write-out is just a block memcpy, programmable blending is super easy and cheap to do) — in fact, I don't believe that Apple even has ROPs. There are of course disadvantages as well — it's very tricky to get right and requires specialised fixed-function hardware, you need to keep transformed primitive data around in memory until all primitives are processed (because shading is delayed), there are tons of corner cases you need to handle which can kill your performance(transparency, primitive buffer overflows etc.). And of course, many modern rendering techniques rely on global memory operations and there is an increasing trend to do rasterisation in a compute shader, where this rendering architecture doesn't really help.
Edit: I just had another look, pretty sure this is standard Tile-Based Immediate Rendering. The documentation sometimes refers to this as "deferred" probably because copying of the final image values to the RAM is deferred. But "deferred" in TBDR means "deferred shading", not just "deferred memory copy". Adreno does not do deferred shading.
So there isn't that much value in the effort to benchmark them.
There are some exceptions for low end gaming systems, some more office use cases and some AIO use-cases e.g. the G-series amd processors like the 5700G. They tend to have GPUs noticable faster then what intel integrated graphics has in the same generation but also noticable less the dedicated graphics.