https://b3148424.smushcdn.com/3148424/wp-content/uploads/202...
https://b3148424.smushcdn.com/3148424/wp-content/uploads/202...
In practice a combined Av1/H264/etc. decoder core is most likely. A lot of logic would be shared.
The die area is very modest, but the hard part is building it in the first place.
Encoding is more area, but should still be peanuts for Apple SoCs.
There are also likely opportunities for additional efficiencies when you make a custom {en,de}coder for your system. I suspect (but haven't confirmed) that the typical Intel/AMD/Nvidia/Apple multi-function media engine isn't just a collection of completely independent encoder/decoder blocks for each codec but a kind of simplified specialized microcoded CPU with a collection of fixed-function blocks which can be shared between different codecs. So it could have blocks that do RGB->YUV conversion, Discrete Cosine Transforms, etc. and you can use the same DCT block for AV1, HEVC, and AVC. Maybe you can also create specialized efficient ways to transfer frames back and forth with the GPU, for sharing cache with the GPU, etc.
I believe the team we worked with at Arm during AV1 standardization is no longer there, which is too bad. They were really great guys to work with.
Your suspicion is mostly correct, though obviously you cannot share too much of the DCTs as these must be bit-exact and are different for each of the standards. But especially things like the compressed tile cache for reference frames used in motion compensation are extremely complicated (to save memory bandwidth and power) and entirely shareable. The SRAM used for line buffers is also a lot of area and shareable. And so on.
On a GPU, you didn't have an option to interleave normal program stream with specialized partial-decoding instruction. You put encoded frame in and you get decoded frame back, the media engine was separate block from compute.
Though this is also changing; see Intel GuC firmware, which has (optionally) some decoding, encoding and processing based on compute.
That being said, it's pretty common to have dedicated silicon for video codecs. It normally takes the form of a little DSP with custom instructions to accelerate operations specific to the codec.
I'm guessing this is a distinct region of the chip and not integrated with CPU/GPU since they scale up by replicating those blocks and wouldn't want to redundantly place that hardware. Having it separate also allows a team to work on it independently.
I think the relative size of the media engines is accurate in that slide, so then it comes down to how large the ProRes parts are in other chips. They are probably a couple of the unlabeled regions next to the performance cores in the M1 Pro die shot below, but I don't know which.
https://images.anandtech.com/doci/17019/M1PRO.jpg Taken from: https://www.anandtech.com/show/17024/apple-m1-max-performanc...
And video decoding/encoding is definitely at least GPU-adjacent, since it usually also involves scaling, color space transformations etc.
Maybe I'm misunderstanding what you're saying, but the slide is of an A17, not an M3 chip.
The A17 floorplan looks nothing like that either.