The primary choices are space savings and power efficiency. Dedicated hardware and serial decoding/encoding often win out as a result.
You can encode the independent groups of pictures (the key frame and all the subsequent predictive frames until the next key frame) in parallel but at least with 4k video you hit memory bandwidth limitations quite quickly.
With adaptive key frame placement you don't know the intervals up front, and they might have wildly different lengths. Some might be hundreds of frames, some might be just a few.
It's the non-linear-algebra stuff that is usually problematic.
However, most formats allow for some parallelism at the "macroblock" level - you can usually decode all 16x16 pixels simultaneously. To some extent you can decode macroblocks separately, but "intra prediction" requires you to have the ones above and to the left available.
It reduces the compression efficiency, but only a little, because the tiles will be quite large anyway, such as 6 tiles per frame (but can be more for 360 video applications). It can be limited by the number of decoding instances a piece of hardware can have concurrently.
Seems like for VR video, if it’s split up longitudinally you would only have to decode the tiles that are in the direction the user is looking at any given time, and just let the other bits pass by undecoded.
mpv for example can use Vulkan for both at the same time.