You wouldn't have to any more than you need 0 through N frames in memory to calculate frame N+1. Whatever your decoding state completes at frame N can be considered a key frame.
I wonder if caching semi-compressed frames would be more efficient in either case (CPU or GPU)
Edit: this doesn’t have to happen synchronously either, it can occur in a background thread or passively.