Awesome video. This helps to show how the Q*K matrix multiplication is a bottleneck, because if you have sequence (context window) length S, then you need to store an SxS size matrix (the result of all queries times all keys) in memory.
One great way to improve on this bottleneck is a new-ish idea called Ring Attention. This is a good article explaining it:
https://learnandburn.ai/p/how-to-build-a-10m-token-context
(I edited that article.)