What
should be doable:
Deconstruct a movie into a set of 3D models, textures, voices, parameters for those 3D models' movement, procedural generation of background objects like trees, grass, sea, sand, weather conditions, etc, etc. Like the data that forms the content of modern, near-photorealistic games.
A game engine-like rendering system would then render a scene, tweak parameters until rendered scene / frames match the original movie closely, and work through the rest of the movie the same way.
When done, one could produce derivations of such movie by changing actors' voices, have them move differently or speak in another language, swap out buildings or other objects, reduce polygon count or resolution of textures, have a cornfield show a little taller stalks, put the sun in a different angle, etc, etc.
Of course this is way beyond current compute capabilities. Not to mention software frameworks to do this job. But in theory this should work. For a 2min trailer it would probably be pointless. But for a 2..3h movie, maybe not.
And yes, of course this would be lossy 'compression'. Just more high-level than current video compression schemes.
For an audio analogy: compare mp3 compression with MIDI + quality sound banks for every instrument under the sun + parameters like how hard a piano key was struck, etc. Vary such high-level parameters until rendered output matches the original music.