What he did was create a specialized compression algorithm that works very well to compress the data that is each frame of Blade Runner, and decompress it (lossily, like mp3) back into a video stream.
To put this into perspective: Blade Runner is 117 minutes long. At 25 frames per second, that is 175_500 frames.
As he says, the input data he used was 256x144 with 3 colour channels, meaning each frame was 110_592 bytes. This results in an input amount of 18_509 MB, uncompressed.
His neural network compresses each image down to 200 floats though, i.e. 800 bytes. So the whole movie as compressed by the NN is 132 MB.
A friend of mine who works with neural networks estimated his NN to be roughly 90MB at 256x144, so the storage needed for the movie he created is about 222 MB.
This means that if this technology can be made fast enough, and the reproduction high enough in fidelity, we could be looking at replacing compressors (Fraunhofer MP3, Lame, MP4, DIVX, JPEG) made to handle all types of input equally well; with compressors that can are so good for one specific set of data that compressor + data is smaller than anything the traditional compressors could create.
And even if its not fast enough, it can still be a very efficient compressor, trading size for CPU.
Plus, lastly, the NN itself can be compressed as well, and classic compressors can possibly be applied to the frame data.
This is well, hopeful but probably wrong ... because we already know how to make it smaller for this type of data.
The algorithms used for interframe/intraframe prediction are chosen to tradeoff speed vs size. If you built a really large-scale predictor that was able to generate very small representations for changes, you would get what he's built.
(Note that encoders already complexly select from tons of different prediction algorithms for each set of frames, etc)
We can already do this if someone wanted to. We just don't.
Because it's not fast enough (and nothing in that work changes this)
Would it be useful to apply NN's to video encoders to better select among prediction modes, etc. Probably. But that's already being done, and is not this guy's work.
"And even if its not fast enough, it can still be a very efficient compressor, trading size for CPU. "
The problem you have is not just compression time. It's decompression time. The bitstream of H265/etc is meant to be decodable fast.
What this guy is building is not. If you were to make it so, it would probably look closer to a normal video codec bitstream, and take up that much space.
In fact, he hasn't built anything truly new, he's just using existing papers and making an implementation. He also says, in his masters thesis, that is primarily an artistic exploration.
Even with hardware decoding, you can only make stuff so fast.
TL;DR While interesting, there are people working on the things you are talking about, and it's not this guy (at least in this work).
I would not expect magic here. We already create video codec algorithms by trading off cpu cost and size. The trick is trying to get better size without increasing CPU cost significantly. As these resources change (and remember, moore's law is pretty much dead), the video codecs will change, and videos will get smaller, but you aren't likely to see serious breakthroughs. We already could produce very small videos by applying tremendous amounts of CPU power.
And it doesn't need to work on every frame, it could pump out I-frames every 15 seconds or so.
e.g. see the screenshots here:
http://www.eteknix.com/blade-runner-gets-trippy-auto-encoded...
Anyway, the article doesn't mention this idea, the first hn discussion had someone claim some other ridiculous use of the fuzzy/muddy images that are supposed to mean something and ultimately I'd say it's an example of "my research project has cool images, I can use them for publicity".
It's not exactly unknown that autoencoders can reconstruct images, so I don't see why it's "impressive that it works at all."
Please don't ignore words other people write, or at least reread before hitting post. I did say possibly.
> I don't see why it's "impressive that it works at all."
I thought it was implicit from my explanation that the "works at all" includes the fact that the data sizes are reasonable already. If we had the same result, but the intermediary data was 18 gigabytes, then there would be nothing impressive about it. As it is, we're a lot below that, before further compression, so it is.
Remember: Prototype, Proof of Concept. You're looking at the first step, not the last step. You're looking at a motor carriage, not a Tesla.
Right, but why is that impressive if it doesn't actually result in a good reconstruction? I can take any collection of numbers and summarize it with the mean value, that doesn't imply that averaging is a good compression method.
If you consider this proof of concept, then what, exactly, concept does it prove? That statistics can represent a dataset?
The idea is you train the network to reproduce the input, but it has to pass the information through the small middle layer which has relatively few neurons. Thus the network learns a good (hopefully) way to represent typical inputs with not much information. You have automatically generated an encoding (hence the name).
You're right it's not really 'reconstructing' the film in the way they imply. But it isn't just a distorted version of the original either. The really important question is how big the middle layer is (and hence how much it has learned / how good the compression is). Unless this is just done for artistic purposes, which seems to be the case.
As for why it's important and not just another encoder like mpeg4, see my response at https://news.ycombinator.com/item?id=11830437
In deep dream, a network whose goal is to predict image categories is "run in reverse" and used as a generative model, but it was not intended to be generative from the outset. Many times autoencoders are intended to be generative, or at least learn relationships between images (sometimes called a latent space/latent variable), where classification tasks care more about just getting the classification correct - no explicit concern on learning relationships between categories/images in those groups.
An autoencoder with a much larger code space could be an improvement, or some newer works such as Adversarially Learned Inference / Adversarial Feature Learning [0, 1] or real NVP [2, 3] could probably do a much better job at the task, at the cost of increased computation.
Also something like inpainting parts of the frames with pixel RNN [4] would be interesting.
[0] http://arxiv.org/abs/1606.00704
[1] http://arxiv.org/abs/1605.09782
[2] http://www-etud.iro.umontreal.ca/~dinhlaur/real_nvp_visual/
In the future we may be able to colouring b/w movies and increase fps via this automatic software.