Autoencoding Blade Runner: reconstructing films with artificial neural networks
medium.com
medium.com
It could be getting this wrong if his error function is calculating linear data from the given image pixels, which are in the totally not linear sRGB colorspace. That would make it badly underestimate any error in a dark image.
Quick check of the PIL docs doesn't mention gamma compensation, so they probably forgot about it. People usually do.
According to the defaults in the code, it uses float32 arrays of the following sizes:
image: 144 x 256 x 3 = 110,592
code: 200
Note that the sequence of codes that the movie is converted to could possibly be further compressed.[ed: does indeed appear from the github page, that the input is a series of png frames, and the output is the same number of png frames, filtered through the neural net. No compression, but rather a filter operation?]
I think it's doing something like this (but more complex) https://cs.stanford.edu/people/karpathy/convnetjs/demo/autoe... where you have a bottleneck on the network
If it is the compression case, I am curious for the size of the compressed movie.
What would be some good resources for 1. getting the bare minimum knowledge required for using existing libraries like Tensorflow 2. going a bit further and having at least some basic understanding of how most popular ML/AI algorithms work ?
I guess I feel like there's no practical result here. It's only interesting from an aesthetic point of view.
Am I being unfair?
If they're trying to create interesting swirly stuff, where do they intend to go after that?
I mean, sure it's aesthetic though not on the level of weirdness of deep dreams modification.
Would be interesting if somebody one of these days can actually reconstruct to a high level of fidelity what our brain is "seeing".. I bet it would look kind of like this..
There's probably newer research out there too.
Now back to the article, can someone explain about how many passes before it gets to near film quality? Can it extrapolate missing frames eventually?
Lossy compression with super high compression rates?
In practice, even the hyper-efficient compression algorithms used in something like zpaq tend to use only very small shallow predictive neural networks because no one wants to wait days for their data to be compressed or ship around big neural nets as part of their archives, so it's more of an information-theoretic curiosity. Few enough people will even use 'xz'.
Professionals like to use something they can understand, and when making a BD or streaming source they know what a compression artifact looks like and which frames' bitrates to tweak to hide it. They pretty much sit there all day and just do that.
Maybe live streaming?