DeepZip: Lossless Compression Using Recurrent Networks [pdf]
web.stanford.edu
web.stanford.edu
1. https://www.facebook.com/notes/blake-ross/aphantasia-how-it-...
For non-images, this idea is pretty old and sounds a lot like compression algorithms that share a pre-defined dictionary. For example: https://en.wikipedia.org/wiki/Brotli, "improved the compression ratio by using a pre-defined dictionary of frequently-used words and phrases." Of course words/phrases are a lot easier to predefine.
I wonder how large the predefined weights of the NN has to be to effectively compress real world images?
On the other hand, in "The Case for Learned Index Structures" [1] paper, they replace B-Trees with neural nets:
> by using neural nets we are able to outperform cache optimized B-Trees by up to 70% in speed while saving an order of magnitude in memory over several real world data sets
Basically ML/DL is coming for the classic data structures, replacing some of them with learning based versions.
The paper doesn't explain how this is possible, how would I go about implementing something like this?
This can be likely optimized (e.g. 7-Zip's arithmetical coder doesn't split data into batches, but rather uses exponentially filtered conditional probabilities that are instantaneously updated each time a new sample is processed). I'm no expert in RNNs, but my hunch is that you can adjust the model incrementally using backpropagation right after it processes another sample instead of completely retraining it.
So, to decode you would start with the unoptimized model and decode the first batch. You then train using that first batch to get a better model and use your improved model to decode the next batch.
Okay, that makes sense; thanks for the response!
This seems like a bold statement. We know they are good at capturing some structure from sequential data at least, but they seem lacking in some regards [0].
However it seems to yield some interesting results. The architecture is like an autoencoder, but with RNNs connected at each layer? Have you tried using NTMs?
[0]: https://rylanschaeffer.github.io/content/research/neural_tur... (section on LSTM Copy Performance)
[2] EE376C Lecture Notes: Universal Schemes in Information Theory: http://web.stanford.edu/class/ee376c/lecturenotes/intro_lect...
> Only random seed is Stored: The model is initialized using a stored random seed, which is also communicated to the decoder. Thus, the effectively the model weights do not contribute to the total compressed data size
> Single Pass: As we do not store the weights of the model, we perform a single pass (1 epoch) through the data