How to run a state-of-the-art image classifier on an iPhone
libccv.org
libccv.org
I thought this was a cool hack for reducing the memory size of the model. Is this an established technique for storing parameters? Are there known better ways to compress the model?
PNG optimizers are not exactly designed for compressing data-as-images.
"Obviously that isn't true for NN parameters." I have little experience with NN's but I wonder if that's true for deep NN's? Is there any gradual, image-like change in parameters in deep learning models?
There may be regularities in deep NNs. For example, many learned features are basically exactly the same but rotated or translated in a different way. This is because NN's aren't inherently invariant to rotations or translations, so they have to learn multiple features for every possible variation. Another possible regularity is that the weights tend to cluster on specific sections of the input image and ignore the rest.
But this entirely depends how the weights are projected into an image. The image provided looks like random static. The standard method is to just create a separate image for every feature, and for every pixel becomes the weights of the NN from that input pixel. Like this: http://fastml.com/images/deep-learning-made-easy/layer2visua...
Standard image compression software should be able to work with that. But this won't work in higher layers, because there are only 3 to 4 color channels in an image.
Dictionary wouldn't work either, unless you think the NN is more likely than random to produce sequences of bits that are exactly the same.
There may be regularities in deep NNs. For example, many learned features are basically exactly the same but rotated or translated in a different way. This is because NN's aren't inherently invariant to rotations or translations, so they have to learn multiple features for every possible variation. Another possible regularity is that the weights tend to cluster on specific sections of the input image and ignore the rest.
But this entirely depends how the weights are projected into an image. The image provided looks like random static. The standard method is to just create a separate image for every feature, and for every pixel becomes the weights of the NN from that input pixel. Like this: http://fastml.com/images/deep-learning-made-easy/layer2visua...
Standard image compression software should be able to work with that. But this won't work in higher layers, because there are only 3 to 4 color channels in an image.
In fact, since the first layer in that example is just gabor filters, you could manually code a program to create gabor filter units for that layer, and get massive compression. Don't know if that's true for all NN's though.
In retrospective, the full connect layer's regularity probably just reflects that some neurons are more dead than others.
You always probably want to have a dictionary-based compression method (LZ-like) for the last step to see how much more you can squeeze. Before that, quantization, residual quantization, some clever transformations probably can carry much longer way in terms of compressing the full connect layer parameters. I haven't explored literally any of these fancy methods yet.
Just finished running xzip / gzip on best setting on my quantized full connect layer output to see if pngcrush's gain is my illusion. For the last two full connect layers, both pngcrush and xzip and gzip produce the about same size files, which suggests very limited compression ratio. On the first full connect layer, pngcrush (8.2M) is marginally better than xzip (8.6M). It is a reasonable choice as last step if you don't want to have library reference to xzip ;)
When I was playing with convolutional neural networks a while ago, I found the following tutorial to be a good introduction to the topic: http://deeplearning.net/tutorial/lenet.html