Asymptotically, yes; prediction = compression (if you have a model for a bitstream which produces probabilities over the next bit, it can be fed into an arithmetic encoder and you now have a compressor). In this case, it's not practically helpful. A VGG is 528MB all on its own, so you need to compress a lot of images to make back that 0.5GB use plus runtime dependencies.