*shrug
Lots of different people on this forum, no offense was intended, and like all nerds I get excited when I get to share knowledge.
>How do you read my question and interpret simply it as "lets send all of the images in full and then give their index and call it compression?"??
Because it struck me as analogous, and yeah I'd call that compression - the message length for one is immensely improved.
Okay, so here's my admittedly piss poor understanding of most compression: you either find more intelligent ways to strip bits from the source in ways the consumer won't mind, or you find more intelligent ways to build reference tables given your problem domain.
I'm sure you know this, but the first one is why jpeg/various mpegs are successful: they have complex quantization models that eliminate gradients we won't notice or frequencies we can't hear.
The way you achieve better results is through building better models for how information in your problem domain is related. If we're compressing text and we know the language we can start referencing letter frequencies and index along that and so on.
The way this works in video to my knowledge is, amongst many other complicated things, they take NxN blocks of images and store only the deltas between Y numbers of frames.
So - perhaps you "image reference blob" could build a reference table for all 16x16 px blocks and transmit only the indices for them, and we're back at my original comment. But to my (again, please correct me) understanding those are kind of the only alternatives? Encoding and decoding is an interesting topic.
I very infrequently have to think in binary and I had perhaps too much fun counting to 10,000 on my fingers; my undergraduate was a long time ago and my knowledge on the topic sparse, so I'm interested in hearing more about it.