Forget JPEG, How Would a Person Compress a Picture?
spectrum.ieee.org
spectrum.ieee.org
It only works because the model allready knows what a 'giraffe' is. It's pre-stored. But if one had to describe a giraffe, one needs to rely on other things which we take for granted that humans know (like what hair and eyes look like).
By that standard I can say that a picture of albert einstein only takes like 30 bytes if i assume the recipient knows what he looked like.
That is a very large algorithm, but technically could be pre-shared knowledge. (And there are lots of ways to improve over keeping a copy of all photos on the internet.)