The question is why an image-generating model needs so much fewer parameters than a text-generating model in order to produce useful results, when our everyday experience teaches us that images need much more storage space than text to convey similar information.
LLMs have been able to generate words for years now. Hell, Tay was back in 2016. Making sure ChatGPT is able to answer the way it does (ie filter out bad things) is part of what makes it hard, and thus bigger, to implement. But having a cohesive readable output is what's hard and takes up more space. Plus, unless I missed it, we don't actually know how many gigabytes the model file for ChatGPT is.
That is because an image is a collection of bit that attempt to represent reality as it is. It probably is easier to relates blue as a "color archetype" when you literally have a collection of bits that mean literal "blue" all the time. In languages "blue" doesn't always mean the color.
Texts are abstractions/coded form of reality. Your mind itself contains the decryption codes to translate text into what it actually means. That decryption code for text apparently is much bigger and harder for a machine to crack than interpreting images.
Also, images are bigger than a text file because of the way data is stored. It probaby has nothing to do with the amount of information stored inside. A book about quantum physics can be smaller than an image of a cat for example.
Conversely, this may underscore how inefficient pixel-like storage is for communicative and artistic images. Ten years from now, many of those kinds of images may only take a few hundred bytes and a good enough generator model to “decompress” them for display.
Text generated though we want not to be just a pile of recognizable words but to follow some pretty strict rules and to actually be true.
Imagine you insisted stablediffusion only produce photorealistic images and judged it for every inaccuracy.
Because a word is worth a thousand pictures.
Sure, an image of standard font/m face/size/weight/color text has low information, because it's only using asymptotically 0% of it's expressive power. Just look at how a logo conveys more information just by styling text.
It's kind of like saying "Why would the factory creating tiny processors for phones (with only tiny bits of raw materials) need to be larger than that other factory that produces loads of big loafs of bread?"