Could you imagine if the mediocre results we currently get from "AI" were mostly from poorly labeled data in every huge dataset and not lack of a scientific or technological breakthrough?, it wouldn't be the first time that too many academics are blinded by the wishful thinking that what they need is an eureka moment when what is needed is tons of dull repetitive work.
So for instance for photography generation what may be needed is huge amount of clean photos with obsessively detailed labels, maybe just the exact same single-point-lighting (the exact coordinates of the light being a data point, plus strength/lumens), with 8 pictures of each subject in black background: front, back, left, right, top, bottom, 3/4 mostly-front (AKA the corner), 3/4 mostly-back, and then the same 8 ones but with white background, then also include all the info possible: weight, height, width and depth, plus versions of the most common states of each object (ball: inflated or deflated; bird: flying, idle or walking), plus photos adding two subjects together (one dataset of woman with hat, another wearing the same clothes but without the hat, and one of just the hat without the woman), with properly labeled relationships so it's clear they all refer to the exact same hat (or lack of), you get the idea...
Then the most important thing for "image generation AI" may not be computing power but the most boredom-resiliant staff you can hire; of course that's just an example for photography, for things like text you may need an equivalent rigorous effort by a multitude of linguists.