In pondering this as a real world application/solution, in which we humans describe a photo, save the description as text, and then rely on AI to render the image.... I could see this being like a game of telephone. Some people will describe the photo very well (as I believe was done to generate the middle image in the article), whereas some people will be very non-descript with something blunt like "my cat", which will then be rendered veeery differently.