Maybe a limit of training these systems on internet text and images, but it seems to me that we’d need something trained on the human experience itself to get expression that doesn’t read hopelessly derivative. Now and likely for a long time, all these tools need a puppet master.
And, of course, celebrity culture requires a celebrity in the flesh. At which point an AI analogue would be so sufficiently similar to us that they might be afforded a SAG membership themselves.