Shared Attention Layer Issues in ChatGPT and DALL-E
Both ChatGPT and DALL-E and pretty much any current text AI models internally use an embedding space where the meaning of text is encoded. And then they decode that embedding with attention layers. The "attention" here is like a fuzzy key-value lookup to determine which words go together.
And both ChatGPT and DALL-E make the same mistakes here. Mistakes that no human would make.
If you ask ChatGPT to "write a funny story about the medieval princess miyu" it'll actually do "write a story about the funny medieval princess miyu". If you ask DALL-E to generate a "black and white photograph of a green orange in the jungle" it'll generate pictures of a "green orange on top of black and white jungle". What that shows is that the attention layer cannot distinguish between "funny" as an attribute of the story vs. "funny" as an attribute of the person inside the story. Or similarly, between "black and white" as an attribute of the image vs. as an attribute of the jungle.
Both AIs are lacking the ability to distinguish between multiple objects. My guess would be because the entire text is temporarily merged into one embedding space.
In line with that, both ChatGPT and DALL-E also cannot count. DALL-E "photo of twelve oranges" will almost never show 12 oranges. And ChatGPT "count the number of A in "BIKLDJJ34c2424v2bv43v2c3xwbves4b5bLSKJ"" will return random numbers if you ask it multiple times.