Last thing I've tried was a Distracted Boyfriend meme, but with people dressed like a cat, a cat tree, and a sofa (cardboard cosplay style). Maybe I simply don't know the right way to explain and make it write a perfect prompt, but it was insurmountable task for DALL-E. It always forgot something - either that there should be only three people, or that there's one cat-costumed person, or that a sofa is a costume and not a piece of furniture, or that I wanted to have the obvious canonical layout of the meme and so on. I just gave up.
With Stable Diffusion at least I can do this using iterative inpainting or ControlNet segmentation. DALL-E simply cannot iterate on existing images.
Overall, DALL-E feels like a dumbed-down version Stable Diffusion without repeatability (can't control the generation, so can't reuse the looking-good seed) and a bunch of extra hoops to jump through.