Illusion Diffusion: Optical Illusions Using Stable Diffusion
github.com
github.com
That prompted me to generate ambigrams with stable diffusion. The results looked odd, as ambigrams tend to, but the "text" was largely illegible. I wonder when the state of the art will be able to handle that request.
(Normally you would feed the output of step n right back in as input to step n+1. That’s what is not happening as usual here.)
This paper [1] shows that giving character-level awareness to the model can improve the "visual spelling".
Here's where it sucked. It seems to have learned the superficial aesthetic of an optical illusion or of "Escher" without learning the relevant component. It spits out things that either aren't optical illusions, or are just random disconnected spattering of geometrical inconsistencies without any overarching theme. A person made optical illusion will generally have a single main loop of impossibly connected objects, or at least some simple overall topology. The illusion is expected to exist on the global scale of the image, not as a weird pocket of a mostly normal image.