Visual Anagrams: Generating optical illusions with diffusion models
dangeng.github.io
dangeng.github.io
I wonder how many permutations could legibly be generated in a single image with an extended version of the same technique. I don't understand the math, but would two orthogonal transformations in sequence still be an orthogonal transformation and thus work?
Here a cat is made from 9 paintings of cats in the style of popular painters:
https://twitter.com/marekgibney/status/1521500594577584141
You might have to squint your eyes to see it.
I made a few of them and then somehow lost interest.
The problem would be this: In the picture at hand, the big cat is rather simple. Just a portrait of a smiling cat. While the 9 smaller cats are doing all kinds of poses to adjust to the form of the big cat portrait. So the subcats are more complex than the main cat.
When doing the recursive cat, it would be hard to make a subcat from 9 subsubcats because the subcat is already a complex image that is not as easy to recognise as the main cat.
Now what would be interesting is a "demixer" which allows you to locate the source image(s) from multiple interations of a given image. Like a reverse image search but for generative images. I suppose it would rely on artefact matching or some other kind of granular pattern matching, along with other more general methods (assuming the source material is actually available online in the first place).
https://dangeng.github.io/visual_anagrams/static/videos/grid...
The rotations on the other hand, wow! It is perfectly visible how the pixels don't change. You can physically rotate the screen and the image "changes". I could not think of a better illustration of how diffusion model images are not just echoes of preexisting images (they certainly are), but solutions to the problem of "find a set of pixels that will match the description of {prompt}". Or in this case, "that will match {A} when oriented this way and {B} when oriented that way".
Per the code, the technique is based off of DeepFloyd-IF, which is not as easy to run as a Stable Diffusion variant.
Is there a good repository anywhere or is it just "wade through twitter"?
https://www.latent.space/p/sep-2023
https://github.com/swyxio/ai-notes/blob/main/Monthly%20Notes...
The ecosystem around Stable Diffusion in general is so massive.
I haven't dug in yet, but it _should_ be possible to use their ideas in other diffusion networks? It may be a non-trivial change to the code provided though. Happy to be corrected of course.
> Our method uses DeepFloyd IF, a pixel-based diffusion model. We do not use Stable Diffusion because latent diffusion models cause artifacts in illusions (see our paper for more details).
You can always reach out and ask for a one-off in good faith.
the penguin/giraffe is probably the best one. The old lady/dress barely looks like either.
* very closely https://www.pinterest.com/pin/giraffepenguin--13398215764267...
* or directly inspired by, but the “young lady” prompt triggered the model to pick a dress, and there’s no way to make an eye and an ear or a month and a chocker photo-realistically identical: https://www.reddit.com/r/RedditDayOf/comments/35cjn5/the_cla...
(I just went and looked on my phone and they are all more effective when they are tiny squares instead of blown up on a big screen.)
If you think of the edges as nodes in a connected di-graph of ears and holes, possible pairs are connected: a swap is a two-pair cluster; further connection is a four-element chain with both ends open-ended. If that connection ties to more pairs, you might have a larger cluster of identical hears and holes. Given graph properties, that’s presumably most of them — see the prisoners paradox for why [0].
That would make the puzzle much more challenging to solve if most ears fit in most holes.
[0] The excellent Matt Parker https://www.youtube.com/watch?v=a1DUUnhk3uE but I recommend the following debate with Derek from Veritasium.
They also “copy” the way those networks seem to do so often that they somehow get copyright strikes; they were either prompted on existing solutions or learned them whole through training:
* The penguin and giraffe one is a previously known ambigram, for example.
* The old lady turning into a dress is obviously based on a classic pencil drawing where a similar old lady hiding in her collar turns into a young lady looking behind her shoulder [0]; however, the network interpreted “young lady” and turned into a white dress because color-matching the two different body parts from the pencil outline and turning it photorealistic wouldn’t have been much harder otherwise. There are photorealistic interpretations, though [1].
I’m more impressed by the radically new ones, like the fire flipping into a face—but most of those rely on having two distinct parts of the image be meaningful in their own context, and not relevant otherwise.
The black-and-white inversion man/woman is impressive because the two interpretations are not on separate parts of the image. That’s where you can interpret the quality of the effect as the model having learned how humans perceive and pay attention to dark and light contrasts differently. That one captures an understanding of perception.
[0] https://www.reddit.com/r/RedditDayOf/comments/35cjn5/the_cla...
[1] https://www.jagranjosh.com/general-knowledge/optical-illusio...
These things aren't a mystery. There are principles you can work from to produce such multi-stable illusions in formulaic, computer generated ways without resorting to the technical debt of a neural net. But, as with so much in modern times, training a neural net gets results faster than distilling a true understanding and then translating your understanding into code.
[2] https://en.wikipedia.org/wiki/Multistable_perception
[3] https://www.isc.meiji.ac.jp/~kokichis/triplyambiguousobjects...
[4] https://i.pinimg.com/originals/6d/07/bd/6d07bd12d34f674ca4de...
Looking at fine elements like hairs (nevermind curly hair) is a disaster, especially when you're used to fine classic german/japanese optics that accurately reproduce every subtle detail of a subject while having extremely aesthetically pleasing sharpness falloff/bokeh.
I've had to swallow the pill though: No one (end users; pros are another story) cares about those details. People just want something that vaguely looks good in the immediate moment, and then it's on to the next thing.
I suspect it'll remain the same for AI generated visuals; a sharp eye will always be able to tell, but it won't really matter for consumption by the masses (where the money is).
https://news.ycombinator.com/item?id=15539373
https://en.wikipedia.org/wiki/Boustrophedon
https://news.ycombinator.com/item?id=15547162
DonHopkins on Oct 25, 2017 | prev | next [–]
Scott Kim has a wonderful talent at designing "ambigrams". Check out his classic book "Inversions" and his gallery of more recent work!
http://www.scottkim.com.previewc40.carrierzone.com/inversion...
An inversion is a word or name written so it reads in more than one way. For instance, the word Inversions above is my name upside down. Douglas Hofstadter coined ambigram as the generic word for inversions. I drew my first inversion in 1975 in an art class, wrote a book called Inversions in 1981, and am now doing animated inversions.
A Scott Kim Ambigram for "George Hart":
https://www.georgehart.com/scott-kim.html
John Maeda's Blog: Scott Kim’s Ambigrams
https://maeda.pm/2017/12/17/scott-kims-ambigrams/
The Inversions of Scott Kim:
https://www.anopticalillusion.com/2012/04/the-inversions-of-...
Channel: An Optical Illusion » scott kim:
https://optical397.rssing.com/chan-26600952/index-latest.php
Scott Kim’s symmetrical alphabet:
https://stancarey.wordpress.com/2012/10/18/scott-kims-symmet...
Typography Two Ways: Calligraphy With a Twist
That's sad, I would've loved to try it.
(Back in Disco Diffusion days I was happy to spend money on Colab Pro. It was fun)
Aren't you sad they don't just let you shoplift it for free?
I'm getting the impression you're just an entitled gamer who wants a free ride from the University of Michigan, not a professional programmer or AI developer who would actually get some tangible value out of subscribing to ChatGPT for $20 a month. I'm thankful to be alive in a time I can so conveniently get so much value for so little cash.
Is that the case? Is $10 really too much to ask to use a high-end GPU for a month? Then it's not really as sad and hopeless as you complain it is. Just be a good boy all year, ask Santa for an GeForce RTX 4090 for Christmas, leave some cookies and milk out for him, and hope your parents get the hint!
people don't pay for things without getting a feel for what they're getting. hence the huge focus in saas on various monetisation strategies. if someone puts these anagrams in a product, it will be freemium or have a tree tier, and then i will play with it.
there are 20 new projects like this every day, i'm not going to pay for all of them just to try them. i'll try the product if/when there is one
You are aware that the University of Michigan is not a startup whose mission is to make an SaaS with a free tier for you to play with funded by their investors, right? Maybe if you enrolled as a grad student they'd let you use their resources for free (once your tuition check cleared). But your chances of being accepted into their PhD program would be higher if you showed more than $10/month in enthusiasm and initiative.
Congratulations!