The uncanny failures of A.I.-generated hands
newyorker.com
newyorker.com
Hands. Chopsticks. Teeth. Sport balls. Etc.
Kinda feels like it gets into loops, but it can't count so it can't tell when to stop.
The invariants that apply top down structures from the low frequency domain aren't as readily recognised or understood.
Classic chatgpt: "The word contains R because it's spelled L-E-M-O-N"
The answer "gnocchi" is derived from the clue "spouted pan I used for dumplings" using a type of wordplay known as anagram. An anagram is a type of wordplay where the letters of one word or phrase are rearranged to form another word or phrase.
In this case, the clue contains an anagram indicator, "spouted," which suggests that the letters of some words need to be rearranged. The words "pan I used" can be rearranged to form the word "nog," and "for dumplings" suggests that we need to add another letter to complete the solution. The letter "c" is a suitable choice here, as it is a common letter in the word "dumplings," and it fits the remaining letters to form the word "gnocchi," which is a type of Italian dumpling made from potatoes.
Therefore, "spouted pan I used for dumplings" is a cryptic crossword clue that requires you to rearrange the letters of "pan I used" and add a letter "c" to get the answer "gnocchi."
This will be true for most AI "failures". People love to harp on some minor issue some new generative AI has in edge cases and act like it's some insurmountable obstacle that will keep humans superior forever: "haha look at the AI try to make a hand what an idiot"
As humans become more and more irrelevant, we will cling to these minor failures more and more to preserve our egos
The truth is we don’t know how far today’s breakthroughs in AI will take us. There’s quite a lot of headroom to explore, but history has shown that there are always some brick walls as well.
The rich discussions can be good, but the cheap flamewars are exhausting.
That's the central question. So far, AI has proven great at ingesting huge amounts of "prior art" and producing variations on the themes it has learnt (more or less) according to the prompts you give it, but it consistently fails to grasp the underlying rules (gravity, perspective, anatomy etc.) that are supposed to govern the images it generates. So it's better for "surrealist" art than for images that are supposed to be realistic...
"Psychic automatism in its pure state, by which one proposes to express—verbally, by means of the written word, or in any other manner—the actual functioning of thought. Dictated by thought, in the absence of any control exercised by reason, exempt from any aesthetic or moral concern."
Everything we do, ultimately, comes down to recognition.
The inflection point will be the first model that is trained to recognize trained models.
Saying that intelligent machines are impossible because of hand-wavy metaphysical reasons (or vanity as you claim) is completely different from saying that this technique can or can't be improved to do things that it can't now.
Jet turbines definitely can and have been improved to exceed the capabilities of birds, and it surely doesn't violate the laws of physics to build a machine that flies by flapping wings, but one technology doesn't inherently lead to the other.
It's not a choice of either unlimited improvement or an insurmountable obstacle if improvements are going in a different direction that what's being considered.
I think there is some deep psychology at play that makes us want to believe we’re about to solve all the things we’re insecure about.
Being good at art, being good at speaking or writing, handling the idea of mortality, being great at maths.
We’ve decided that we want machines that can outdraw everyone, in all cases, with zero limitations, and we might be, but we’re also happy to overlook some of the floors and setbacks of such systems to believe there we’ve solved art.
As someone who’s played with these tools and as someone who has done a lot of illustrating in my life, as a hobby, I can still tell you that there ares setbacks, which are, these tools don’t read minds…I still don’t get to produce what’s in my minds eye. I can come close but it is a different product that’s produced. The pros are they do make some great art.
TL;DR nothing is perfect…
I'm just extrapolating that kind of wishful thinking to anyone who thinks that AI will just be a tool for limited uses, always inferior to the less naïve human whose multi-modal efficiency could never be matched.
There's nothing against the laws of physics that prevent AI from becoming superhuman in all domains. There are people who nevertheless claim AGI is impossible because of handwavy reasons.
And it told me firmly "yes" followed by fiction, with a plausible link that went nowhere. Honestly, I'm not entirely sure it was complete fiction, it was so plausible and what I wanted to hear. But I couldn't find it with Google.
What's bothering me is not just people saying this is like a human being, but people saying the next generation, the progression will bring us closer to a human mind. "Better", if it is more deceptive, will be worse.
Improving what it is already good at will produce something but not a human-like mind and not a superhuman oracle.
Humans have eyeballs, but an artificial eye is never going to be a human-like mind no matter how much you develop it and whether it is better or worse than a human eye.
2) Why the need for a relatively pejorative term, "potheads"?
Another which I discovered myself as a kid as a first sign was that if I go back and reread the same text it never stayed the same. Discovered other lucid dreaming signs later. The brain continuously generates/halluicnates stuff while dreaming, keeps too little context.
I'm going to say... no, this isn't true.
I wish I could remember to do this :-O
Or looking twice in a clock and seeing if the mind is making up the position of the clock hands on the spot.
Or looking text in a book and seeing if it's gibberish
In computer graphics rendering hands and hands gestures was considered very challenging for a long time (This was, for example, a PhD thesis of one of the Pixar co-founders, back when computer graphics was just starting). And it is still very challenging to render a close-up of a hand performing a natural looking gesture.
https://www.youtube.com/watch?v=ptEZQrKgHAg
Now from personal experience, it's not as easy as it seems in the video. It still takes a lot of tweaking and many iterations don't come out well, but hands are definitely not an insurmountable problem.
that's actually true for all the art generated by (the current generation of) AI. With some things (like hands) it's more obvious than with others, but the closer you look, the more "wrong" things you notice...
They definitely do and with pretty decent (albeit certainly far from perfect) results. I feel you have a bit of cognitive bias based on your sources. A few hundred images of hands well tagged used as fine tuning into an existing SD model produces a dramatic improvement, based on personal experience, including grasping objects. Are they photograph level reproductions, certainly not, but easily qualify as skilled artistic renderings. The part I’ve been experimenting lately on is getting fingernails and skin texture to both be realistic at the same time (I can get one or other, but not both so far).
If you really want to know, you could just try it yourself. Setup WebUI, download some CAI or HF SD models that have been fine tuned with hands, learn about prompting, and then see that the field has dramatically improved by just not using stock settings.
Zero to hand model hero in around an hour!
links2 https://www.newyorker.com/culture/rabbit-holes/the-uncanny-failures-of-ai-generated-hands
(in this case i was more interested in not getting paywalled)i was still able to load the images by typing 'i' to launch an external viewer
It seems like it’s caused by the massive number of permutations possible with hands: all the joints, combined with all the angles at which they can be seen.
I’ve wondered if this could be solved by creating a massive permutation of hand images using rigged 3D hand models (using, for example, Blender, or coupling with Unreal), and programmatically putting them in all possible combinations and angles possible from the rigging and then rendering millions of images of these combinations.
Then the image models could learn from that artificially created dataset.
Anyone with actual knowledge know how people are trying to tackle this?
1. Humans are terrible at drawing hands and feet (ask any illustrator, amateur or professional).
2. Advanced Interpolation software reads and generates images based off of both real-life photos and illustrations of hands drawn by humans, depending on what materials are fed to it.
3. The images generated carry forth the failings of humanity.
Here's what you missed. The generation problems all happen under a certain length scale. Its things under a certain band of fine detail where the distortion happens. Normally you won't notice it, because normally there isn't anything but noise in that band anyway. Shove a SD generated image through a fourier transform to see what I mean.
And here's my very informed conjecture why it happens. The hint is right there in the name. Stable Diffusion. The generator network trains to de-noise images. The adversary adds Gaussian noise. That's a perfectly reasonable noise method, but it comes with its own spread and distribution in frequency space.
A very similar problem exists with dithering processes. One way to use a b/w screen for grayscale is to treat the gray scale values as coin flip probabilities. This is known as random dithering. Its simple, its obvious, it should work, and it does work surprisingly well most of the time. But it runs into the exact same distribution problem. When there are closely spaced stripes in the image you want to show, that stripe pattern gets completely washed out by the noise being added to fake brightness levels. Put another way, overlaying TV static on any image makes it impossible to see blobs of the same size as might happen by random chance. The generator network can't see certain patterns, because the length scale of those patterns are coincident with the length scale of the noise being added by the adversary.
This problem will never be solved by substituting data for comprehension. That's a one shot lesson worth generalizing.
90% solved it appears. Where the baseline I'd say has been "almost impossible to get right."
I don't know how well if at all the adversary side of the process adjusts the probabilities.
I still don't understand the problem, if you ask model trained on a noise pattern "trees" for a forest it will still give you a random forest, that's what it was trained on, also: https://arxiv.org/abs/2208.09392, to see the diffusion process applied to processes other than Gaussian noise.
https://surma.dev/things/ditherpunk/
Imagine asking Dakke2 for a picture of the bridge through the fog. The problem is an image of fog is graphically indistinguishable from random noise. Whats the difference between a fog patch and something the adversary drew in? Good luck training out of that one.
Less tortured understanding: think about my playing cards example. Consider a face down card with the ordinary geometric lacing patterns drawn on the back side. So we are all looking at the same thing: https://www.wopc.co.uk/images/countries/usa/standard/standar...
Now think about what happens to that red card as I randomly add white noise static on top. The original fine patterning on the red side is soon distorted beyond recognition because those thin white lines are obscured by equiprobable white dots and coincidental random patterns.
I don't see the problem, the pattern will remove the static white noise at the top and generate a new fine patterning on the red side (in case it did not go through the forward process to the end), of course it will not recover the lost information.
>Imagine asking Dakke2 for a picture of the bridge through the fog. The problem is an image of fog is graphically indistinguishable from random noise. Whats the difference between a fog patch and something the adversary drew in? Good luck training out of that one.
Ignoring the fact that Gaussian noise is really different from a patch of fog, it is obvious that if you have a high noise-to-signal ratio you will not be able to fully recover the signal, during the forward diffusion process all information is lost, however, the model will learn the distribution of samples. Also Dalle 2*
You said it yourself, it erases information. Pivotally it isn't just a blunt eraser. It erases a very particular grain of detail. The problem is it looks like your on hallucinogenics. https://i.imgur.com/YlPHzgY.png
My hypothesis here is that the distortion manifests when you have fine grain patterns like you would find on a playing card because that is the length scale that is most affected by the noise process. The presence of noise doesn't effect long range objects. The same process would never cause the red card to conmpletely go away. It destroys information about that fine, thin, white line and patterns of a similar characteristic. That's really the key to what I'm saying.
The forward diffusion process completely destroy all information, in fact when the model is trained we simply sample directly from the gaussian distribution instead of actually applying the forward diffusion process on a sample, so the fact that you lose information in the diffusion process is not really a problem because with the diffusion process you lose all information but the model clearly can still generate images.
Its exactly the same sort of problem as in the dithering example. An original image of 2-pixel-thin stripes is very recognizable to us as a "thin striped object". Trying to color it in with random dithering will render it unrecognizable. However a much wider 10-pixel-stripe pattern can be safely colored with random dithering.
The commonality is long range ordering of short range features. Thats what the SD process struggles with.
"when increasing the image size, the optimal noise scheduling shifts towards a noisier one (due to increased redundancy in pixels)"
That's highly related to what I'm saying. Its a statement that may be true in the vast majority of cases, and thus would certainly be validated as true-on-average by whatever benchmark, but there's certain types of patterns that aren't redundant among nearby pixels despite being macroscopically visible.
A similar thing happens with compression artifacts. For pictures of normal environments and objects jpeg works great. Save a 1 pixel wide line as a jpeg and it is surrounded by artifacts. jpeg compression is also largely built on the assumption that high frequency components in large images won't contain anything important to the image content.
In both cases, the result is a visible digital artifact with a "tripy" patterning effect.
> These variations in the coarseness of noise pattern will each obscure patterns of the same coarseness.
is not accurate. During the forward noise process the high frequency information is obscured first, and the coarse information last. During the reverse process, the neural net learns to denoise the coarse visual elements first, and adds high frequency details at the end. With the original clip guided diffusion notebook you can skip the last 10% of the denoising steps to get a smoother image.
This is also the root cause of the noise offset problem (the coarse information is not completely obscured by the forward process)
I'm glad someone else sees the frequency space folly of the diffusion process. If I had time, I would test the hypothesis of doing all the learning in frequency space rather than trying to shape the profile of the noise. But I don't have time so feel free to steal my idea.
I don't think "people" actually tackle this. Ask a kid or non-professional to draw animal hands/feet without learning them first, e.g. crocodile hand, or feet of a seal. "People" ends up with uncanny drawings as well.
There are tons of face photo on the Internet, but few of them are with hands clear visible.
Also, if there is text, it will never read the same twice if you look away from it.
SD still has the hand problem, but a lot of the newer checkpoints are getting pretty good at hands.
As far as keeping up with things, the best way to do this is actually to be in AI related discords. I'm not sure why, but people involved in AI are terrible about having a single source of info...or even just info in general. I would start by joining the LAION discord then finding others from there.
Another great way to learn about what's going on in AI is actually through threads on 4chan's /g/ board. If you're willing to look past the fact that it's 4chan you will be the first to know about lots of neat stuff.
Btw, I'm not a researcher. I'm just talking about ways to get a quick scoop on what the current AI meta is and how far it can be pushed.
edit: I just realized you don't care to keep up...lol. Oops. I'm still leaving this up because I think some people will find it useful.
I think the overarching story (subtitle of the article) is what is the message: "machines can grasp small patterns but not the unifying whole."
Amateur human illustrators might also struggle to draw accurate representations of hands, but, unlike machines, can critically evaluate and KNOW whether they got it right or not...
but yeah humans rarely draw 7-fingered hands or smiles with four rows of 40 teeth
Guess you haven’t seen hand drawings done by young children?
thanks
That's because yourself as a human being already owns a pair of hand, you can look at it everyday.
Now try draw the hands/feet of some animal you've heard of without a photo. e.g. hands of a sloth
I view this result as further evidence that these algorithms are approaching the target.