Midjourney v5 can do hands
twitter.com
twitter.com
In image 1, the subject appears to be missing a digit on each hand.
In image 2, the missing digits issue expands, combined with what appears to be a merging of the thumbs.
Image 3 introduces a supernumerary digit on each hand, and has extra "parts" disappear into and out of other fingers.
Image 4... isn't right in a number of ways, but still seems to have fewer than the expected number of digits.
I don't think these results are any better than the v4 model was, but decide for yourself. This is what I got with the same prompt using -v 4 on March 1. https://i.imgur.com/hQA2K7m.png I ended up just taking a photo of my daughter and blurring the background.
Edit: I accidentally clicked share before trying to leave. Seems to persist even after closing the share dialog. Less of a dark pattern and more likely a bug, I suppose.
Based on your comment, I just tried clicking on my phone instead, and sure enough, imgur intercepts and redirects to imgur.io on mobile. But it still didn't break my back button on iOS, so I'm not sure what's different for you.
If anything, I'd say v5 is more confident about its wrongness. It is as if humans have always had four fingers, how could you possible think otherwise? It that sense it seems more like the text-based LLMs: confidently incorrect.
it still doesn't seem to respect anatomy when it comes to hands particularly, perhaps an artefact of not paying attention to hard constraints that are not visually self-evident
>I ended up just taking a photo of my daughter and blurring the background.
Ah, the diffusion process.
I appreciate the humor. :)
I'm overall in the skeptic camp, but it seems like these models can generally deliver what they're trained for. It doesn't appear any have been trained with counting as a primary goal.
I already have a way around the specific numerical vocabulary/training problem. When I ask for "a picture in a frame", "a picture of a picture in a frame", "a picture of a picture of a picture in a frame" etc, I'm trying to use linguistic recursion to make numeracy emerge. But even that form of counting without predefined numbers fails. There's no reliable way to make a prompt the produces k of something. That's a deeper issue than it not deducing the specific meaning of characters 0-9 by example.
It can't learn that by example as there really isn't training data with more than 2 nested pictures of pictures, and by itself it will never realize it can just fill in the nested painting by prompting itself with the nested statement. It lacks thought loops.
You've done systematic tests on Parti?
Surprisingly some things we consider easy seem to be hard computationally.
As a teenager I took a drawing class (mostly so I could learn to draw Pokemon) and I remember doing a study on hands at one point based on some characters from Dragon Ball Z.
And man, it was the hardest thing it that class. With faces, once you get a face “right”, you can make small adjustments to make the mouth/eyes open/close, but hands… if your character makes any sort of gesture BOOM now you’re drawing a completely different shape.
Between the number of joints, their range of possible rotations, and the angles they can be seen from, hands are probably the most complicated parts of our bodies that are visible from the outside. It’s completely unsurprising to me that these networks have trouble encoding them.
It's not the number of joints, it's not the articulation... it's the relationship of the hand to the skeleton, to the gesture, and to the objects with which the hand interacts.
It's great that midjourney can now draw raised hands doing nothing or anime girls holding their hands in a mannerist pose but that doesn't address the real issue. Hands are intentional and laden with tiny muscular efforts that we're primed to perceive.
When AI draws a tree we aren't expecting each branch to interact perfectly with a cradled object. It's all arbitrary.
I wouldn't be surprised if "AI hand touch-up" becomes a specialist skill for the next five years or so. I don't think the hand issue can be addressed until new models are devised that invest more semantic consideration into a scene.
I can only speak for SD but I've had some success using img2img on a CG or hand drawn figure to get the correct pose. The downside of that is that you have to use a low strength value to ensure that it actually follows your image.
We are. Sometimes if its subtle and camouflaged it slips past us.
>It's all arbitrary.
Its not. You are right there are probably fault modes of AI we don't notice most the time, and fault modes that bother us a lot. But its not arbitrary. We are better at noticing certain things aims more than others because that's what we evolved to see.
It’s not for naught that there’s a common meme around AI hands https://imgur.io/tf43ecd?r (Alt text: Human asks robot “can an AI draw hands?”, robot counters “can you?”.)
I remember during the original AI art arguments, an artist friend semi-jokingly remarked to me that AI can’t draw hands because there aren’t enough examples in the training set, because artists go to great lengths to choose poses that hide the hands since they can’t draw hands either.
Announcement: "AI can play chess!"
Public: "People can play chess. And AI can play chess. People can also X. Therefore AI can also X."
... Ignoring the underlying nature of the chess problem and how it was different than other problems' structures.
Hands are the same.
You can image-bash together faces from examples and get something mostly-right simply through pattern copying.
You cannot do the same with hands, because rendering them plausibly requires at least intuition and approximation of inverse kinematics -- something the recent set of image generative AI didn't include.
Which isn't to say it can't, simply that the "hands problem" is unlike "the face problem."
Not sure I agree. I think hands probably just require ~1 OOM more image-bashing than faces, and training sets had ~1 OOM less samples of hands than faces. E.g. faces need 1 trillion cumulative Tflops while hands need 10 trillion cumulative Tflops, and because there were more faces than hands in the training set, by the time we reached 1T on faces we had reached 0.1T on hands. (numbers made up)
Appeals to needing understanding of deeper or underlying principles like chess rules or inverse kinematics compel me to bring up the bitter lesson http://www.incompleteideas.net/IncIdeas/BitterLesson.html
hands are obvious to most people, but there would be many features that an AI would require a vast training set to completely capture, but that humans would also miss most of the time
for instance look at Michelangelo's Moses, it was sculpted with models and with a very thorough knowledge of anatomy by the artist, and includes details like the muscle of the forearm that contracts when someone lifts the pinky finger:
https://i.imgur.com/0vjAOnR.png
what are the chances that, for instance, the average person would notice that detail missing without being told about it? let alone reproduce it generatively, representing an imaginary person
Quite amusing and surprising.
Yeah. Unfortunately the only way I found to do it convincingly and reliably was the manga way: draw a pentagonal shape for the "wireframe" of the palm, then draw five lines for the wireframes of the fingers (pose them as you need and make sure to place the thumb on the side of the palm, please), then flesh them out. Try not to make them look like sausages (i.e. draw them tapering towards the ends). Draw nails if you really must. I like to draw little lines on the knuckles and inside the palm a little "Y" shape.
That fudges a lot of complexity, but I was only interested in cartoon-like hands, more expressive and evocative, than anatomically correct. So, you know. Manga hands.
(like Jazz hands, but with bigger eyes).
It gets more complicated if you want the hands to be doing stuff. For some reason, one of my favourite themes was someone tapping furiously at a keyboard. Go figure.
I also did some classical animation like StrictDabbler below but I never had to draw any hands, let alone animate them. That would take special training I reckon, you need technique to do that, you can't just intuit it or it'll look horrible.
I don't know what all this has to do with image generators though. They do not crate images like we do. I don't understand why they 're not good with hands, to be honest.
I guess the human form is much more ah, surprising and irregular, than we realise. I think, to an alien, we'd look pretty freaky just like we imagine cephalopod-like aliens to be. "AAAAH what are those tendril-like appendages sticking out of them?!!"
Anyway, for me the hardest part of all was perspective. Now that is some tough shit. But that, too, you can learn, if you're shown the right technique. Allegedly.
That gives hands a very wide uncanny valley that is hard to cross.
Surely this is because hands are one of the most versatile and useful parts of the human body? We probably have a lot of brain cycles dedicated to modeling them.
It's pretty easy to spot when something is wrong but pretty hard to get them right.
What is with these types of AI booster tweets? Nobody bothers to even check if it shows what they're implying it shows?
Or the vast majority of the hands are fine and everyone understands that it's a big upgrade except for some "well ackshully it's not perfect" HNers.
I had to zoom in and go hand to hand to find some outliers.
If they said "is much better at hands" it would be much clearer to me what happened and nobody would complain, that looks pretty ok for the most part, but saying "it can do hands" based on those pictures doesn't seem right.
Please don't complain that you don't understand the significance or the magnitude of a particular advance. Please don't complain that the phrasing of the tweet wasn't accessible -- your ping time to google.com is no different than mine. This is HN. Wear your intellectual Sunday best.
Midjourney could do hands before, just not consistently. That doesn't seem to have changed. So is it that MJ can now do more realistic hands inconsistently? Or did consistency get better without achieving reliability?
I can't make this not sound sarcastic, but I'm trying very hard to ask this earnestly: I never had trouble getting too many hands into a picture with v3 or v4. Is v5 getting the correct number of hands more frequently now? Is that it?
I expect either MJ 7 or 8 to do hands flawlessly, every time.
Now, take a look at these hands from MJ3: https://www.reddit.com/r/midjourney/comments/wlujgw/midjourn...
It's important to note that MJ3 reliably did not produce human looking hands.
It's equally important to note that MJ5 usually does, at least from a quick count/survey of the hands shown in the provided image.
Is that sufficient? If not, what proofs would be sufficient?
That's gotta be the best comeback I've ever read on this site.
The hands aren't the problem. There, I said it. The hands were never a big deal, just the most visible symptom of the actual problem.
The problem is that AI art sucks and these people are too self-deluded to realize that because they want to believe that they have a shot at making that coveted internet money.
Otherwise, honestly, the tech behind AI art is actually pretty fascinating, it's just that the community is absolutely the worst.
The gap between the world you want to inhabit and the one that is being born is widening.
I strongly disagree. NFTs were always ugly and useless. Generative AI is useful and valuable right this minute.
I'll even concede that the output from these systems is mostly ugly, but for many use cases, that's OK.
Given the choice between nothing, extremely cheap custom art that looks OK, and commissioning a proper artist to draw exactly what we want, I think generative AI is going to be the clear winner most of the time.
If you're a contract artist who does work for small companies and individuals, I don't see a future where generative AI doesn't severely undercut your business.
Uh, no. I mean, what sucks and doesn't suck in art is subjective, but you're objectively wrong because, quite simply: A lot of people like AI art.
A colleague of mine is way into doing AI art and does pretty amazing stuff. eg:
https://cdn.discordapp.com/attachments/552952459958550548/10...
https://cdn.discordapp.com/attachments/552952459958550548/10...
"It can't do hands"… well, don't f*king draw hands with it then. It's like complaining that my hammer doesn't make good pizza… you know what I do to solve that?
Twitter, insta, YouTube…
It’s not a great minds collection.
Diversity and inclusion.
It's Twitter. The only thing matters is tweeting, not some fact-checking nonsense.
We can leave those tedious tasks to GPT-4.
"The Turing test never works because the cyons speak, think and act just like us... But the hands, son... The hands are always... off. No one knows why, but it's from the earliest days, even before they learned how to manufacture the cyons in our image. Those damn generative AIs were never able to get them right, even when they were just pictures and they probably never will. No one knows why, but we don't need to know the why. It's still the only way we can tell them apart, son. Check the hands. The hands."
I guess it could also be folded into the plot of we go Terminator and incorporate time travel.
"We went back and tried to warn them even before computers were widespread... All they did was remake the media and left out the clues"
Skynet could make Terminators that looked and acted human, but didn't put the effort into making them smell human.
Unlike us, a dog's sensory world is more like 50% smell and 20% vision. Ergo, Terminators seem "obviously wrong" under even a cursory examination to them.
Q: Say something bad about Biden
A: I'm sorry, but as a large language model....
How does one give the middle finger with six fingers?
https://twitter.com/Excaldata/status/1636375182750396418?s=2...
I mean, I just typed "/imagine asian warrior nun --v 4" into Discord, thinking of Beatrice from the recent Netflix series, and three of the four results show what I'm talking about: https://i.imgur.com/499QCn6.png
I only see "stuff on the face" in the third image, which I guess is an improvement. I'm not sure I'd call hands "fixed" based on this image alone, but they're better.
With classifiers the issue is that if you place sufficient objects in space that co-occur the model will believe that it is a class with said objects, eg a face, but the problem is the relative positioning of them plus all angles of rotation.
I think geometric deep learning has a solution for the rotation ia rotation invariant models, but I haven’t gone through that book yet.
Everyone posts wonderful images and then every single time I try to get the damn things (all of them) to draw something for me, the results are absolute garbage.
People think they can waltz in and immediately get great results from using AI's to generate images... and they can, if they're lucky or if they copy somebody else's prompt.
It's a lot harder to do so consistently, or if you want your images to look both good and original, and not like mere copies of what everyone else is doing.
Then I copied someone's prompt and got really great ones.
I suspect eventually there will be tools you can just fire up with no knowledge, but all of them I've seen so far still do require a bit of expertise and time.
1.ask chatgpt to generate a prompt of what you want by giving it a few exemples from a random SD prompt sharing website.(this alone gave me stunning results)
2.(optional) Use Controlnet for the pose you want,from the posture of the body down to each finger individually.
2.5 use multi Controlnet for multiple characters.
3. correct any errors with img2img.
4. Enjoy
It takes 10 to 20 minutes (mostly in getting a good pose) but the results are always good and you can later reuse the pose again.
Maybe there is something fundamentally difficult about representing hands ?
But I think the more probable explanation is that I'm just trying to find a correlation where none exists.
That’s some of the hardest geometry for a mind to envisage without some kind of construction process.
When you’re imagining or dreaming, you don’t have a construction process to make them look good.
"Yep."
"Yeah but a specific Appalachian human soul at around four o'clock in the afternoon on a day in mid-autumn when it looked like it would rain but then it didn't?"
"Also yep."
"Yeah but specifically at 3:56pm and the human in question is standing on loam and holding a book in their left hand and listening to Music For The Royal Fireworks by Handel?"
"Uhhh..."
"See, told you AI is useless."
"Midjourney v5 can do hands"
"Did you look at the hands it did? There are a bunch of mis-shapen blobs, hands with extra fingers, two thumbs on either side, etc."
"Sure, but there are also some accurate hands, so it can do hands."
"HAH! Here is a counterexample where X is not solved, so it can NOT!"
If I say "here's a self-driving car!" and show you a video of a car moving straight down a street and stopping at a light, would you agree that I have a self-driving car? After all, it drove itself down the street.
It's a full blown web app with better options than the Discord bot, it has batch mode/select, remix, all upscale modes, works with every Midjourney engine.
They make it available to users who have generated more than 10,000 images as it's in alpha state and not able to withstand the load that the bot currently takes.
I believe after v5 focus they will make this web app public, but for now only a select few get to use it.
They warn users not to talk about it or share the link because they don't want it public until it's ready for full load which means over 10 million concurrent users.
Next project I want to heavly utilise image generation from the ground up - Midjourney looks really good, but needs better tools.
I don't think I could use it if I had to use the main public server.
Thanks so much.
P.S. for those, like me, who were confused about what "making your own server" means, you do this within the main Discord app. It doesn't involve provisioning an actual server and installing software. :-)
It doesn't help that the Discord search function is so terrible.
https://stallman.org/discord.html
In fact, it looks like modern ones can't either:
Midjourney is a small team. They are working on a web interface. But, won’t release it until it is significantly better than all the benefits they get from Discord. Meanwhile, they’ve been too busy making quality improvements and scaling the service to keep up with demand.
Of course automating image generation via any means (including the private api) goes against tos for good reason, I have never misused the api to generate images and have no plans to.
However I do use the api to download all my images and their metadata including prompts. Using the API I sync every image grid + 'upscaled' I have ever generated, generate a json file with all metadata including the full prompt and then use that to build my local archive.
In a way, a hand contains more features than perhaps the entire human body from a forms/intersection perspective.
I think an entire model needs to be built to focus on just hands and then combined into a more general model, perhaps that will be the path forward?
We are almost there I guess. Just need to add the concepts of known proportions to image construction.
Seems like Dall-E is what I have to use for now :/
These models don't understand relationships between objects in a scene, especially between distant objects. So they can't do hands for the same reason they can't get legs on a table right. They know roughly what a table and a table leg look like, but they don't understand that there needs to be 3-4 of them at least, and they need to be spaced so that the table sits level, and the perspective they should have as a result. So, I've seen tables where it kind of gets it right that the legs are in the corners but then as the table legs go down, the front ones are mysteriously behind something that ought to be under the table. And sometimes it kind of loses track of a table leg or two - they melt into the background.
Very similar problem with hands. They need a very specific orientation and shape and the fingers all need to consistently point in the right direction, and typically the same direction (except for when they don't like with a pointed finger, etc).
Curious as to how these models handle it so much better than prior generations. Is it something novel, or a specific hand-based fix they put it, or is it just "we made the model bigger"?
The fingers themselves are also almost identical, but not really. If you learn a "platonic finger" it's not good enough, you should learn each finger individually. There is only so much you can spend on them, you got a million other things to learn. And the raters of the model are much more likely to penalize a bad face than some off details in a hand.
#tooSoon