Now there's a lot of AI results coming up and it's not performing much better in that regard either though.
You can see that some people have a fundamental misunderstanding about how bicycles work, and this misunderstanding does not stem from the complex 3D shape of bicycles.
Don't you think drawing a hand is a bit of an unfair "gotcha" for an AI that can accurately draw thousands of other things? Humans have hands and they're one of the most important parts of the body; humans use hands for everything. And yet, all but the most artistic humans cannot draw human hands accurately. An AI doesn't have hands. Hands make up a small part of its training data. Why should we consider the hand the standard upon which to judge the quality of the AI?
I guess it would take someone who can look at and interpret the code to figure it out.
We use hands, recognize hands, some of us can draw or sculpt accurate representations of hands, we can talk about aspects of hands such as grip and fingerlength. All of these things encompass the idea and physical instance of "hand" in much more than just the artistic sense.
Also, we certainly know when we draw a hand poorly. Unless we are children, or have a brain condition. And if we are children or have a brain condition other adults wouldn't say that we have a good understanding of what a hand is.
Perhaps the AI knows it's drawing hands poorly. Perhaps it can't communicate this to us. Even so this should be in the code run somewhere.
> "while simultaneously saying the AI draws hands poorly because it doesn't understand hands?"
For two reasons:
1) The AI is just drawing. Even humans draw many things that they don't understand. If we draw something that we completely don't understand (such as through random scribbling) we don't even call it a representation. It's a fluke. I used to scribble and then trace images in my scribbling (if possible). Often I ended up tracing things that looked like a child's bad drawing of Donald Duck, but once, without having to trace particular lines at all, my scribbling was a perfect seeming of a rose flower (with some minor additional flourishes). I recognized the rose flower, but I certainly didn't set out to draw it.
2) I'm assuming the AI wasn't trained on medical and other information pertaining to a hand, the way humans are (even if the training isn't formal). It is trained on images. At best it is trained on images of discrete parts of the body, but this only allows it to understand the shape and relative position of each body part, at best. Not to understand the body part.
Ultimately it would take something looking at the code run to determine whether or not the AI brings an understanding of the concept of "hand" into the literal picture. If all it's bringing is #1 then it's not an understanding of "hand".
Sure, but the AI doesn't use hands so why would hands be of particular importance to the algorithm? Let's say the algorithm did draw perfect hands every time. Would you accept that it has understanding then? I doubt it. Your argument is essentially a slippery slope: no matter what it can do, you'd find something it didn't excel at and say "See, it has no understanding."
> Also, we certainly know when we draw a hand poorly. Unless we are children, or have a brain condition. And if we are children or have a brain condition other adults wouldn't say that we have a good understanding of what a hand is.
I've seen plenty of people claim they produce great art, when in reality it's terrible.
This "hand understanding" could probably be simplified for an artistic concept of "hand".
I believe a computer "understands" basic arithmetic on a non-conscious level. I haven't been convinced it understands the human hand, even in an artistic sense.
> "I've seen plenty of people claim they produce great art, when in reality it's terrible."
Yes. But other people typically don't go around saying that those people understand great art. Understand what makes it great.
As for the bikes, I'm not surprised most people cant recall the exact shape or design of a mechanical object they don't use or work on every day. I'm sure a hobby cyclist could accurately lay out one.
Is it? Humans understand that humans have hands, yet most are incapable of drawing a hand well. The AI also clearly understands that humans have hands. It does put something hand-like where a hand typically goes. But, like most humans, the AI is not good at accurately reproducing a hand.
What is the definition of "understanding on an intellectual level?" I have a friend who is an artist. She is doing a series of paintings with figures. She cannot paint hands to save her life. The hands she paints look very similar to the eerie vignettes produced by Stable Diffusion. I fail to see what makes the hands my friend paints somehow superior on an "intellectual level."
Edit, to add some context. How many of the hands your friend paints will look like the ones in this link?
https://huggingface.co/spaces/stabilityai/stable-diffusion/d...
I think you're getting way too hung up on the number of fingers. It's clear that the algorithm could do a better job with hands given more training data. There is nothing special about hands. So, if you're right that it doesn't understand hands, then this lack of understanding stems purely from a lack of training data. In which case, I think it's fair to say that it understands the other things it has sufficient training data to draw well.
This is like the joke about training a neural net on arithmetic where you get the wrong answer repeatedly until it remembers to answer 5+5=10. but then, until there are more data, 10+5 is also 10, because it didn’t actually understand arithmetic (to be fair, a human wouldn’t understand it by a single example either).
And you can see this in action by making ChatGPT believe 5+3=7.
ChatGPT manipulates symbols, and it captured the rules to do that very well. That is one of the abilities of general intelligence, but it’s not the only criterion for intelligence. You can do more than just manipulating symbols, you can also abstract over them, form your own thoughts and questions about them, be curious, reflect on your reasoning and explain it, deduct patterns from very few data because you have all the context from your previous knowledge and the abstractions built on top of it. Besides of the abstractions and functions programmed into it ChatGPT only has probabilities of symbols being related to other symbols (and the rules implied by that), but it cannot reason about these symbols and cannot form creative thought. Its „intelligence“ is limited to a finite order/level of abstraction (the features and parameters that define the model and allow it, for example, to capture shading and geometry, but not the concept of a human hand) while yours is basically limitless. You can always put another abstraction on top of what you just thought or experienced. The magic of deep learning was basically increasing the order of abstraction a neural net can capture, but it’s still limited.
On the other hand, I have my pet theory that the weirdness of dreams or psychedelics arises from the brain basically sampling the connections in the brain / piping random noise through its neural net (as a side effect of all the reorganization it’s doing).
Yet it correctly understands that faces typically have two eyes, one mouth, one nose, etc. So clearly this "lack of understanding that hands have five fingers" is unlikely to be inherent to the model.
Let's say I ask you to draw a lady bug. You'll draw a red shell with some black dots haphazardly strewn about. However, the most common lady bug in Europe always has 7 spots. It's unlikely that your drawing will reflect that. Why? Because you lack understanding of Coccinella septempunctata. But does that you mean you lack understanding in general? Of course not. Lady bugs simply aren't important to you.
So again, why are we elevating hands to be the litmus test of understanding? Yes, hands are important to humans. But this algorithm is not a human, so hands are no more important to it than anything else it can do. Like let's say if could draw perfect hands 100% of the time. Does that mean you would concede that it has understanding? I doubt it. You'd pick some other thing it didn't do well and say "See, it can't accurately draw eggs stacked in a pyramid, therefore it lacks understanding." The issue with your argument is that is a slippery slope without a specific reason why the correct rendering of hands is important.
And I'm not arguing that GPT-3 or Stable Diffusion are omnipotent. Clearly they're not. But that doesn't mean that can't understand things in their domain. As others have mentioned in adjacent comments, the only test we have for understanding, in humans or ML models, is measuring the correctness of an output for a given input. Essentially, your argument is that "It's an algorithm, it can't understand like a human," which is begging the question.
I'm not claiming that ChatGPT or any other ML algorithm is "generally intelligent." Just that it has an understanding of certain concepts.
It is, we just don’t know it’s exact features. It might very well be optimized for recognizing faces (and therefore to identify the features that make up a face). A general AI doesn’t have to be retrained on specifics. Sure, you „can“ (in a very generous hypothetical sense of the word) train a model like this on all pictures and movies in existence and then some, so that it has seen everything and never fails to give the wrong answer for any prompt that only involves things that were depicted at some point in time. You don’t have to show a child all hands on the planet for it to recognize hands have 5 fingers. You don’t even have to show children pictures of every body part once for every skin color. They only need to see 1-2 different skin colors once to make the deduction that every body part can come in different skin colors. That‘s understanding, a general intelligence. Try this with a model like ChatGPT and you get a racist model.
>And I'm not arguing that GPT-3 or Stable Diffusion are omnipotent. Clearly they're not. But that doesn't mean that can't understand things in their domain. As others have mentioned in adjacent comments, the only test we have for understanding, in humans or ML models, is measuring the correctness of an output for a given input. Essentially, your argument is that "It's an algorithm, it can't understand like a human," which is begging the question. I'm not claiming that ChatGPT or any other ML algorithm is "generally intelligent." Just that it has an understanding of certain concepts.
We can also inspect the model. And even the ouputs are obviously different from what a human would be able to output, so it fails even that test.
If you argue that’s still intelligence, just on a lower level, you can absolutely do that. But at that point you’re basically saying everything is intelligent/conscious just on varying levels. In the sense that consciousness is what consciousness does. Which is a stance I generally agree with, but it’s also unfalsifiable and therefore meaningless in a scientific discussion.
Children do all the time, and I'm pretty sure children have a rudimentary understanding of hands.
ChatGPT knows that hands have five fingers because it's trained on text and text will almost always say that.
DALL-E doesn't know that hands have five fingers, it just knows what hands generally look like, and the number of fingers on a hand is just one of the elements it tries to match. At a glance, there are far more important elements of what a hand looks like, such as the shading.
Neither of these mean either AI is stupid. DALL-E doesn't have a concept of what a hand is, it just has an idea of what is looks like, and it's decent at recreating that.