No. You know what a giraffe is, Dall•E simply creates pixel groups which correlate to the text pattern you submitted.
Watching people discuss a logical mirror scares me that most people are not themselves conscious.
No. You know what a giraffe is, Dall•E simply creates pixel groups which correlate to the text pattern you submitted.
Watching people discuss a logical mirror scares me that most people are not themselves conscious.
They just assert it as axiomatic, whistling-past all the ways that they themselves – unless they believe in supernatural mechanisms – are also the product of a finite physical-world system (a biological mind) and a finite amount of prior training input (their life so far).
I'm beginning to wonder if the entities making this argument are conscious! It seems they don't truly understand the issues in question, in a way they could articulate recognizably to others. They're just repeating comforting articles-of-faith that others have programmed into them.
Though, I guess maybe that is all “knowing what a giraffe is” is?
Would a person who only ever read articles and looked at pictures of giraffes have a better understanding of them than Dall-e does? At some level, probably, in that every person will have a similar lived experience of _being_ an animal, a mammal, etc. that Dall-e will never share. Is having a lesser understanding sufficient to declare it has no real understanding?
This is basically on the same level as our visual cortex and some form of memory. But our visual cortex isn't sufficient to give us consciousness.
I took a quick look at the Stanford Encyclopedia of Philosophy entry for philosophical zombies ( https://plato.stanford.edu/entries/zombies/ ) and I can't see evidence of this argument having been seriously advanced by professionals before. I think it would go something like:
"Yes, we have strong evidence that philosophical zombies exist. Most of the laypeople who discuss my line of work are demonstrably p-zombies."
("A laser is trying to find the darkness...")
But saying "It’s simply a high dimensional latent mapping between characters and pixels" is clearly a very bad argument. Your brain is simply a high dimensional latent mapping between your sensory input and your muscular output. This doesn't make you not intelligent.
It definitely does more than that.
They'll find some morsels of fragmentary hints-of-meaning in the junk, or just act from whatever's bouncing around in their own 'ground state', and make something interesting & coherent, to please their interlocutor.
So I don't see why this corner-case impugns the level-of-comprehension in DALLE/etc – either in its specific case, nor in the other cases where meaningful input produces equally-meaningful responses.
In what ways are you yourself not just a "very complex & impressive mirror", reflecting the sum-of-all external-influences (training data), & internal-state-changes, since your boot-up?
Your expectationthat random input should result in noise output is the weird to me. People can see all sorts of omens & images in randomness; why wouldn't AIs?
But also: if you trained that expectation into an AI, you could get that result. Just as if you coached a human, in a decade or 2 of formal schooling, that queries with less than a threshold level of coherence should generate an exceptional objection, rather than a best-guess answer, you could get humans to do so.
Of course. And the equivalent of that explanation is baked into DALL-E, in the form of its programming to always generate an image.
> but DALLE can not act like humans
No, not generally, but I don't think anyone has claimed that.
Given the model has only 1 input and 1 output and training is essentially surviving that order, it's not dissimilar.
So is a single-purpose AI equivalent to the entirety of the Human Experience? Of course not. But can it be similar in functionality to a small sliver of it?
One great example of this phenomenon is "Häagen-Dazs".[0]
Admittedly that's a brand name, rather than a specific dish, but I assume that Dalle-2 would generate an image of ice cream if given a prompt with that term in it (unless there is a restriction on trademarks?).
[0] https://funfactz.com/food-and-drink-facts/haagen-dazs-name/
This is probably because your string is a low-entropy keyboard-mash.
Asking it to produce noise, or raise an objection that a prompt isn't sufficiently meaningful to render, is a silly standard because it's been designed, and trained, to always give some result. Humans who can object have been trained differently.
Also, the GPT models – another similar train-by-example deep-neural architecture – can give far better answers, or give sensible evaluations of the quality of its answer, when properly prompted to do so. If you wanted a model that'd flag nonsense, just give it enough examples, and enough range-of-output where the answer your demanding is even possible, and it'll do it. Maybe better than people.
The circumstances & limits of the single-medium (text, or captioned image) training goals, and allowable outputs, absolutely establish that these are different from a full-fledged human. A human has decades of reinforcement-training via multiple senses, and more output options, among other things.
But to observe that difference and conclude these models don't "understand" the concepts they are so deftly remixing, or are "just a very complex and impressive mirror", does not follow from the mere difference.
In their single-modalities, constrained as they may be, they can train the equivalent of a million lifetimes of reading, or image-rendering. Objectively, they're arguable now better at composing college-level essays, or rendering many kinds of art, than most random humans picked off the street would be. Maybe even better than 90% of all humans on earth at these narrow tasks. And, their rate of improvement seems only a matter of how much model-size & training-data they're given.
Further: the narrowness of the tasks is by designers' choice, NOT inherent to the architectures. You could – and active projects are – training similar multi-modality networks. A mixed GPT/DALLE that renders essays with embedded supporting pictures/graphs isn't implausible.
Retrain dall-e and give it a choice whether it generates an image or does something else, and you will get a different outcome.
The argument boils down to this: is a human brain nothing but a mapping of inputs onto outputs that loops back on itself? If so the dall-e / gpt-3 approach can scale up to the same complexity. If not, why not?
So the interesting part is this, why did one random prompt fail in a consistent way and the other in a random way? Perhaps the encoding of meaning into vocabulary has patterns to it that we ourselves haven't noticed. Maybe your random string experiment works because there is some amount of meaning in the syllables that happened to be in your chosen string.
For example, the "yon" is immediately reconizable to me (hither and yon), so "yon corpti" could mean a distant corpti (whatever a corpti is). "becross" looks similar to "across" but with a be- prefix (be-tween, be-neath, be-twixt, etc.), so could be an archaic form of that. "chronly" could be something time related (chronos+ly). etc...
That morse code gives nothing useful probably just indicates some combination of – (a) few morse transcripts in training set; (b) punctuation-handling in training or prompting – makes it more opaque. It's opaque to me, other than recognizing it's morse code.
https://i.dailymail.co.uk/1s/2019/04/25/14/12707994-6959547-...
Below a certain level of complexity, perhaps only "vaguely like beetle or bicycle/race names" arbitrarily survived all the other competing influences on the model's internal weights. Above, much more subtle patterns – like many more latinate word roots – might start to survive.
Compare also what I linked in another thread branch – https://twitter.com/gojomo/status/1540095089615089665 – where simply growing the size of a similar (non-public Google PARTI) model leads to a phase-change in its ability to render meaningful human text.
You can easily see that these language models are in some sense working on fragments as much as they are on the actual words isolated by spaces in your sentence. Just take a test sentence and enter as a prompt to get some images. Then take that same sentence, remove all spaces and add new spaces in random locations, making gibberish words. You will see that the results will retain quite a few elements from the original prompt, while other things (predominantly monosyllables) become lost.
To me, I have not seen a single example that cannot just be explained by saying this is all just linear algebra, with a mind-bogglingly huge and nasty set of operators that has some randomness in it and that projects from the vector space of sentences written in ASCII letters onto a small subset of the vector space of 1024x1024x24bit images.
If you then think about doing this just in the "stupid way", imagine you have an input vector that is 4096 bytes long (in some sense the character limit of DALL-E 2) and an output vector that is 3 million bytes long. A single dense matrix representing one such mapping has 6 billion parameters - but you want something very sparse here, since you know that the output is very sparse in the possible output vector space. So let's say you have a sparsity factor of somewhere around 10^5. Then with the 3.5 billion parameters of DALL-E 2, you can "afford" somewhere around 10^5 such matrices. Of course you can apply these matrices successively.
Is it then so far fetched to believe that if you thought of those 10^5 matrices as a basis set for your transformation, with a separate ordering vector to say which matrices to apply in what order, and you then spent a huge amount of computing power running an optimizer to get a very good basis set and a very good dictionary of ordering vectors, based on a large corpus of images with caption, that you would not get something comparably impressive as DALL-E 2?
When people are wowed that you can change the style of the image by saying "oil painting" or "impressionist", what more is that than one more of the basis set matrices being tacked on in the ordering vector?
https://en.wikipedia.org/wiki/Bouba/kiki_effect
There's other evidence in support of this idea, but it would be a hassle for me to dig it up now.
>Your first random prompt is far from random. It contains the fragments "sublim", "chr", "cross" and "corpt" in addition to the isolated "E", which all project the solution down towards Latin and Christianity.
Oh right, I omitted a sentence. DallE gave me churches. DallE-mini gave me bicycle racing and beetles. Both of them behaved self-consistently, but were not consistent with each other. I did test changing some of the words that seemed like they might be having the steering effects like you pointed out. The new phrase was "E sublumary widge fraus chronly non estoi". DallE mini gave me mostly doctors/medical procedures with a few indiscernible scenes mixed in. Real DallE gave me scrabble, word tiles, pages from a book.
Its like this. If you talk to a dog with the "who's a good boy" voice, the dog will understand "who's a good boy", even if you're actually saying "who's a little asshole". Likewise, a pronounceable string will generally feel like it belongs within a certain envelope of meanings, even if its not actually a phrase in any real spoken language.
I showed some maps that DallE generated which happened to have some places labelled with its usual gibberish. People said the place names felt vaguely like they might be in Eastern Europe. Interpretation of nonsense words works both in the direction of human to DallE and DallE to human!
I suggest you look at the parent article.
Defining "understanding" in the abstract is hard or impossible. But it's easy to say "if it can't X, it couldn't possibly understand". Dall-E doesn't manipulate images three dimensionally, it just stretch images with some heuristics. This is why the image shown for "a cup on a spoon" don't make sense.
I think this is a substantial argument and not hand-waving.
True, it has some problems fully abstracting, and then logically-enforcing, some object-to-object relationships that most people are trivially able to apply as 'acceptance tests' on candidate images. That is evidence its scene-understanding is not yet at human-level, in that aspect – even as it's exceeded human-level capabilities in other aspects.
Whether this is inherent or transitory remains to be seen. The current publicly-available renderers tend to have a hard time delivering requested meaningful text in the image. But Google's PARTI claims that simply growing the model fixes this: see, for example: https://twitter.com/gojomo/status/1540095089615089665
We also should be careful using DALL-E as an accurate measure of what's possible, because OpenAI has intentionally crippled their offering in a number of ways to avoid scaring or offending people, under the rubric of "AI safety". Some apparent flaws might be intentional, or unintentional, results of the preferences of the designers/trainers.
Ultimately, I understand the practicality of setting tangible tests of the form, "To say an agent 'understands', it MUST be able to X".
However, to be honest in perceiving the rate-of-progress, we need to give credit when agents defeat all the point-in-time MUSTs, and often faster than even optimists expected. At that point, searching for new MUSTs that agent fails at is a valuable research exercise, but retroactively adding such MUSTs to the definition of 'understanding' risks self-deception. "It's still not 'understanding' [under a retconned definition we specifically updated with novel tough cases, to comfort us about it crushing all of our prior definition's MUSTs]." It obscures giant (& accelerating!) progress under a goalpost-moving binary dismissal driven by motivated-reasoning.
This is especially the case as the new MUSTs increasingly include things many, or most, humans don't reliably do! Be careful who your rules-of-thumb say "can't possibly be coceptually intelligent", lest you start unpersoning lots of humanity.
Your argument overall seems to take "you skeptics keep moving the bar, give me a benchmark I can pass and I'll show you", which seems reasonable on it's face but I don't think actually works.
The problem is that while algorithm may be defined by theory and tested by benchmark, the only "definition" we have for general intelligence except "what we can see people doing". If I or anyone had a clear, accepted benchmark for general intelligence, we'd be quite a bit further towards creating it but we're not there.
That said, I think one thing that current AIs lack is an understanding of it's own processing and an understanding of the limits of that processing. And there are many levels of this. But I won't promise that if this problem is corrected, I won't look at other things. IDK, achieving AGI isn't like just passing some test, no reason it should be like that.
You can read OpenAI's paper or try using it. They've intentionally not taught it many things; it doesn't know most celebrities' names, or copyrighted characters, and there's several filters before and after the big model to prevent NSFW generations of any type.
(Oddly, it does know "Homer Simpson" and "Hatsune Miku".)
That sound data sanitizing rather than the AI safety/danger that Bostrom and company worry about. For them, it's "OMG, must limit it's capabilities" rather than "Must keep it from doxing or offending people". The parent I was replying to was kind of ambiguous what sort of crippling they meant (We also should be careful using DALL-E as an accurate measure of what's possible, because OpenAI has intentionally crippled their offering in a number of ways to avoid scaring or offending people, under the rubric of "AI safety".)
It seems one of their techniques is to silently add extra words to some prompts to increase the variety of returned images in certain dimensions, eg: https://twitter.com/rzhang88/status/1549472829304741888
While I can't find the link right now, someone also tweeted examples of asking DALL-E2 for certain historical roles that would have been, by actual history and likely representation in training data, fairly uniform in race/sex – but the new OpenAI 'bias mitigations' generate ahistorical race/gender-varied examples. That's interesting – I'm a big fan of non-traditional castings like 'Hamilton!' etc – but serves to make the AI look ignorant of history, when in fact it has just been design-limited to simulate ignorance.
Prompts like "a movie still of ancient philosopher Heraclitus", "ancient philosopher Aristotle", and "famous physicists of the 17th century" render largely plausible portraits – except with an added sprinkling of ahistorical, 21st Century "AI safety" race/gender variety.
This is less a common-sense failure of the AI than an artifact of its creators' imposed harnesses.
And thus I'd suggest that more generally, unless a model's full design/training-data/constraints are fully described, it remains possible that other "failures" could sometimes be side-effects of undisclosed designer choices.
And as agents increasingly surpass that, the focus shifts, or retreats, to even murkier standards. "Sure, but it's just a mechanistic mirror, it doesn't really have [X, Y, Z] inside, even if it's externally indistinguishable from human excellence." So, an originally 'black-box' internally-oblivious test then shifts to a 'white-box' internally-dependent test, just to maintain the same comforting conclusion.
And by subtly shifting the grounds-of-evaluation, & definitions, it becomes easier to ignore massive improvements, & disturbing potential impacts. It creates an illusion of "zeno's paradox" where agents are never reaching the (shifting) endpoint, when in fact they're speeding past every robust test that can be devised. That's a really important thing to accurately notice!
Also, regarding: "current AIs lack… an understanding of it's own processing and an understanding of the limits of that processing"
People only have vague, incomplete understanding of their own reasoning, & limits, too. (Experts in restricted domains may be better – and those domains are often easier to train in an artificial agent.) The mere act of "thinking-about-thinking" can sometimes improve the quality of reasoning in humans – forcing something to be explainable – and it turns out to help in these modern large models too. Asking a GPT-like Large Language Model (LLM) to explain its reasoning step-by-step can improve its answers; prompting it with the info that it is an agent that may make errors, and asking it to state its confidence, also often generates more truthful, properly-qualified output.
So are the current LLMs that much behind, or inherently incapable of, human-level self-understanding in this sort of test? It seems an open question to me, and even if current SotA models miss a mark, perhaps 5x or 1000x larger models coming soon will outdo humans on any measurable test of "self-understanding", too.
The burden of proof is not on the one claiming logically consistent interpretations of events.
Examples of prompts that don't really work are "spherical giraffe" and "the underside of a giraffe".
We'll need more 3Dness to generate other media like models and video, so it should be coming.
If we assume this AI or a successor can win that evaluation, in what way would you say you know what a giraffe is better than the AI?
Perhaps a more interesting question could be: [how] do we know what consciousness is?
Dall-E will likely be similar in that it is effectively doing that perception step where you fix the text description from the classifier output and run that in reverse to show what the neural network is "seeing" when it is "thinking" about that given output. So it won't be able to describe features of a giraffe, or information about where they live, etc. but it will be able to show you what it thinks they look like.
[1] https://www.youtube.com/watch?v=AyzOUbkUf3M [2] https://youtu.be/AyzOUbkUf3M?t=1293
How would you tell the difference though? Can you think of a test to distinguish between those two abilities ?
*Don't mistake skepticism for knowledge*
This is a major problem on this site and elsewhere.