No, DALL-E doesn’t have a secret language
twitter.com
twitter.com
DALL-E 2 has a secret language - https://news.ycombinator.com/item?id=31573282 - May 2022 (109 comments)
Think about the "bouba"/"kiki" effect [1], or how the spells in Harry Potter sound recognizably "spell-like" even though spells aren't a real thing. (In the latter case, it's because they're phonetically Latin-adjacent.)
For example, here are some nonsense terms:
* Swith'eil Aerveid
* Karflon B43
* Hooboo skramp
If I were to ask you all which is a dangerous gas, a sex act, and an Elven city, I suspect there would be fairly wide agreement. Inventing words that are evocative this way is a fundamental part of fiction writing, especially in fantasy, horror, and sci-fi. There's nothing profound going on here, it's just that language recognition is fuzzy and highly associative even at the level of individual phonemes.
Cryptonomicon is also great.
When I use or think in English, my brain has weird rules of its own. Some person can have his/her pronoun constantly wrong. For some it's random. Context and rhythm can also can create pattern.
Abra kedabra
“Abracadabra belongs to Aramaic, a Semitic language that shares many of the same grammar rules as Hebrew, says Cohen in Win the Crowd.
‘Abra' is the Aramaic equivalent of the Hebrew 'avra,' meaning, 'I will create. ' While 'cadabra' is the Aramaic equivalent of the Hebrew 'kedoobar,' meaning 'as was spoken.”
https://www.haaretz.com/2013-05-08/ty-article/word-of-the-da...
From the book: "about a head shorter than Harry [...] swarthy, clever face, a pointed beard [...] very long fingers and feet”, indeed.
I lol'd at the "Jewish culture" tag.
(granted, having the goblins run the banks is... not great)
Involving genetics and evolution in etymology makes no sense.
Ask a person off the street to come up with an "alien planet in a science fiction TV series" or "deadly gas/virus from a comic book" and you'll get similar answers.
Your average person on the street could probably successfully come up with trope-fitting "evil alien species" and "peaceful alien species" names, too.
I love doing the bouba/kiki to ppl who know nothing of it. So far none of the "test" have given anything different that what the studies show.
Besides, here is one instance where they failed to reproduce this result in Papua New Guinea.
But when perhaps 90% or more of languages do it a certain way on every continent, one does wonder if our thinking tends a certain way. General word order is another an example like that. The indirect object, or patient, or similar category usually is not the first word in a phrase or sentence. "The dog I see" for "I see the dog". "The money him the programmer gave." Like Yoda's species. Only 1 - 2% of languages do it that way consistently, and usually for recent grammatical development reasons, not a long-preserved feature.
Similar story in phonetics. Nearly all languages have at least three stop consonants, usually p, t, k. Closing against the lips, teeth and back of the mouth in the (perhaps obvious?) positions in the mouth to make stop consonants. A few languages only have two of these positions. Some have four or even five, using the throat and using the palate, too. But >95% of them, all over the place, have at least the p, t, k stops. We aren't hardwired, I suspect, with that particular set of noises like we are with crying as a baby. You have to learn them all. It is cultural. Yet almost every culture has converged on at least a somewhat-similar subset. And a few have not.
There are multiple ways to interpret such results, but I lean to something like convergent and parallel cultural evolution.
So far as I'm aware, that's a very fringe idea in linguistics, and pretty firmly rejected, despite many investigations.
Rather, what you're seeing is propogation from older language influences. Things like Proto-Indo-European have very, very widespread influence - but are still very much not universal.
So there's no consensus on why these patterns have these tendencies. The sample size is quite large. There's several hundred documented, not-known-to-be-related language families, and language isolates, that must have been separated historically from each other for at least a couple thousand years (or their relation would be pretty obvious). It does look like the tendencies mentioned are retained for a long time in separate populations despite language change, for some reason. That it might be because it relates to how we think is speculation, admittedly.
You can read more about the phenomenon here: https://en.wikipedia.org/wiki/Word_order#Distribution_of_wor... It is a rather striking distribution, isn't it?
No, not at all striking. [0] Merely a reflection from the older influences. Many, many languages that are not grouped in the same families, _are_ descended from PIE, and its children. That is to say, they may not traditional grouped in the same family, but would probably still be considered to be in the Indo-European family.
For an example of this "same family" thing, Japanese, Tamil, Quecha, are all probably descended from Proto-Uralic [1], despite not traditionally being considered sister languages.
[0] https://en.wikipedia.org/wiki/Proto-Indo-European_language#S...
If we're just going to assume on very loose basis from typological comparison, we might as well assume proto-World, because that's what the sum evidence suggests. But that guess (an idea I do take seriously) is very different from demonstrating their relationship through solid comparative and historical linguistics.
[1] https://en.wikipedia.org/wiki/Den%C3%A9%E2%80%93Yeniseian_la...
I assume, then, that you've never heard of the Borean languages [0]. Which link the people of southern Asia, northern Europe and South America, culturally, with a common origin around 40-45,000 years ago. The Borean hypothesis is a claim that is made quite seriously, but does not presuppose a universal origin. It is wide reaching, but there are people groups outside of it.
It's an interesting idea, I cannot disprove it. But I don't think it has been proven, either. Even the possible relationship between Uralic, Afroasiatic and Indo-European is not widely accepted yet. Those language families are extensively documented and we can reconstruct their proto-languages convincingly back to somewhat around 10,000 years ago. They look kind of similar in some ways, maybe areal effects? I think to hope for another 10,000 years further back is too much. Borean is a claim about probably further back than even that. The Nostratic and Altaic subsets of the Borean hypothesis, presumably with its proto-language somewhere in Asia around 10,000 years ago, alone is controversial and is not generally accepted.
I think a phycist would feel similarly if it was attributed to "human energy" or something like that.
Sure almost all languages share consonants, but there are only so many sounds our mouths can make that has nothing surprising. Yet there are exceptions like Dahalo which is a language made of clicks.
Sure there are similar words from different roots. Say in the 3 languages I speak, mom/maman/mama, dad/papa/tata, no/non/ne. But there's no reason to involve genetics. Some sounds are easier to make for babies/toddlers, it makes sense they would converge toward for universal words like parent appellations.
Protein synthesis doesn't have to play a part in it no matter how magical genetics may seem.
Swith'eil Aerveid, which I took to be the Elven city, immediately made me think of Morrowind. I think it's a perfectly fitting name for a Dunmer or Dunmer city. Maybe a dwarven ruin.
[1] https://giannisdaras.github.io/publications/Discovering_the_...
[2] https://twitter.com/giannis_daras/status/1531693093040230402
This is largely some embedding of semantics that we currently do not fully have a mapping for, precisely because it was generated stochastically.
Saying it was "not true" seems like clickbait.
So, it's certainly an "interesting" result in the sense that it shows how these kinds of systems work, but it's definitely not a language.
They're not nonsense words, they're words with high probability which are not seen in the dataset.
Are there higher-resolution images to be had?
Like those AIs that guess what you draw, and recognize random doodling as "clouds", DALL-E is probably using the least unlikely route. That a gibberish word is drawn as a bird is maybe because it was "bird (2%), goat (1%), radish (1%)".
It's almost as-if its an illusion designed to fool, we, the users.. by only providing inputs meaningful to us, we come to the foolish idea that it understands these inputs.
Exactly. This could be profound. I'm looking forward to further work here. Sure, the examples here are daft, but developing this approach could be like understanding a talking lion [0] only this time it's a lion of our making.
[0] https://tzal.org/understanding-the-lion-the-in-joke-of-psych...
“Watch those sea creatures.”?
None of this makes DALL-E any less impressive to me. High quality image generation is a truly amazing result. Results from foundational models (GPT-3, PaLM, DALL-E, etc) are so impressive that they're forcing us to reconsider the nature of intelligence and raise the bar. That's a sign of a job well done to me.
As much as people would like there to be, there really does not seem to be anything here. The original author doesn't think so, either (would need to refind the tweet).
Just nitpicking here a bit. It has been trained to ensure the validity of mappings, but only for mappings of valid prompts, where "valid" is vaguely "things that appear in the training set". On the other hand, it wasn't trained to ensure the validity of mappings of invalid prompts to images. It's an open question what that would even mean - someone in another thread here suggested it should output "not a valid prompt" in this case.
This is by far the worst.
I also think you could probably hire a team of linguists pretty cheap compared to a team of AI engineers.
Also, there is no way for vocabulary to exist on its own without grammar, as these are two sides of the phenomenon, we call language. Some signs of grammar had to emerge together with this, at once. However…
----
Edit: Let's imagine a typical movie scene. Our nondescript individual points at himself and utters "Atuk" (yes, Ringo Starr!) and then points at his counterpart in this conversation, who utters "Carl Benjamin von Richterslohe". This involves quite an elaborate system of grammar, where we already know that we're asking for a designator, that this is not the designator for the act of pointing, and that by decidedly pointing at a specific object, we'd ask for a specific designator not a general one. Then, C.B. von Richterslohe, our fearless explorer, waves his hand over the backdrop of the jungle, asking for "blittiri" in an attempt to verify that this means "bird", for which Atuk readily points out a monkey. – While only nouns have been exchanged, there's a ton of grammar in this.
And we haven't even arrived at things like, "a monkey sitting at the foot of a tree". Which is mostly about the horizontal and vertical axes of grammar, along which we align things and where we can substitute one thing for another in a specific position, which ultimately provides them with meaning (by what combinations and substitutions are legitimate ones and which are not).
Now, in light of this, that specific compounds are changing their alleged "meaning" radically, when aligned (composed), doesn't allow for high hopes for this to be language.
I wonder what level I would be able to share ideas I lack the words for, my perceived bitrate at creating "random" noise is certainly higher than when verbally communicating an idea to another human. Will we even share a common language in the future? Or will we have our own language that is translated to other people?
> 5.62 (…) For what the solipsist means is quite correct; only it cannot be said, but makes itself manifest. The world is my world: this is manifest in the fact that the limits of language (of that language which alone I understand) mean the limits of my world. [1]
So, something could become apparent, but you would still haven't said anything (as it's not part of that conversation). ;-)
[1] https://www.masswerk.at/digital-library/catalog/wittgenstein...
(I deem this edition to be somewhat appropriate in context.)
To clarify: that these triggers do produce (radically) different results when provided in varying compositional context, like "as cartoon", "as painting", etc, (i.e, "a as b") suggests that these are just random alliterations that serve merely by accident as synonyms, rather than having a specific value in that position. The latter being a requirement for language. (And it wouldn't be too far fetched to supposed that these are just some "residual" values, when showing up in pseudo-textual compositions. As there had to be some trigger for that.)
I believe the original claim has substance behind and it was a very interesting non-trivial observation.
Also, it proposes a very exciting idea for an emergent phenomena that, if understood, could have deep consequences to our understanding of knowledge, language and a lot of other related topics.
Though I'll also point out that even evidence for that weaker claim is tenuous. It definitely knows how to move an image closer to "3D render" in concept-space, but it doesn't seem to understand the linguistic composition of your request. For example, you'd have an extremely hard time getting it to generate an image of a person using 3D rendering software, or a "person in any style that isn't 3D render"; it would probably just make 3D renders of persons.
I haven't played around with it myself, I'm going off the experiences of others. For example:
https://astralcodexten.substack.com/p/a-guide-to-asking-robo...
These absurdly big, semi-supervised transformers are predicting what the next pixel or word or Atari move is. They’re strikingly good at it. To accomplish this they build up a latent space where all the pictures of sunglasses and the word “shades” are cosine similar, and quite different to “dog” or a picture of a dog, and have an operator (in word2vec, addition, in DALL-E, something nonlinear) that can put sunglasses on a dog.
Is that latent space and all the embeddings into it a “language”? Who cares? It works and it’s fucking cool.
I dont think the latter have much interesting to say about the former; but, having done no research, they think they do.
Also, I'm wondering if there is some way that these models could have a decent error response rather than responding to every input?
It is trivial to make it reject gibberish prompts. Just use a generative model to estimate the probability of the input, it's what language models do by definition.
Would still be interesting to see how the output changes with little changes to these inputs. If my vague understanding is at all close, this will reveal the “faces” that are more “noisy” than the others. Not sure what that gives though.
1. Tokenized text embeddings can map to similar points in latent space. The text encoder is autoregressive so this won’t work for all random sequence of tokens but it can work for the right ones. I wonder if anyone has tried reverse decoding the embeddings of interest to see if they cluster around known words that are relevant.
2. Diffusion models are trained by pushing off manifold points onto the manifold so to speak. It is not surprising that off manifold points map onto known concepts during the reversal process.
IMO these words were part of some training images (e.g. taken from nature atlases) and DALL-E learned to associate them with birds, although in gibberish form.
For example "hedge" combined with "hog" is neither a "hedge" nor is it a "hog" nor is it some sort of horrific hybrid mixture of hedges and hogs. A hedgehog is tiny rodent. Most likely this is what's going on here.
The domain is almost infinite. And the range is even greater. Thus it's actually realistic to say that there are must be hundreds of input and output sets that form alternative languages.
But if one of those probes towards unrewarded input produces a correlation then SOME side effect is influencing it. It means there can be side effects ALL over the unrewarded space.
That being said the reward space VS. unrewarded space is tiny. 1 over infinite for all intents and purposes. It's basically all combinations of letters in reality vs. all possible grammatically correct English combinations.
The unrewarded space is massively huge. Within that unrewarded space given how massive it it is, there is actually a very high probability that there is at least several sets of inputs in there that form a grammar and a consistent language with consistent outputs.
But these sets are hard to find you can't just pick anything. If you pick one secret word that has a correlation with birds, then you mash it up with english expecting it to stay coherent... well that's simply an invalid set.
You need to find the OTHER words that work in conjunction with the secret bird word. There are millions likely, but they will likely be impossible to find. Library of Babel vibes, https://maskofreason.files.wordpress.com/2011/02/the-library....
This could actually be an interesting project. Some algorithm that explores this space attempting to map out connections. It would have to be another ML algorithm, but likely that search is still never ending; but like the library of babel, in terms of probability something must be out there that works.
I think this is a large logical leap. The unrewarded space may be huge, but it is not large enough that it's almost guaranteed (or even close) that we'll find something that looks like language in there. If we did find something that looked like language, even in the unrewarded space, it would be very surprising, which is why the initial post that inspired this response was so talked about! But we have not.
It's like if I gave you a lottery ticket and you won. The lottery ticket is more likely to be rigged then you actually winning. Or in other words, the fact that we even found a secret word is indicative that there's a lot going on in the unrewarded space.
Something is going on if even some random word produces a correlation of some sort.
> Tries a lot of prompts that generates things to a common theme.
> "To me this is all starting to look a lot more like stochastic, random noise, than a secret DALL-E language."
Some of the whale dialogue is clearly transcribable, but he regenerates it again until he gets "Evve waeles" and answers that resemble "Wales".
https://nitter.net/benjamin_hilton/status/153178089297217536...
A bunch of people didn't read the original study and just saw the pictures and assumed the gibberish is the only result being discussed.
https://www.newscientist.com/article/2114748-google-translat...
I am more likely to believe celebrity gossip than AI news articles.
For example, I could write a heuristic algorithm to product the same thing using a Google image search, but it would look like MS word clip art.
The answer is that the training process literally has to make the results smooth. That’s how training works.
Imagine you have 100 photos. Your job is to classify them by color. You can place them however you want, but similar colors should be physically closer together.
You can imagine the result would look a lot like a photoshop RGB picker, which is smooth.
The surprise is, this works for any kind of input. Even text paired with images.
The key is the loss function (a horrible name). In the color picker example, the loss function would be how similar two colors are. In the text to image example, it’s how dissimilar the input examples are from each other (Contrastive Loss). The brilliance of that is, pushing dissimilar pairs apart is the same thing as pulling similar pairs together, when you train for a long time on millions of examples. Electrons are all trying to push each other apart, but your body is still smooth.
The reason it’s brilliant is because it’s far easier to measure dissimilar pairs than to come up with a good way of judging “does this text describe this image?” — you definitely know that it isn’t a bicycle, but you might not know whether a car is a corvette or a Tesla. But both the corvette and the Tesla will be pushed away from text that says it’s a bicycle, and toward text that says it’s a car.
That means for a well-trained model, the input by definition is smooth with respect to the output, the same way that a small change in {latitude,longitude} in real life has a small change in the cultural difference of a given region of the world.
gwern helped too. He has an intuition for ML that I’m still jealous of.
Your best bet is to just start building things and worry about explanations later. It’s not far from the truth to say that even the most detailed explanation is still a longform way of saying “we don’t really know.” Some people get upset and refuse to believe that fundamental truth, but I’ve always been along for the ride more than the destination.
It’s never been easier to dive in. I’ve always wanted to write detailed guides on how to start, and how to navigate the AI space, but somehow I wound up writing an ML fanfic instead: https://blog.gpt4.org/jaxtpu
(Fun fact: my blog runs on a TPU.)
I’m increasingly of the belief that all you need is a strong desire to create things, and some resources to play with. If you have both of those, it’s just a matter of time — especially putting in the time.
That link explains how to get the resources. But I can’t help with how to get a desire to create things with ML. Mine was just a fascination with how strange computers can be when you wire them up with a small dose of calculus that I didn’t bother trying to understand until two years after I started.
(If you mean contrastive loss specifically, https://openai.com/blog/clip/ is decent. But it’s just a droplet in the pond of all the wonderful things there are to learn about ML.)
ML is simply curve fitting. It's a applied math problem that's quite common. In fact I lost a lot of interest in intelligence in general once I realized this was all that was going on. The implications really say that all of intelligence is really some form of curve fitting.
The simplest form of this is linear regression which is used to derive an equation for a line from a set of 2D points. All ML is basically a 10,000 (or much more) dimensional extension of that. The magic is lost.
Most of ML research is just to find the most efficient way to find the best fitting curve given the least amount of data points. A ML guys knowledge is centered around a bunch of tricks and techniques to achieve that goal with some N-D template equation. And the general template equation is all the same: A neural network. The answer to what intelligence is seems to be quite simple and not that profound at all... which makes sense given that we're able to create things like DALL-E in such a short time frame.
One of the big mysteries of the universe (intelligence) and the thing I always wondered about was essentially answered within the last 2 decades which is pretty cool.. but it's like knowing the secret behind an amazing magic trick.
I made that by using ML as a guitar. I chose instruments and style the way a guitarist’s fingers chooses frets.
And saying “give me this style with these instruments” is far easier than recording it yourself.
For what it’s worth, I agree with you about AGI. https://twitter.com/theshawwn/status/1446076902607888385?s=2...
But for me, that means it’s far more interesting than AGI. Everyone has their eye on AGI, and no one seems to be taking ML at face value. That means the first companies to do it will stand to make a fortune.
What was your point here? ML is like a guitar? What you said doesn't seem to contradict anything I said other then you find curve fitting interesting and I don't.
Not trying to be offensive here, don't take it the wrong way.
On a side note your music example also basically destroys the question of what is music? Well the answer to that question is that music is the set of all points on some sort of N-dimensional curve. The profoundness of the question is completely gone.
Human thought and and human intelligence on the other hand was a great and epic concept on the scale of the origin of the universe. It was truly this mysterious epic thing that seemed like something we would never crack. ML brought it down and completely reduced this concept and simplified it by a massive scale. The entire field is now an extension of this curve fitting concept. And the disappointing thing is that the field is correct. That's all intelligence is in the end.
This is all I mean. Not saying ML is less interesting or easier then any other STEM field. All I'm saying is the reduction was massive. The progress is amazing but 99% of the wonder was lost. The scale at which we lacked understanding was covered in a single step and now the average person can understand the basics easier than they can understand something like quantum mechanics. There's still a lot going on in terms of things to discover and things to engineer, but the fundamentals of what's going on are clearer than ever before.
So I think what happened here is that you mistook what I wrote and took offense as if I was attacking the field, I'm not. I'm writing this to explain to you that you're mistaken.
So dial your aggressive shit back. Is everyone from Romania like you? I certainly hope not.
The interesting thing to me about the secret language is that it seems to imply that when DALL-E fit words to concepts, it created extrapolations in its curve fit that are more extreme than the actual training samples, ie. its fit has out-of-domain extrema. So there are letter sequences that are more "a whale on the moon" than the actual text "a whale on the moon." Linguistic superstimulus.
Regarding the gibberish word to image issue - CLIP uses a text transformer trained by contrastive matching to images. That means it's different from GPT, where it trains to predict the probability of the next word. GPT would easily tell apart gibberish words from real words, or incorrect syntax because they would be low probability sequences. CLIP text transformer doesn't do that because of the task formulation, not because of an intrinsic limitation. It's not so mysterious after realising they could have used a different approach to have both the text embedding and filter out gibberish if they wanted.
A good analogy would be a Rorschach test - show an OOD image to a human asking him to caption it. They will still say something about the image, just like DALL-E will draw a fake word. It's because the human is expected to generate a phrase no matter if the image makes sense or not, and DALL-E has a similar demand. The task formulation explains the result.
The mapping from nonsense word to image is explained by the continuous embedding space of the prompt and the ability to generate images from noise of the diffusion model. Any point in the embedding space, even random ones, fall closer to some concepts and further from other concepts. The lucky concept most similar to the random embedding would trigger the image generation.
There is for sure a set of consistent words that produce output that makes sense to us. He just picked the wrong set!