Is there "meaning" somehow inside the novel on my desk? Can you extract it and manipulate it? If I had no body, environment, temporality, what even could the book mean to me? Even more, if nobody could teach me the meanings of the words themselves in the book, what could it ever be to me but a collection of certain patterns?
I feel you might enjoy Chapter VI of Douglas Hofstadter's "Godel, Escher, Back", being "The Location of Meaning".
What is a word if it doesn’t reference anything, if it has no context? Isn’t this like a pointer without referencing an address?
I’ll give you that words can reference tokens of other modalities like hearing and such, but that is still the same problem.
For example, consider a word that means two different things in two different contexts that would otherwise be very far apart in whatever latent "meaning space" they might exist in. A naive model might end up averaging across both contexts, resulting in a representation of that word that represents neither meaning of the word, being located vaguely in between both regions of space. Alternatively, this could pull "groups" of words that otherwise are not related into adjacent regions of space, creating the appearance of semantic similarity when there shouldn't be any. In this example, the physical surrounding word context is actually misleading, and addition information is needed in order to make the learned embeddings more realistic.
In any case, it should be obvious to people trained in statistical practice that word models do not encode "meaning" per se, but rather an estimate of a model of meaning, which is not a complete model, but it does a pretty good job most of the time. We as the experts should be careful to stick to language along these lines, so as not to promulgate fanciful ideas about AI. I am not even going to bother with the "it's not really intelligent!" angle here. For all I care, the model is fully sentient. But the fact is that the token embeddings do not encode "meaning", it encodes "an estimate of model of meaning". They are our best effort attempt at encoding meaning parsimoniously, but that is not the same thing as meaning itself. Map is not the territory, all models are wrong, etc.
I think this one will get fuzzy real fast when you talk about cognition itself. Otherwise, thanks for the thoughtful response. Food for thought.
This is word salad. I don't even know what you mean by "meaning," but an embedding is simply an array filled with a bunch of floating point numbers[1] and, importantly, these embeddings are trained, so "doctor" will have an embedding similar to "hospital" because it's seen in close proximity in the training data.
The length of the number array is referred to as dimensions for math reasons. That's half the statement right there.
> importantly, these embeddings are trained, so "doctor" will have an embedding similar to "hospital" because it's seen in close proximity in the training data
It is trained to put together words with similar... meaning. So the embedding is a representation of... meaning.
This might be waxing philosophical, but my position is that the meaning of a word is intrinsic, not extrinsic. The meaning of the word "ball" has nothing to do with the fact that it's next to the word "net" on a page (which is what LLMs do). But rather, "ball" is the round rubbery thing I kick around in my back yard.
Based on you calling that phrase earlier “word salad”, I think you’re just not that familiar with word embeddings. https://en.m.wikipedia.org/wiki/Word_embedding
I suggest you read some Quine and Kripke; the argument that word embeddings encode meaning (particularly via locality) is an incredibly philosophically-naive position. People have been debating what language is for a very long time.
Maybe the thing that will land is to say that embeddings are an attempt at representing meaning.
There’s lots of other stuff you could represent about words. You could represent popularity by giving each word a number between 0 and 1 representing what portion of the training text was that word. That would tell you if a text might be harder for someone new to the language to understand. You could represent the spelling of the word as a list of integers. You could represent the pronunciation difficulty with a score based on ambiguous and infrequent n-grams. Or, you could try to represent the concept that the word is pointing to so you can follow synonyms or categorize text by topic. This would be an attempt at representing the word’s meaning.
I’ve been explaining as if you didn’t understand what this was, but if you disagree that’s another thing.
From the conversation so far it's clear to me that the OP has at least some idea of how word embeddings work and they described as "word salad" the expression that related them to "meaning", i.e. they quibbled about the "... representation of meaning", not the "high dimensional representation... " part.
I hope the OP can give a reference to the Quine and Kripke source they refer to.
As far as I remember it, the point about arithmetic with word verctors (man, king, woman, queen and all that) is a motivational example for Word2Vec by Tomas Mikolov, which wikpedia tells me was published in 2013. I can't find the reference right now, but in 2014 I was taking a Master's in all that jazz and I remember one of my tutors having the paper on his desk, and commenting that the claim that word embeddings model meaning was "well, that's what he says". So there has been debate and disagreement about the ability of word embeddings to represent meaning for at least as long as Word2Vec has existed, and as far as I know the idea of word embeddings existed earlier than that (I was taught the concept unconnected to word2vec in my Master's).
The bottom line is, just because somebody says their algorithm does a thing, doesn't mean that everyone has to accept it immediately, and without critical discussion.
So I, for one, want to see that Quine and Kripke stuff the OP is referring to.
That critical discussion was 10 years ago. Embeddings are vector representations of meaning. They’re not perfect, but they can be used for enough meaning-dependent tasks that this statement shouldn’t be controversial.
All this AI/ML stuff is so exhausting because the philosophers come out of the woodwork. They don’t prove anything or disprove anything, they just kinda whine.
When I was a flight instructor and said the airplane wanted to float down the runway, no one tried to point out that airplanes don’t have desires. I can say a server knows another server is offline and that’s fine too. But say “reason”, “think”, or “meaning” with something AI related and it’s like all work has to stop until we stop using that precious word.
I’ve tried to engage with that stuff less, but got tricked cause I thought this person just literally didn’t understand.
The unwillingness to admit to AI capabilities because they are expressed in terms of human capabilities is what I deem an argument that "boats can't swim".
True, in that what it does is different, but irrelevant in almost every facet that most people discussing it would care about. The boat, if anything, "swims" better than humans in basically every way. The argument against its use is only that swimming doesn't refer to boats.
(I admit I might have picked this argument and phrase somewhere, but I did recently a cursory glance back through old comments, and think I might have coined it based on having then recently learned that Russian uses the same term for swimming and sailing, which I had found interesting at the time. But perhaps not, my first use appears almost a decade ago, and my mind is ever forgetful)
And what I want to know is where is the code for "meaning". I don't care why it's the code for "meaning", but if you point me to the code for "word embeddings" and you say "that's the code for 'meaning", then I'm going to wonder whether, if I ask for the code for "quicksort" you'll point me to the code for "bubblesort".
So, no, what we call things that do things is important, otherwise we don't know what things do what, and what things are being done.
And lest we forget:
However, in AI, our programs to a great degree are problems rather than solutions. If a researcher tries to write an "understanding" program, it isn't because he has thought of a better way of implementing this well-understood task, but because he thinks he can come closer to writing the implementation. If he calls the main loop of his program "UNDERSTAND', he i s (until proven innocent) merely begging the question. He may mislead a lot of people, most prominently himself, and enrage a lot of others.
What he should do instead is refer to this main loop as "G0034", and see if he can convince himself or anyone else that G0034 implements some part of understanding. Or he could give i t a name that reveals its intrinsic properties, like NODE-NET- INTERSECTION-FINDER, it being the substance of his theory that finding intersections in networks of nodes constitutes understanding. If Quillian <1969> had called his program the "Teachable Language Node Net Intersection Finder", he would have saved us some reading. (Except for those of us fanatic about finding the part on teachability.)
https://cs.fit.edu/~kgallagher/Schtick/Serious/McDermott.AI....
Edit: And this is not right at all:
>> That critical discussion was 10 years ago.
In AI, there are lots of people who say lots of things. They keep saying them, even when people point out that they are just things that they say. They continue to say them even long after the other people get tired and give up trying to make sense of what is being said. That doesn't mean that the "discussion" is over, it's just that there is no real way to stop people saying things if that's what they really want to do.
Remember that the real claim is that embeddings represent meaning. In a flawed but working way. You could say they are a map, not the territory, so arguing about the true meaning of meaning isn't relevant.
> I'm going to wonder whether, if I ask for the code for "quicksort" you'll point me to the code for "bubblesort"
Embeddings are meaning (noun) like a dictionary, not meaning (verb) like a qualia.
The code sorts. It doesn't matter what algorithm because the claim was just that it gives you sorted data. Just like a claim that embeddings represent things is not a claim about methods or qualia.
The claim (that I’m making at least) is not that these systems can understand meaning. It’s that they can store a representation of meaning and do something with it. I think there’s a way lower bar for this. (I also think the bar was cleared quite a while ago and we should consider it settled.)
I could write down the meaning of a few words on some pieces of paper. (I guess if someone disagrees with that then talking about vectors is pointless). I could stuff each of those papers into a different dog toy and train a dog to get them on command. Then I think you could truthfully say something like “this dog can retrieve the meaning of the word ‘independence’, ‘hammer’ or ‘sing’ for you”. It wouldn’t be true to say the dog understands any of those words.
>> Maybe the thing that will land is to say that embeddings are an attempt at representing meaning.
And I don't disagree with that, neither with what you say in this comment.
Or at least, I think I agree. What I would have said is that you can write down words on a paper, and the words are a representation of meaning, but the paper doesn't understand meaning, nor do the words. You need a human, with a human understanding of words, and language, and what those things mean, in order to decode the meaning from the words. In other ah words, the words on the paper are a representation of meaning, but only for a human. For a cat, say, they don't represent anything.
Wat's more, just because a language model is trained on word embeddings, doesn't mean it encodes meaning. It's certainly an attempt to do that, but just because someone made an attempt doesn't mean they've done it.
If you change your perspective on language to mean any arbitrary capture of useful information (such that it can be used in the future), then you can see that the boundary between words and the world is not the heart of the issue. For example, your perception of the world works in a similar manner in that your sensory organs cannot comprehend the world, almost like your sensory organs interpret the world using their own language. Or maybe if that example is not very intuitive, then how about imagining an alien species that has sensory organs that act on linguistic structures. In some way, the aliens will figure out a coherent structure of their own, even though they cannot "experience" the world through senses like ours. What "intelligence" is doesn't seem to be bound by how "close" someone is to reality, and "close" might not even be the right word, since different perceptions can have different capabilities. I think the tricky part about "intelligence" is that there is always some "meaning" captured, it is just alien to those who do not share the same interpretive capacity. A cat could extract meaningful information from words on a paper, but certainly not in the same way we do.
Now if we want to make an AI that acts and thinks like us, then understanding our own machinery (the relationship between the world and language) is certainly important. But I think the bitter lesson rears its head here, and I believe that something that is truly worthy of being called AGI will be able to thrive given any set of arbitrary senses, even linguistic ones. In other words, I do not think embodiment will naturally lead to AGI, rather that embodiment is a necessity if we want make AI in our image. And making an AI in our image is the fastest way to get an AI to do useful work for us.
I don't think that it is. What I think is that language encodes meaning, that can be decoded only by an entity that knows how to decode meaning from language; which is a bit of a tautology, but that's the point, you can't expect an entity without the ability to decode meaning from language to understand what language means. Or at least I don't expect that.
Here's an analogy, and I shouldn't be making it because it's about cryptography and I'm not an expert there. Suppose I sent you an encrypted message and you knew how to decrypt it- I encoded it with your public key and you decrypted it with your private key, or whatever. If you could do that, then you could read my message and know what it says.
If you couldn't decrypt my message, then you could stare at it for as long as you liked, you could make copies of it, you could make variations of it, you could even learn a model of the structure of encrypted messages like mine, and be able to produce many more of those messages, but you would still not know what those encrypted messages say. Because they're encrypted and you can't read them, you can only read their encoding.
That's what I'm driving at. I think that's how language works, in practice. Not that it's some form of encryption, but that it only makes sense to humans, so only humans can decode meaning from language. With language modelling, we're reproducing the encoded message, but we haven't yet found out how to equip the machines modelling language with the ability to decode meaning from it, so for all intents and purposes it might just as well be encrypted.
I'm not arguing for embodiment, either. I'm perfectly fine with the idea of a "brain in a jar". But I agree that if we want our AI's to behave like humans, they will have to have some experience in the world.
Like when you encrypt an image in ECB mode: https://upload.wikimedia.org/wikipedia/commons/c/c0/Tux_ECB....
I think that no, because you can manipulate the structure of an encrypted string and swap around pieces of it etc, without having to decrypt it.
As to making convincing messages, as you say, the entity making the messages and swapping around the data is a language model, but the entity reading the generated messages is a human. It is the human that finds the messages "convincing" as you say. We have no doubts that humans can decode meaning from text, even if we have no idea how we do it, yet. The question is whether language models can do the same thing. And what I say above is that there is no indication that they can, because all they do can be done by language generation, without any decoding of meaning needed.
> You need a human, with a human understanding of words, and language, and what those things mean, in order to decode the meaning from the words.
By this definition I think you’ll always have doubt about these systems.
You can use (transformer) LLMs to get a "word sense" aka its predicted probability distribution, which also induces a network of associations (not via context, but via whatever the LLM has learned as fitting words). This gives a structural definition of meaning the ye olde sense (v. Humboldt and so forth).
This of course starts with the assumption that LLMs can encode meaning, albeit not in a simplistic linear fashion as Word2Vec or GloVe did
The meaning is purely extrinsic. There's nothing about a ball that makes it named that, as shown by the fact that a ball is called other things in other languages.
A word without someone ascribing it meaning means nothing.
A word has meaning because we give it to it, and we give it that meaning based on context.
Some of that context is remote in time and space (we learned the word ball a long time ago), some is near (that it occurs near "back yard" and "round rubbery thing" makes it more likely it's that thing we use to play games rather than a testicle or a party where people dance), but it is wholly dependent on context.
Put it alongside words in a different language, for example, and it might have a whole other set of potential meanings.
The objection I though you were making, with the "ball" etc, is that all these words that sit next to other words according to what the words mean, must mean something in isolation. For example, even though "ball" can be used in different contexts, it can only mean so many things, and it would never mean, say, "a cooking utensil made of aluminum where I fry my eggs each morning" no matter what other words you put it next to.
And, as Young et al (1976) have demonstrated, it is possible to move words out of their expected context for great fun and profit, but it's not clear that a word can always be moved in any context, and still make sense. I would even go so far as to say that can probably not be done, at least not without a drastic reconfiguration of all of English (where "ball" is used).
So the question is not only "what other words do we find the word 'ball' close to" but also "why do we find 'ball' next to those other words?". The latter question can't be answered just by looking at what words hang out with what other words, unless we already know what words mean on their own.
And that is, at least for me, the objection about LLMs "meaning" and "understanding" anything. No matter what text an LLM is generating, the entity that is decoding the meaning of the generating text is always a human. The LLM can't do that on its own. Because there is no mechanism that it is equipped with that could ever do that.
P.S. "words hanging out with other words" are what are known as "token collocations" in technical jargon. They are an idea as old as linguistics itself, a basal concept whose modern implementations in NLP systems are just, well, modern implementations. We still don't know how humans decode meaning from token collocations even though it's such an ancient idea. And it might even be wrong, at this point.
_________________
Bibliography
1. Angus Young, Malcolm Young, and Bon Scott (1976) Big Balls. In Harry Vanda, George Young (prod), Dirty Deeds Done Dirt Cheap, Albert productions - Atlantic Records, Sydney, Australia, side 2, track 2. URL: https://youtu.be/4WwJ6OVSwkM
The fact that we can have conversations with GPT where these models are able to correctly e.g. write working code and symbolically evaluate code is sufficient proof of that assertion that it is possible to figure this out that I don't think there is any reasonable basis for debating this any more. Claiming it's not possible is just demonstrably wrong.
> And, as Young et al (1976) have demonstrated, it is possible to move words out of their expected context for great fun and profit, but it's not clear that a word can always be moved in any context, and still make sense. I would even go so far as to say that can probably not be done, at least not without a drastic reconfiguration of all of English (where "ball" is used).
I don't see what this has to do with anything. Yes, we can shift around, alter their meanings, use them in unusual contexts and still make sense of it, and yes of course they won't always make sense in every context. How is that relevant?
> The objection I though you were making, with the "ball" etc, is that all these words that sit next to other words according to what the words mean, must mean something in isolation. For example, even though "ball" can be used in different contexts, it can only mean so many things, and it would never mean, say, "a cooking utensil made of aluminum where I fry my eggs each morning" no matter what other words you put it next to.
If I use it to mean "a cooking utensil", then it means a cooking utensil when I communicate with others that have that shared context. It has no meaning separate from context.
> So the question is not only "what other words do we find the word 'ball' close to" but also "why do we find 'ball' next to those other words?". The latter question can't be answered just by looking at what words hang out with what other words, unless we already know what words mean on their own.
"What words mean on their own" makes no sense. No word has a meaning separate from context. Without context there's nothing to assign them meaning. And that context is interactions. For GPT purely words. For us, words and other sensory input. In either case, without context assigning attributes to the word "ball" it is nothing more than a meaningless a sequence of letters. What does "JMw3rfd" mean? What's its intrinsic meaning? Until I tell you that this is my new word for "ball" it has none. Once I do, it has an extrinsic meaning. It's one that isn't widespread and will soon be forgotten, but the meaning exists the moment we label it, and only once we label it, and it is purely extrinsic.
> And that is, at least for me, the objection about LLMs "meaning" and "understanding" anything. No matter what text an LLM is generating, the entity that is decoding the meaning of the generating text is always a human. The LLM can't do that on its own. Because there is no mechanism that it is equipped with that could ever do that.
We don't know enough about human reasoning, and so how close how LLMs are to how human reasoning works to be able to even begin to determine whether this is true or false.
Then there's no reason to continue this conversation.
I think so too, but how is that intrinsic? It’s still contextual.
I think constricting it to text only is a mistake. Tokens can represent anything.