>"neurons"
This is literally the accepted term for the nodes wherein sums are produced of the set of weighted products which connect such nodes together. If you know a better term, or would prefer a better term that "neuron" to describe these points in a "neural network", by all means speak your mind, but don't pretend some offense at it.
>and that's because you restricted yourself to facts.
I restricted myself to writing from a point of view concordant with your own because that is what one does when explaining a position well, not because I agree with your point of view.
>You could not give a factual description with the word "meaning" in it (in the sense of "the meaning of...").
If we're going to play a game of semantics, than so be it.
If your argument is simply that symbols and numbers can't contain meaning because they are symbols and numbers, then you've worked your way into a tautological box that I, nor anyone else, will be able to free you from.
It is known that memories in biological neuron-based networks are not locally stored or unique to some neuron ( there is no "grandma" neuron in your head to keep track of dear old grannie ), but that instead memories are spread across the set of neurons there available.
This sort of representation in known as a "holographic" representation, which I expect our neural net will also utilize. Information will not be stored in any specific neurons in the trained model, but will be spread across the whole of them.
But even if true, that is merely data, you will no doubt point out. Where is the "meaning"?
As you want to delve philosophically into it, we must question the very nature of the word meaning here. Our model lacks any form of experiential qualia, and so cannot associate words with experienced reality, which is the normal basis for human concepts of meaning, which then underlie our development out of concrete language and into metaphoric.
This leaves meaning within the model as only being the derived relationship between words.
Words, on their own and bereft of qualia, can present our model with structure, grammar, punctuation. Words will have certain uses within these structures and grammars, they will have certain places. Words will modify words, phrases will modify larger sentences. Trues can be asserted and falsehoods. Truth and falseness itself can be demonstrated through sentences making claims about other sentences.
Would the model know what a red ball is? Of course not. But it would be aware of how the term ball is used, what sort of type it is given, the qualifications and usages of words of that type. The modifiers it is reasonable to use on it, etc.
And incredibly, with no basis in anything other than being slightly corrected for each incorrect token generated, our training harness manages to discover a mathematical neural net model of these relationships between words such that it can manipulate them as well or better than many people.
The use of backpropagation and gradient descent allows us to search the space of all possible models starting from a randomized model, and find a model that does this by nudging our model to be a little more capable at it over and over again.
However, English is not just a fixed mathematical set of symbols that can be rotely manipulated.
Our language is metaphoric from its very foundations. Our hearts soar, sink and swell in ways that have nothing to do with the reality, but only our perceptions of ourselves.
To manipulate English as it does, the traversal of the model-space must advance inexorably on a model capable of handling such non-literal language.
The model contains meanings for car related actions, for curling related actions, it contains the ability to translate between them when the only connection is a meaning that is implied by the words but not either of their literal meanings.
What is a meaning here? It is a relationship between the words. The relationship between the car words and curling words only exists as the relationship to a common implied meaning that can be transformed under the context of other words.
Where is the meaning? We don't even know for certain where the meaning is inside ourselves. I expect it to be similarly holographically stored across the set of neurons and provide the appropriate context to twist the swirl of mathematics that occurs between input vector and output.
>So I still do not know why you think all this is going on inside a language model.
Because it is the simplest explanation of what it is doing. Call it "latent spaces", "embeddings" and the assume the compression I discuss is via "manifold hypothesis".
>What do language models have to do with meaning, and why do we absolutely need to explain their behaviour by talking about meaning? Where is this information in your comment? I cannot see it.
So, where is the meaning? It is everywhere in the model once trained. It creates a holographically stored mapping of reality as presented through the corpus of human thought in the form of writing based solely on being iteratively nudged through model-space.
An amazing feat by the engineers that dreamed it up.
If you are unwilling to call this mapping model of reality built only of tokens/symbols/words "meaning", then I have no argument for you, except to again conclude that you've locked yourself in a tautological box.
If your argument is merely "ha ha, you're merely conjecturing this", then yes, of course I am. Why? Because it behooves someone to hypothesize, test, and study things which interest them.
Do you have some competing idea of how it works? It would behoove you to share it with the thread.
As it is, I only see you making demands for proof of things that you are well aware that no human has yet completely modeled, even after decades of studying such phenomena. What you gain from this incessant nay-saying, I do not understand.