The slightly more complicated one is that a compressed text dump of Wikipedia isn't even 70GB, and this is lossy compression of the internet.
The slightly more complicated one is that a compressed text dump of Wikipedia isn't even 70GB, and this is lossy compression of the internet.
ie: given "just wikipedia" what's the best score people can get on however these models are evaluated.
I know that all the commercial ventures have a voracious data-input set, but it seems like there's room for dictionary.llm + wikipedia.llm + linux-kernel.llm and some sort of judging / bake-off for their different performance capabilities.
Or does the training truly _NEED_ every book every written + the entire internet + all knowledge ever known by mankind to have an effective outcome?
I have the same question.
Peter Norvig’s GOFAI Shakespeare generator example[1] (which is not an LLM) gets impressive results with little input data to go on. Does the leap to LLM preclude that kind of small input approach?
[1] link should be here because I assumed as I wrote the above that I would just turn it up with a quick google. Alas t’was not to be. Take my word for it, somewhere on t’internet is an excellent write up by Peter Norvig on LLM vs GOFAI (good old fashioned artificial intelligence)
the 60-70B parameters of models is basically like... just stored patterns of "if these 10 tokens in a row input, then these 10 tokens in a row output score the highest"
Is that a good summary?
> The model uses its learned statistical patterns to predict the probability of what comes next in a sequence of text.
based on what inputs?
1. previous tokens in the sequence from immediate context
2. tokens summarizing the overall topic/subject matter from the extended context
3. scoring of learned patterns from training
4. what else?
I think it's more likely that LLMs are actually learning and understanding concepts as well as memorizing useful facts, than that LLMs have discovered a compression method with that high of a compression ratio, haha.
LLMs are just completing patterns of text that have been given before, 'everthing ever written' is both a lot for any individual person to read; but also, almost nothing, in that to propertly describe a table requires more information
text is itself an extremely compressed medium which lacks almost any information about the world; it succeeds in being useful to generate because we have that information and are able to map it back to it
I should make it clear that my comparison there is unfair and mostly just funny – you don't need to store every possible combination of 10 tokens, because most of them will be nonsense, so you wouldn't actually need that much storage. That being said, it's been fairly solidly proven that LLMs aren't just lookup tables/stochastic parrots.
Well i'd strongly disagree. I see no evidence of this; I'm am quite well acquainted with the literature.
All empirical statistical AI is just a means of approximating an empirical distribution. The problem with NLP is that there is no empirical function from text tokens to meanings; just as there is no function from sets of 2D images to a 3D structure.
We know before we start that the distributions of text tokens are only coincidentally related to the distributions of meanings. The question is just how much value that coincidence has in any given task.
(Consider, eg., that if I ask, "do you like what i'm wearing?" there is no distribution of responses which is correct. I do not want you to say "yes" 99/100, or even 100/100 times. etc. what I want you to say is a word caused a mental state you have: that of (dis)liking what i'm wearing.
Since no statistical AI systems generate outputs based on causal features of reality, we know a priori that almost all possible questions that can be asked cannot be answered by LLMs.
They are only useful where questions have cannonical answers; and only because "cannonical" means that a text->text function is likely to be conidentally indistinguishable from a the meaning->meaning function we're interested in).
I’m not saying you believe that, but I fail to see how that situation is structurally different from what you claim. If it’s a matter of degree, how do you feel things change as the situation becomes more complex?
But, to be clear, the reason it could ever work at all has nothing to do with the methods or the data itself, it has to do with the properties of the data generating process (ie., reality, ie., what's being measured).
You can never build representations from measurement data, this is called inductivism and it's pretty clearly false: no representation is obtained from just characterising measurement data. Theres no cases where I can think of that this would work -- temperature isnt patterns in thermometers; gravity isnt patterns in the positions of stars; and so on.
Rather you can decide between competing representations using stats in a few special cases. Stats never uncovers hidden representations, it can decide between different formal models which include such representations.
eg., if you characterise some system as having a power-law data generating process (eg., social network friendships), then you can measure some parameters of that process
or, eg., if you arrange all the data to already follow a law you know (eg., F=Gmm/r^2) then you can find G, 'statistically'.
This has caused a lot of confusion histroically: it seems G is 'induced over cases', but all the representaiton work has alerady been done. Stats/induction just plays the role of fine-tuning known representatios. it never builds any
If there were such a thing, it'd be interesting to propose how our own minds, at least to the degree that they can be seen as statistical learners in their own right, achieve semantics. And how that thing, whatever it might be, is not itself a learned representation driven by statistical impression.
We move, and in moving, grow representations in our bodies. These representations are abstracted in cognition, and form the basis for abductive explanations of reality.
We leave plato's cave by building vases of our own, inside the cave, and comparing them to shadows. We do not draw outlines around the shadows.
This is all non-experimental 'empirical' statistics is: pencil marks on the cave wall.
If someone else crafted an experiment, and you were informed of it and then shown the results, if this was done repeatedly enough, would you be incapable of forming any sort of semantic meaning?
The meaning of the measures is determined by the experiment, not by the data. "Data" is itself meaningless, and statistics on data is only informative of reality because of how the experimenter creates the measurement-target relationship.
There’s no doubt in my mind that experimental learning is more efficient. Especially if you can design the experiments against your personal models at the time.
At the same time, it’s not clear to me that one could not gain similar value purely by, say, reading scientific journals. Or observing videos of the experiments.
At some point the prevalence of “natural experiments” becomes too low for new discover through. We weren’t going to accidentally discover an LHC hanging around. We needed giant telescopes to find examples of natural cosmological experiments. Without a doubt, thoughtful investment in experimentation becomes necessary as you push your knowledge frontier forward.
But within a realm where tons of experimental data is just available? Seems very likely that a learner asked to predict new experimental results outside of things they’ve directly observed but well within the space of models they’ve observed lots of experimentation around should still find that purely as an act of compression, their statistical knowledge would predict something equivalent to the semantic theory underlying it.
We even seemed to observe just this in multimodal GPT-4 where it can theorize about the immediate consequences of novel physical situations depicted in images. I find it to be weak but surprising evidence of this behavior.
You are correct to observe that science, as we know it, is ending. We're way along the sigmoid of what can be known, and soon enough, will be drifting back into medieval heuristics ("this weed seems to treat this disease").
This isnt a matter of efficiency, it's a necessity. Reality is under-determined by measurement; to find out what it is like, we have to have many independent measures whose causal relationship to reality is one we can control (through direct, embodied, action).
If we only have observational measures, we're trapped in a madhouse.
Let's not mistake science for pseudoscience, even if the future is largely now, pseudoscientific trash.
it is always trivial to take one of these models and expose it's failure to operate semantically, but these cases are never in the marketing material.
Consider an associative model of addition, all numbers from -1bn to 1bn, broken down into their digits, so that 1bn = <1, 0, 0, 0, 0, 0, 0, 0, 0>
Using such a model you can get the right answers for more additions than just -1bn to 1bn, but you can also easily find cases where the addition would fail.
It's never adding.
On the other hand you can look at statistical model identification in, say, nonlinear control. This can absolutely lead to unboundedly long predictions.
This is like saying a brain operates wholly on electrochemical states and knowns nothing about meaning in the real world, though; the mechanistic description is accurate, the cognitive conclusion attached to it is, at best, based on unsupported conjecture about the relation of mechanism to understanding.
Modern LLMs are able to transfer knowledge between different languages, so it's fair to assume that some mapping between human language and a more abstract internal representation happens at the input and output, instead of the model "operating" on English or Chinese or whatever language you talk with it. And once this exists, an internal "world model" (as in: a collection of facts and implications) isn't far, and seems to indeed be something most LLMs do. The reasoning on top of that world model is still very spotty though
> Is that a good summary?
No - there's a lot more going on. It's not just mapping input patterns to output patterns.
A good starting point to understand it are linguist's sentence-structure trees (and these were the inspiration for the "transformer" design of these LLMs).
https://www.nltk.org/book/ch08.html
Note how there are multiple levels of nodes/branches to these trees, from the top node representing the sentence as a whole, to the words themselves which are all the way at the bottom.
An LLM like ChatGPT is made out of multiple layers (e.g. 96 layers for GPT-3) of transformer blocks, stacked on top of each other. When you feed an input sentence into an LLM, the sentence will first be turned into a sequence of token embeddings, then passed through each of these 96 layers in turn, each of which changes ("transforms") it a little bit, until it comes out the top of the stack as the predicted output sentence (or something that can be decoded into the output sentence). We only use the last word of the output sentence which is the "next word" it has predicted.
You can think of these 96 transformer layers as a bit like the levels in one of those linguistic sentence-structure trees. At the bottom level/layer are the words themselves, and at each successive higher level/layer are higher-and-higher level representations of the sentence structure.
In order to understand this a little better, you need to understand what these token "embeddings" are, which is the form in which the sentence is passed through, and transformed by, these stacked transformer layers.
To keep it simple, think of a token as a word, and say the model has a vocabulary of 32,000 words. You might perhaps expect that each word is represented by a number in the range 1-32000, but that is not the way it works! Instead, each word is mapped (aka "embedded") to a point in a high dimensional space (e.g. 4096-D for LLaMA 7B), meaning that it is represented by a vector of 4096 numbers (cf a point in 3-D space represented as (x,y,z)).
These 4096 element "embeddings" are what actually pass thru the LLM and get transformed by it. Having so many dimensions gives the LLM a huge space in which it can represent a very rich variety of concepts, not just words. At the first layer of the transformer stack these embeddings do just represent words, the same as the nodes do at the bottom layer of the sentence-structure tree, but more information is gradually added to the embeddings by each layer, augmenting and transforming what they mean. For example, maybe the first transformer layer adds "part of speech" information so that each embedded word is now also tagged as a noun or verb, etc. At the next layer up, the words comprising a noun phase or verb phrase may get additionally tagged as such, and so-on as each transformer layer adds more information.
This just gives a flavor of what is happening, but basically by the time the sentence has reached the top layer of the transformer it has been able to see the entire tree structure of the sentence, and only then have "understand" it well enough to predict a grammatically and semantically "correct" continuation from which it is able to predict continuation words.
Since unicode has well over 64000 symbols, does that imply models, trained on a large corpus, must necessarily have at least 64000 ‘branches’ at the bottom layer?
The linguistic sentence structure tree for any input sentence is a useful way to think about what is happening as the input sentence is fed into the model and processed through it layer by layer, but doesn't have any direct correspondence to the model. The model has a fixed number of layers of fixed max-tokens width, so nothing changes according to the sentence passing through it.
Note that the bottom level of the sentence structure tree is just the words of the sentence, so the number of branches is just the length of the sentence. The model doesn't actually represent these branches though - just the embeddings corresponding to the input, which are transformed from input to output as they are passed through the model and each layer does it's transformer thing.