Pretraining data enables narrow selection capabilities in transformer models
arxiv.org
arxiv.org
This paper uses GPT-2 transformer scale, on sinusoidal data:
>We trained a decoder-only Transformer [7] model of GPT-2 scale implemented in the Jax based machine learning framework, Pax4 with 12 layers, 8 attention heads, and a 256-dimensional embedding space (9.5M parameters) as our base configuration [4].
> Building on previous work, we investigate this question in a controlled setting, where we study transformer models trained on sequences of (x,f(x)) pairs rather than natural language.
Nowhere near definitive or conclusive.
Not sure why this is news outside of the Twitter-techno-pseudo-academic-influencer bubble.
Learning in High Dimension Always Amounts to Extrapolation
https://arxiv.org/abs/2110.09485
What you're asking for is not "generalization" but magic and humans would also fail.
What you might be asking for is a system that simply continually learns.
Because I'm not convinced humans can do this.
Or that it reasonably means anything.
I read an interesting paper recently that had a great take on this: If you add enough data, nothing is outside training data. Thus solving the generalization problem.
Wasn’t the main point of that paper, but it made me go ”Huh yeah … I guess … technically correct?”. It raises an interesting thought that yes if you just train your neural network on everything, then nothing falls outside its domain. Problem solved … now if only compute was cheap.
Having said that, the real question is what percentage of the learned representations do generalize. For a perfect model, it would learn only representations that generalize and none that overfit. But, that's unreasonable to expect for a machine *and* even for a human.
Maybe we just don't know. We are staring at a black box and doing some statistical tests, but actually don't know whether the current AI architecture is capable enough to get to some kind of human intelligence equivalent.
The paper is making the rounds despite being a weak result because it confirms what people want, for non-technical reasons, to be true. You see this kind of thing all the time in other fields: for decades, the media has elevated p-hacked psychology studies on three undergrads into the canon of pop psychology because these studies provide a fig leaf of objective backing for pre-determined conclusions
They haven't read the other papers either. It's really striking to me to watch people retweet this and it get written up in pseudo-media like Business Insider when other meta-learning papers on the distributional hypothesis of inducing meta-learning & generalization, which are at least as relevant, can't even make a peep on specialized research subreddits - like, "Pretraining task diversity and the emergence of non-Bayesian in-context learning for regression", Raventós et al 2023 https://arxiv.org/abs/2306.15063 (or https://arxiv.org/abs/2310.08391 ) both explains & obsoletes OP, and it was published months before! OP is a highly limited result which doesn't actually show anything that you wouldn't expect on ordinary Bayesian meta-reinforcement-learning grounds, but there's so much appetite for someone claiming that this time, for real, DL will 'hit the wall' that any random paper appears to be definitive to critics.
1. The test mechanism is to use prediction of sinusoidal series. While it's certainly possible to train transformers on mathematical functions, it's not clear why findings from a model trained on sinusoidal functions would generalize into the domain of written human language (which is ironic, given the paper's topic).
2. Even if it were true that these models don't generalize beyond their training, large LLMs' training corpus is basically all of written human knowledge. So then the goalpost has been moved to "well, they won't push the frontier of human knowledge forward," which seems to be a much diminished claim, since the vast majority of humans are also not pushing the frontier of human knowledge forward and instead use existing human knowledge to accomplish their daily goals.
LLMs are trained on a tiny subset of the written human knowledge which we've proven to probably not be garbage and which is nicely formatted in simple text formats and which was published without too many paywalls on the web and on and on and on. It's a lot, and it definitely includes enough facts that no one person knows all the things the LLM "knows", but the average child knows plenty of things which never made it into that sort of a corpus. Yes, it's probably true that the vast majority of humans are also not pushing the frontier of human knowledge forward, but the vast majority of humans are working with a slightly different (partially overlapping) set of information than what the LLMs see.
An LLM right now is still just a facsimile of one of human's cognition. Arguably, without the other senses, it can't experience or understand things the same way we can. This leads to the AI unable to form a comparable sentience to a child, at least, not in the same way and not in a way that would be recognized by most people.
Someone already gave you the sense of time as an example. These senses are things we were born with and critical to our perception of the surroundings. Until the AI is trained on at least the majority of these senses, a child will always have something they can do that the AI can't.
But every single day I am using OpenAI GPT4 to handle novel tasks. I am working on a traditional saas vertical, except with a pure chatbot. The model works, is able to understand which function to call, to extract which parameters, and to know when the inputs will not work. Sure, if you ask it to do some extraneous task, it fails.
Google/Deep Mind need to start showing up with some working results.
Where. are. the. models. google.
That said, even GPT4 certainly has pretty significant limitations on what it manages to reason about. But without comparing their capabilities in other aspects, arguably so do most humans. We tend to force our way past those limitations by learning incrementally by doing over and over. Current models don't get that luxury without complicated fine-tuning steps, so if anything what should surprise us is how well they do with the limitation of only context to act as short-term memory.
This is only relevant for the AI-bro fantasy of AI becoming exponentially smarter than humans.
(And btw., if that is what you consider the "AI-bro fantasy" then I'm firmly in that camp. There is no logical reason to assume that's not possible; short of identifying a non-deterministic, non-materialistic source of intelligence in human brains that violates known physics and that we can't emulate, it's just a question of when, not if, and the reason for not going full tilt on that right now is down to a variant of the Wait Calculation, not a need for intelligence to kick it off - that is, trying to kickstart that now is likely to be more expensive and not be any faster than trying to kickstart it next year, and so on, and the question is guessing at when that stops being true)
Another important thing to keep in mind is one paper(wish I could remember which one it was) that showed even larger scale llms have trouble understanding that A=B is same as B=A if they have not seen A or B before
So - and I say this as someone who writes NLP papers too - who cares?
https://arxiv.org/abs/2110.09485
"Supercharged Interpolation" is not a real thing.
They generalize fine when the data incentivizes that.
If you trained it on one function class, of course that's all it learned to do. That's all it ever saw!
If you want to learn arbitrary function classes to some degree, the solution is simple. Train it on many different function classes.
Untrained models are as blank slate as you could possibly imagine. They're not even comparable to new born humans with millions of years of evolution baked in. The data you feed them is their world. Their only world.
---
Here is a grammar for a made up language:
sentence = "sob" subject (verb "ob" object | verb "ad" sentence | "ploj").
verb = "ab" | "pom"
object = "foo" | "bar"
Construct a few sentences, and show the sentence so far, and which rule and step of the rule you are currently applying. E.g. sentence= ""; rule=sentence; step="sob".---
Notice I did not explain the BNF syntax, which is reasonable as it'll have plenty of that in its training data. Notice that I also "forgot" to give it a production for the "subject" rule - leaving it with two choices: It can give up and tell me off, or it can do as best it can and let me know it's had to improvise.
Here's the first sentence construction:
---
First Sentence Construction:
sentence = "";
Rule: sentence; Step: "sob". sentence = "sob";
Rule: subject; Step: (choosing object as a subject for simplicity) "foo". sentence = "sob foo";
Rule: verb; Step: "ab". sentence = "sob foo ab";
Rule: (verb "ob" object | verb "ad" sentence | "ploj"); Step: "ob". sentence = "sob foo ab ob";
Rule: object; Step: "bar". sentence = "sob foo ab ob bar"; (Complete Sentence)
---
It has correctly applied the rules I did give, and opted to "fill in" "object" as a replacement for my missing "subject" rule. There are few enough rules this could be a fluke, but it's a better choice to guess at that substitution than verb or sentence, though hardly takes a genius.
Second:
---
Second Sentence Construction:
sentence = "";
Rule: sentence; Step: "sob". sentence = "sob";
Rule: subject; Step: "bar". sentence = "sob bar";
Rule: verb; Step: "pom". sentence = "sob bar pom";
Rule: (verb "ob" object | verb "ad" sentence | "ploj"); Step: "ad". sentence = "sob bar pom ad";
Rule: sentence; Step: "sob" (this starts a nested sentence). sentence = "sob bar pom ad sob";
Rule: subject; Step: "foo". sentence = "sob bar pom ad sob foo";
Rule: verb; Step: "ab". sentence = "sob bar pom ad sob foo ab";
Rule: (verb "ob" object | verb "ad" sentence | "ploj"); Step: "ob". sentence = "sob bar pom ad sob foo ab ob";
Rule: object; Step: "foo". sentence = "sob bar pom ad sob foo ab ob foo"; (Complete Sentence)
---
This is also correct, and it correctly points out that the arm it picked triggered a nested application of the `sentence` rule.
It stumbled on the third attempt:
---
Third Sentence Construction:
sentence = "";
Rule: sentence; Step: "sob". sentence = "sob";
Rule: subject; Step: "foo". sentence = "sob foo";
Rule: verb; Step: "pom". sentence = "sob foo pom";
Rule: (verb "ob" object | verb "ad" sentence | "ploj"); Step: "ploj". sentence = "sob foo pom ploj"; (Complete Sentence)
---
Note that there is a possible interpretation of the "|" operator which would make its sentence valid, so "technically" it could be, but it would involve (ver "ob" object | verb "ad" sentence | "ploj") to be interpreted as (verb (("ob object)|("ad" sentence)|"ploj")) and that's contrived.
So to answer you: It can, and it does, but not always consistently. One can choose to be disappointed it's not perfect, but for just handing it a grammar and asking it to reason about it with no example text or chances to "practice" whatsoever, I think it's doing quite well.
I've interviewed many software developers who'd struggle with a task like this.
When prodded about the third example, it picks a more reasonable interpretation of the "|", though it's explanation starts out with an error, but then given a correct explanation:
---
Q: How do you figure this? The full rule is "sob" subject (verb "ob" object | verb "ad" sentence | "ploj"). How do you figure the verb is allowed without being followed by either "ad" or "ob"?
A: Upon closer examination, you are correct to point out that the structure following "sob" subject must indeed involve a verb according to the rule provided. The full rule:
"sob" subject (verb "ob" object | verb "ad" sentence | "ploj"). indicates that the sentence must start with "sob", followed by a subject, and then either:
a verb followed by "ob" and an object,
a verb followed by "ad" and another sentence,
or the standalone term "ploj".
---
It goes on to give two correct alternatives to the erroneous one:
---
* "sob foo ploj" — where "ploj" follows directly after the subject without a verb, or * "sob foo ab ob bar" — where "ab" is the verb followed by "ob" and "bar" as the object.
---
While it's not perfect, it can very clearly "learn and reproduce multistep algorithms" to some extent. It also clearly needs to get better at it.
I personally think the more theory trained folks wonder if there will turn out to be some analogy with the corresponding towers of languages and machines recognizing them (regular languages, finite automata, push down automata, turing machines, for examples).
is an example covering the historical language classes and recognizers for them.
Perhaps LLMs and various architectures will turn out fitting what you describe.
Or, you can simply ask it to use a tool, like a calculator for you. This is more reliable than fine-tuning a current technique BPE tokenized model ever will be.
Largely two things come into play: 1) Some part of the neural net is emulating more traditional logic, but it may not always be the most activated part or tuned to be perfect in the answers 2) There isn't really a "jmp" equivalent in a single iteration, so the neural net has to learn to not only do decimal division but do it based on iterating output tokens, continuing perfectly each token output, and choosing to put the right stuff into context and keep that context activated at the right time.
"Activated" in this case means, more or less, the group of neurons specializing on this task are being both fed and listened to.
You can even train a neural net to emulate a traditional addition circuit directly, it's just less efficient than one would think if you're trying to build a general purpose model instead of a specialized one.
https://gwern.net/scaling-hypothesis And specifically the part where it discusses addition under the heading Blessings of Scale.
If this is true, I suppose based on what we know of the curse of dimensionality it makes sense. Very thought-provoking work.
Meant to link this. More pertinent to my point on Tokenization https://arxiv.org/abs/2310.02989
A better answer: every time the model needs to predict the digit of a sum the model needs to solve the entire addition to know the carries, a bigger sum requires a bigger model to solve them.
With those structured numbers will the LLMs be 100% accurate on new prompts or will they just be better than chance (even significantly better than chance)?
Because this is one thing, it has to learn the structure and then create probabilities based on the data, but does that mean it's actually learning the underlying algorithm for addition for example or is it just getting better probabilities because of a narrowing of them? If it can indeed learn underlying algorithms like this that's super interesting. The reason also this is in an issue if it _can't_ learn those, you can never trust the answer unless you check it, but that's sort of a sidepoint.
As it happens, conditioning even people to actually apply rules properly tends to take a lot of repetition. How many individual examples of step-by-step working out basic math problems do you think they have been in their training data?
Prompt them the way you'd prompt a child who is sloppy into working step by step and explaining how it applies the rules it has learned, and it will tend to do better. While tokenization might not help, I don't think there's an inherent problem there beyond feeding them enough training data. Whether that's worthwhile vs. having it resort to tools, is another matter.
However, I'm not sure why being bad at math (if even true based on other comments here), is a legitimate criticism. We've already got a lot of machinery that is very good at math. So, use the tools that make sense for the domain. Or (what I'm really waiting for) better yet find novel ways to stitch the tools that we do have together. Can LLMs turn a problem statement into a matlab program? If so then it doesn't really matter how bad at math they are.
You can ask GPT-4 arbitrary arithmetic it could never see in it's training set. Even when it's not completely correct, it's extremely close. It is clearly computing algorithms even if those algorithms are not quite right.
Being able to infer answers to math problems is a thing that humans can do, and that's fine. It's not as good as doing the math though.
The question is whether it is computing the right ones.
>Asking a token prediction model to do math is like asking a human to do math without doing the math. What's 9 times 9? I can tell you it's 81 from sheer memorization. I can probably invoke 9 x 9 + 1 = 82 without needing to do any calculation either. But if you ask me 32 * 64, that's very difficult to do without doing calculations. Implicitly doing math is not sensible.
It's not about memorization. It's trivial to test on instances that would not appear in training and see GPT-4 be better than a human who would attempt the problem without a tool or pad.
The biggest problem with LLM arithmetic is tokenization. https://arxiv.org/abs/2310.02989
Other than that, the algorithms it uses for arithmetic will continue to converge during training until it is correct.
Humans doing things "without a pad" is not the same thing as doing mental math vs. trying to intuit an answer.
Token predictions will not converge on correct arithmetic algorithms.
>What's 9 times 9? I can tell you it's 81 from sheer memorization. I can probably invoke 9 x 9 + 1 = 82 without needing to do any calculation
Performing math how you suggest is little more than memorization except you've just memorized a different chain in the process.
>That is a close equivalent to trying to do with math token prediction.
This isn't going to become true no matter how much you repeat it. You've consistently made assertions that are trivially proven false.
>Token predictions will not converge on correct arithmetic algorithms.
Yes it will. It literally will. This isn't some debate. This is something that has been researched.
https://www.alignmentforum.org/posts/N6WM6hs7RQMKDhYjB/a-mec...
You think you have an understanding of Language Models and token prediction. Unfortunately you don't.
It's synthesis of memory perhaps, but it's not memorizing specific answers; which is different yet from executing steps of an algorithm you have memorized.
> This isn't going to become true no matter how much you repeat it. You've consistently made assertions that are trivially proven false.
Then agree to disagree because I think you're full of shit. actually, not even that, I think you just don't even understand what I'm saying. which is agreeing with you 90% of the way. This article is nice. But Memorization -> Generalization -> Cleaning up Noise -> Stability is STILL NOT THE ALGORITHM. It will never be the algorithm. It's always just "generalization-ish". Which is again, a modality that human brains can use to process problems and works pretty well, and can presumably work very well, but is ultimately inferior to established, perfected algorithms.
but nah. you chose petty arrogance so yay, we get get to be dicks. fuck off.
Since the thing is a computer, why can’t it answer queries by writing and executing a program? Does a transformer-based AI have to be 100% transformer?
But it's no longer a token prediction model.
I haven’t tried this, but I think I could describe a very basic logical CPU to it and have it execute opcodes and tell me the CPU state at each step (like Knuth’s MIX). Couldn’t I ask it to write programs and execute programs for this computer when doing so would help it give better answers?
But even if it did, you are basically describing programming in a contrived language, not the llm learning to do math. You can explicitly teach a model to follow an algorithm. But that's not the same as the algorithm being "learned" during training and "invoked" when asked what's 2 + 2?
And GPT4 indeed does just fine with things like 32 * 64 in multiple different ways that humans can also easily memorise the rules for. When I asked it to calculate it step by step, and use easily memorisable shortcuts, it first suggested the "doubling and halving method (though it stupidly started by doubling 32 and halving 64...), and got it right.
I then told it I know the powers of two up to 2^24 by heart, and asked if that changed things.
It then reasonably pointed out this means I know 2^5=32 and 2^6 = 64, and 2^5+2^6=2^(5+6) = 2^11 = 2048 and got the rules right (that was exactly what I intended when I pointed out I remember the powers of two).
So it's not all that awful at these things. It does badly when you effectively try to get it to do maths by blind recall and without nudging it to work step by step, sure.
Where it then falls down tends to be when you ask it to do calculations which involves repetitively applying the same rules many times over, where it will tend to start out well, but occasionally make stupid little mistakes.
If anything the type of mistakes it makes are scarily close to the same kind of lapses in focus humans get when doing the same, where we just get sloppy and fail to add two numbers < 10 correctly for no good reason in the middle of doing it correctly many times, and fail to go back and verify each step.
Where some see LLMs struggling with math, I see LLMs trying to do math in a way that is disturbingly close to how a human school child would, and making the same types of mistakes.
But it's a clearly suboptimal approach. Humans and AI alike can do well with bad approaches if they must. But we can find alternative ways and we need not shoehorn AI into being LLMs.
I'm not convinced it's better than humans in general at a "math without math approach" yet, but it's certainly better than a lot of people at it. They also do make really trivial mistakes sometimes. But I also don't think there is any indication that any of this is down to things that can't easily be trained out of it. I