> These seem to generalize.
I'm not convinced. Zhou et al explicitly is teaching the model to decompose the longer chain addition into a windowed problem. Which yeah, that is good and how we humans do this. But that is limited, see Figure 10. It is a pretty heavy prompt. Figure 3 is showing that the method is not very robust (aka, not generalizing). We see changing symbols really hurts performance, subtraction and multiplication are not great. But even a 90% accuracy is not suggesting great generalization when we're talking about a pretty simple algorithm, especially one that is natural to computers. Chen et al is doing a better job, but their prompting still shows a lack of generalization and appears to be dependent on how much the model was originally tuned for mathematical tasks in the first place, with GPT having explicitly been tuned for this (if we are to believe OpenAI claims of working on making the model better at math). But this is convincingly a better method, just not convincingly generalized.
I'll note that an important aspect of generalization is not just extending to the OOD but also: not requiring fancy prompts at inference, working with arbitrary symbols in few shot exampling (<5 examples -- not batches -- it should know), and without significantly affecting the rest of the knowledge base. If you taught a LLM to only be good at math and nothing else I would both be impressed but also not call it a generalized model.
I'm not a NLP person so I don't know the nuances of the specific datasets used here, but I'd also be careful at evaluation, especially with GPT. Datasets like these are incredibly difficult to generate in ways that do not result with training data in the test set. As an example, I'll point to our classic HumanEval dataset, which is often used for testing code performance. A paper with 60 authors thought that simply by writing code by hand that it would result in unspoiled data. But they chose code problems that were similar to the style of interviews/leet code. Guess what, you can search github for similar strings and find them (pre-cuttoff). You can either explore yourself or search my chat history if you'd like specific examples. Clearly such an evaluation is not a great one.
> This was a retrieval from training issue not an inference one.
Actually I disagree. We know (likely) why this happens -- because the autoregressive nature biases a model towards a specific sequence direction -- but that doesn't mean it isn't generalization issues. And I very much disagree that humans will fall prone to these same issues. You keep using this claim with very little evidence. An explicit example from that work is the LLM getting correct the answer to "Who is Tom Cruise's mother?" (Mary Lee Pfeiffer) but not getting the correct answer to "Who is Mary Lee Pfeiffer's son?" Humans naturally handle this difference trivially. The thing is that this specific knowledge is probably not known to most people so the latter question is more likely to confuse the person. But this is different.
And we need to also recognize that humans are even explicitly evaluated this way. Learning history, in a textbook or lecture you'll be presented information like "On December 7, 1941
the Japanese attacked Pearl Harbor." But then on a test you'll be asked "What day did the Japanese attack Pearl Harbor?" That's specifically a reversal. When was X? Why is the day X important? These are explicit ways that people test other people, at a very early age, so I disagree that it is a highly common occurrence for people to fall prone to this. I'm sure you'll find example, but I'm willing to bet that those examples have another factor (you also won't be able to test the examples I gave on LLMs because they will have specifically seen both directions __because__ we do this kind of testing).
> It just has no incentive to communicate this. Don't blame the LLM here. Blame the very human data that doesn't encourage this.
You're right that we can probably do a better job at training the models to respond this way. But they naturally don't. But that is the clear demonstration of not understanding. I'd equally claim that a person doesn't understand something if they rambled off bullshit and were trying to talk their way into a solution or post hoc explain why their wrong answer is correct. Either way, it doesn't challenge the claim that it doesn't demonstrate a lack of understanding. Especially when you see these ramblings not be self consistent.
> Links
I'm not sure what you're showing here. They're kinda supporting my explicit point of RLHF messing with the distribution. I'll add more that this is a crazy hard problem to even evaluate because what is being shown as metrics are extremely aggregated. Aggregation is the bane of evaluation but so is dimensionality.
> I'm not trying to say that LLMs are perfect. and as long as people understand that saying an LLM doesn't understand x doesn't mean an LLM doesn't understand anything at all, i'll happily state weaknesses.
Just to be clear, no one has claimed that an LLM has 0 understanding. Those of us critiquing the understanding claim are pretty well aligned with that the model has some understanding. Of course it does. That's what a fitting function is. But we're challenging the claim of understanding in the generalized context which is what you've been pointing towards. Things like an LLM having a world model. Things like an LLM understanding how addition works (compared to understanding how addition works for 2-6 digit numbers). These things are quite different and I've been trying to be very clear that I'm not claiming LLMs are a load of bullshit. I've explicitly stated as much in other comments, but the prior comment wasn't as necessary given the context. (not that this statement should be necessary in the first place. It is only a result of over hype that people binarize critique. I really don't want to see more of this religious nature around ML, it is harmful to our community)