- Test 1: First he can just draw squiggles on the math test
- Test 2: Then he can do arithmetic correctly
- Test 3: He fails on the last details on the algebraic calculation.
Now, event though he fails on all tests, any reasonable parent would see that he improving nicely, and would be able to work in his chosen field in a year or so.
Or alternatively, if we talk about AI, we can set the Test as a threshold, and we see the results are continuously trending upwards, and we can expect the curve to breach the threshold in the future.
That is; measuring improvement, instead of pass/fail, allows one to predict when we might be able to use the AI for something.
When you actually do these millions of tests, I don't think it really matters what the exact success metric is - an AI which is 'closer to correct, but still wrong' on one test will still get more tests correct overall on the dataset of millions of tests.
I'm not saying LLM will achieve AGI (I don't know if it will, or when it does we'll even know). But somehow people seem to be judging AI's intelligence with this simple procedural:
1. Find a task that AI can't do perfectly. 2. Gotcha! AI isn't intelligent.
It just makes me question humans' intelligence if anything.
1. Choose a few good blunt instruments we use to gatekeep students on the premise that it tests their "intelligence" (or wait, do we mean subject matter comprehension with this one?)
2. Apply a big ol' machine learning model to those tests
3. Woa it's smarter than a third grader! OMG it's smarter than a lawyer! You guys this must be ASI already!
Rhetoric and selective rigor can justify any perspective. Smart and stupid arguments can be made for any position. Water is wet
I also can't claim to know with certainty whether transformers are going to end up being AGI in some meaningful sense, but I will definitely say that we've created a lot of rubrics for assessing human intelligence that mostly exist for expediency, and a cursory glance at education should tell you there's a lot of Goodhart's Law going on with all of 'em. I know for a fact I can do a damn good job on your average multiple choice test on knowing some etymology and being good at logical elimination, and I can bullshit my way through an essay, both without taking the class, and I view this more as a flaw in the instrument than evidence that I'm a godlike superintelligence that can just know anything without studying it. Humans make a lot of tests that are soft to bullshitting with a little pattern-recognition thrown in
> Human beings do arithmetic problems wrong all the time
Humans built cars and planes and massive ships before we had calculators, that requires a massive amount of calculations that are all perfect to be possible. Humans aren't bad at getting calculations right, they are just a bit slow. Today humans are bad since we don't practice it, not because we can't. LLMs can't do that today, can learn and can't is a massive difference.
Is it?
LLMs fails to figure out that this is what it has to do, instead it looks like it has a ton of specialized rules to handle arithmetics that results in a lot of errors in the output and are extremely expensive to run.
Because an LLM is a neural network and neural networks contains neural networks. There is nothing stopping it from having an embedded neural network that learned how to do computations well, except an inability to identify such structures and patterns well enough to train for it.
Humans are good at the former, but not the latter.
See below which I have just run on GPT4: https://chat.openai.com/share/3adb3aa2-8aec-474f-bdb0-4d761d...
That was how everyone did it back then, it really isn't that hard to do. Most people today never tried to do it so they think it is much harder than it actually is.
This is why prompting LLM's to show their steps works so well, it makes them work through the problem "in their head" more efficiently, rather than just spit out an answer.
However, you can give LLM's external access to tools. Ask GPT4 a particularly challenging math problem, and it will write a python script and run it to get a solution. That is an LLM's "pen and paper".
No, that is an LLM's calculator or programming, it doesn't actually do the steps when it does that. When I use pen and paper to solve a problem I do all steps on my own, when I use a calculator or a programming language the tool does a lot of the work.
That difference is massive, since when I use a calculator that doesn't help me learn numbers and how they interact and how algorithms works, while if I do the steps myself I do. So getting an LLM that can reliably execute algorithms like us humans can is probably a critical step towards making them as reliable and smart as humans.
I do agree though that if LLMs could keep a hidden voice they used to reason before writing they could do better, but that voice being shown to the end user shouldn't make the model dumber, you would just see more spam.
Maybe we should be giving the LLM's MS paint instead of python to work out problems? There is nothing unique or "human" about running through a long division problem, it is ultimately just an algorithm that is followed to arrive at a solution.
Yes, which is why we should try to make LLMs do them and that way open them up to learn much more complex understanding of algorithms and instructions that humans has yet to build a tool for.
> You need to do a lot of "steps" to write a program that solves your question. Debatably even more steps and more complexity than using pen and paper.
What does this have to do with anything? I am highlighting a core deficiency in how LLMs are able to reason, you saying that what they currently do is harder doesn't change the fact that they are bad at this sort of reasoning.
And no, making such a program doesn't require more steps or understanding. You Google for a solution and then paste in your values, that is much easier to teach a kid than to teach them math. I am sure I can teach almost any 7 year old kid to add two numbers by changing values in a python program in about an hour, much faster than they could learn math the normal way. Working with such templates is the easiest task for an LLM, what we want is to try to get the LLM to do things that is harder for it.
"I have a problem for you to solve. Muffins sell for $3/each. rick bakes 30 muffins a day. Tom bakes 2 muffins monday, 4 tuesday, 6 wednsdays, up to 14 on sunday. On days which tom and jerry combined bake more than 41 muffins, the price of the muffins drops to $2.50. How much total revenue do rick and tom take in during a full week, combined."
Please tell me how ChaptGPT4 writing a script to solve that is not logical reasoning, while a human pulling out pen and paper to do it is...
I changed the prompt a bit (made all the numbers 3-4 digits) and gpt-4 answered with this, it just made up numbers for the days that you didn't add numbers for so it failed before it even came to arithmetics. Here is what it said, after I said this about tom "Tom bakes 2911 muffins monday, 491 tuesday, 699 wednsdays, up to 149 on sunday.", it just assumed sundays number was for all other weekdays not given a human wouldn't do that, and it missed the "up to" statement. Maye the large numbers I gave threw it off, but if that is enough to throw it of just shows that it can't really reason.
So thanks for that, more evidence these models are bad at reasoning.
Here is the first part of what it responded with, it is wrong already here:
First, let's calculate the number of muffins baked by Tom during the week:
Monday: 2911
Tuesday: 491
Wednesday: 699
Thursday: 149
Friday: 149
Saturday: 149
Sunday: 149
Edit: Here it made an arithmetics error just below, the error is that 4062 is not greater than 4199, so two critical errors, I taught math at college for years and you wouldn't find many students making mistakes like this: Let's determine the days when Tom and Rick combined bake more than 4199 muffins:
Monday: 2911 (Tom) + 3571 (Rick) = 6482
Tuesday: 491 (Tom) + 3571 (Rick) = 4062
Wednesday: 699 (Tom) + 3571 (Rick) = 4270
On Monday, Tuesday, and Wednesday, they bake more than 4199 muffins combined, so the price of the muffins drops to $2851.50 on those days.Unless of course you didn't realize that tom has a pattern to his baking, at which point to irony becomes palpable.
And on top of that, I am willing to bet if you give me your prompt, I would be able to restructure it in such a way that GPT4 would be able to answer it correctly. More often than not, people are just really bad at properly asking it questions.
I used your exact quote and just changed the numbers, it is still a perfect information problem.
Or, ah right you mean you gave me an imperfect information problem since you assumed the reader would guess those values. Yeah, I read it as a perfect information problem where all values were given, and then you would give the income as a range of possible income values based on how many muffins were baked on Sunday. None of the LLMs I sent it to managed to solve it entirely, it is a pretty easy problem.
Reasonable way to parse your sentence is:
Monday: 2, Tuesday: 4, Wednesday: 6, Sunday: 0-14, rest: doesn't work so 0
> Unless of course you didn't realize that tom has a pattern to his baking, at which point to irony becomes palpable.If you didn't say he baked on those days then he didn't bake on those days. The specification is clear. If I say "I will bake 2 muffins on Tuesday and 6 muffins on Sunday" the reasonable interpretation is that I wont bake anything the rest of the days. Why would you assume he baked anything at all those days?
Or if I say "Emily will work Mondays and Thursdays", do you just guess the rest of the days she will work? No, you assume she just works those days.
Is that a standard problem you wrote from memory? Not sure why you would assume there were muffins baked in the days you didn't list.
For example, if I say Tom bakes up to 14 muffins on Sunday, then the reasonable interpretation is that Tom will bake 0-14 muffins on Sunday. Maybe you should write the prompt clearer if you mean something else? Because as written anyone would assume that he didn't bake the other day, and on sundays he baked up to 14 muffins.
Anyway, it failed even with your "up to" interpretation meaning the reader should fill in the values, it still made that math error. But it using your "up to" interpretation there is a huge red flag, since in a real environment nobody would give that kind of information as a riddle with hidden values, you would specify all the values for each day each person worked and the rest you assume the person just isn't working and baked 0 muffins. If the LLM starts to guess values for some patterns and words where it doesn't make sense then it is really unreliable.
I don't have a stake in this muffin game, but that's indeed how I interpreted the instructions when reading them.
Had it said "and so on up tp 14 on Sunday" I would assume he baked each day.
Thankfully GPT4 has strong reasoning skills and knew exactly what I meant.
https://chat.openai.com/c/b0ed06f1-c0d3-46a6-b07c-289b328417...
I encourage you to see the chat yourself, and would love to here how it's not reasoning.
Edit: Fixed Link: https://chat.openai.com/share/991ca8af-f735-436f-bfc2-5df929...
you can click the [>_] at the end for the code generated.
Seem to have hit reply cut off
Unable to load conversation b0ed06f1-c0d3-46a6-b07c-289b328417bbNo, that's an LLM's Python playground.
An LLM's "pen and paper" is "think step by step" where it gets to see it's own output to keep track of what it is doing.
I'd expect that with appropriate prompting one could get a good model to one/few-shot learn how to do addition this way.
Speak for yourself. Even though I've always been strong at my conceptual understanding and problem solving in math, I always found it difficult to avoid arithmetic mistakes on pen and paper and could never understand why I was assessed on that. I could have done so much better in high-school math if I was allowed to use a programmable computer for the calculations.
And I think it's the same for LLMs, we should assess them on doing the arithmetic in a single pass, but rather on writing the code to perform the calculation, and responding based on that.
But I do acknowledge that there are probably some or many humans that maybe can't reach that level of reliability with arithmetics.
Compare that to understanding arbitrary base64-encoded strings; that's much harder for humans to do without tools. Tokenization still isn't _the_ greatest fit for it, but it's a lot more tractable, and LLMs can do it no problem. Even understanding ASCII art is impressive, given they have no innate idea of what any letter looks like, and they "see" fragments of each letter on each line.
So I'm not sure if I agree or disagree with you here. I'd say LLMs in fact have very impressive capabilities to learn logical structures. Whether grammar is the problem isn't clear to me, but their internal representation format obviously and enormously influences how much harder seemingly trivial tasks become. Perhaps some efforts in hand-tuning vocabularies could improve performance in some tasks, perhaps something different altogether is necessary, but I don't think it's an impossible hurdle to overcome.
The tokens are just the input - the internal representation can be totally different (and that format isn't tokens).
The conversion to a helpful form is required anyway (also lets remember that computers don't work in base 10, and there isn't really a reason to believe that base 10 is inherently great for LLM's either)
hundreds | tens | ones
1 2 3
+ 2 1 5
-----------------------
3 3 8
Rather thanunoDOOOOS(third) {}{}{} [512354]_ = three"ate
* replace {}{}{} with addition, {}{} is subtraction unless followed by three spaces in which case it's also addition * translate and correct any misspellings * [512354] look up in your tables * _ is 15 * dotted lines indicate repeated numbers
Technically they're doing the same thing. One we would assume is harder to learn the fundamental concepts from.
The issue is not the fact that the model "thinks or doesn't think in tokens". The model is forced at the final sampling/decoding step to convert it's latent back into tokens, one token at a time.
The models are fully capable of understanding the premise that they should "output a 5-7-5 syllable Haiku", but from the perspective of a model trying to count its own syllables, this is not possible, as its own vocabulary is tokenized in such a way that not only does the model not have direct phonetic information within the dataset, but it literally has no analogue for how humans count syllables (measuring mouth drops). Models can't reason about the number of characters or even tokens used in a reply too for the same exact reason too.
The person you're replying to broadly is right, and you are broadly wrong. The internal format does not matter when the final decoding step forces a return of tokenization. Please actually use these systems rather than pontificating about them online.
IDK, there was an article posted here about yet another LLM that performed very badly on the math tests because they mistakenly left out all the math training data.
What impressed me was that it could learn any math at all from just 'reading' books or whatever. Though, perhaps, any correct answer could be attributed to pure luck, dunno.
To add multi-digit numbers requires short term memory (are we on the units, or tens? was there a carry?), which LLMs don't have, so that's really the issue.
The normal workaround for lack of memory in LLMs is "think step-by-step" to use it's own output (which gets fed back in as an input) as memory, and I'd assume that with appropriate training data and prompting an LLM could learn to do it in this fashion - not just giving the answer, but by giving all the steps.
I suppose in theory LLMs could do limited precision math even without memory, if they did it in a single pass through their stack of transformer layers (96 for GPT-3) - use first few layers to add units and generate carry, next few layers to add tens, etc. I'm not sure how, or even if, one could train them to do it this way though - perhaps via a curriculum training agenda of first adding single digit numbers, then two-digit ones, etc ?
That'd depend on the design of the neural net and training objective.
It's certainly not something that comes naturally to an LLM which neither has numbers as inputs or outputs, nor is trained with an arithmetic objective.
Consider inputting "12345 * 10" into GPT-4. First thing it is going to do is tokenize the input, then embed these tokens, and these embedding vectors are then the starting point of what the transformer has to work with...
https://platform.openai.com/tokenizer
You can use OpenAI's tokenizer tool (above) to see how it represents the "12345 * 10" character sequence as tokens, and the answer is that it breaks it down into the token ID sequence [4513, 1774, 353, 220, 605]. The [4513, 1774] represents the character sequence "12345", and "605" represents the character sequence "10".
These token ID's will then be "embedded", which means mapping them to points in a very high dimensional space (e.g. 4096-D for LLaMA 7B), so each of those token ID's becomes a vector of 4096 1's and 0's, and these vectors are what the model itself actually sees as input.
So, for "12345 * 10", what the model sees during training is that whenever it sees V1 V2 V3 V4 it should predict V5, where V1-5 are those 4096-D input token embeddings. The model has no idea what any of these mean - they might represent "the cat sat on the mat" for all it knows. They are just a bunch of token representations, and the LLM is just trying to find patterns in the examples it is given to figure out what the preferred "next token" output is.
So, could you build (and train) a neural net to multiply, or add, two numbers together? Yes you could, if that is all you want to do. Is that what an LLM is? No, an LLM is a sequence predictor, not an NN designed and trained to do arithmetic, and all that is inside an LLM is a transformer (sequence-to-sequence predictor).
To solve this you would need some sub networks that are pretrained to handle numbers and math and other domains, and then you start training the giant LLM it can find and connect those things. But we don't know how to do that well yet afaik, and I bet all the big players has already tested things like that. As you say adding capabilities to the same model is hard.
If you want the LLM to be better than a human at math, then give it a calculator, or access to something like Wolfram Alpha for harder problems. Your proposed solution of "give it a specialized NN for math" is basically the same, but if you are going to give it a tool, they why not give it a more powerful one like a calculator ?!
A more serious byproduct of the tendency to talk about machines in anthropomorphic terms is the companion phenomenon of talking about people in mechanistic terminology. The critical reading of articles about computer-assisted learning —excuse me: CAL for the intimi— leaves you no option: in the eyes of their authors, the educational process is simply reduced to a caricature, something like the building up of conditional reflexes. For those educationists, Pavlov’s dog adequately captures the essence of Mankind —while I can assure you, from intimate observations, that it only captures a minute fraction of what is involved in being a dog.
https://www.cs.utexas.edu/users/EWD/transcriptions/EWD09xx/E...
This is definitely not true in the real world. Approximate solutions are often good enough to answer a question.
I don't think everything has uncertainty and thresholds to it, especially, when it actually resides outside of a technical implementation.
Somewhere between "it's always wrong" and "it's always right unless the bits got flipped by cosmic rays" we deem the accuracy to be good enough.
Any implementation (or write down etc.) of something can have errors, but the errors are in the implementation and do not give rise to uncertainty outside of the implementation. There is no uncertainty as to what the sum of two integers should be (within the usual mathematics).
And ChatGPT is definitely able to get improve its answer by iterating, it just depends on the toughness of the problem. If it's too difficult, no amount of iteration will get it much closer to the correct answer. If it's closer to its reasoning limits, then iterating will help.
Edit: was answering to gp, no idea how my post got here
> But if you stop them just there, an error persists
But humans doesn't stop there when they are making things that needs to be reliably correct. When errors aren't a big deal humans make a lot of errors, but when errors costs life humans become very reliable by taking more time and looking things over. They still sometimes makes mistakes that kills people, but very rarely.
Back-of-the-envelope / mental math often works like that, and it's something that humans regularly use, so clearly it has some use.