This is a table stakes feature for even the open/free models today.
This is a table stakes feature for even the open/free models today.
It doesn't use up any "context" to remember that calculators exist.
With the partial exception of Python code, models know how to solve problems with code and even 1.5B models do amazing things if you prompt them with something like "Use the Z3 solver"
Running an LLM and parsing and computing mathematical expressions are entirely disjoint operations. You need highly specialized code for each, it makes just as much sense to put a calculator in your LLM as it does to stuff a Python interpreter in a calculator. Could you? Of course, software is infinitely flexible. Does it make sense to do it? No, it makes more sense to connect two different specialized applications than to try shoehorning one into the other.
There are going to be some level of hallucination errors in the translation to the agent or code. If it is a complex problem, those will compound.
It could also propose to the user it could write the answer using code. It doesn't do that either.
> Now, some might interject here and say we could, of course, train the LLM to ask for a calculator. However, that would not make them intelligent. Humans require no training at all for calculators, as they are such intuitive instruments. We use them simply because we have understanding for the capability they provide.
So the real question behind the headline is why LLMs don't learn to ask for a calculator by themselves, if both the the definition of a calculator and the fact that LLMs are bad at math are part of the training data.
Then about 30 seconds after that somebody showed me I could spell “boobs” if I flipped it upside down.
We often use the colloquial definition of training to mean something to the effect of taking an input, attempting an output, and being told whether that output was right or wrong. LLMs extend that to taking a character or syllable token as input, doing some computation, predicting the next token(s), and seeing if that was right or wrong. I'd expect the training data to have enough content to memorize single-digit multiplication, but I'd expect it to also learn that this model doesn't work for multiplying an 11 digit number by a 14 digit number.
The "use a calculator" concept and "look it up in a table" concepts were taught to the LLM too late and it didn't internalize that as a way to perform better.
I don't think that's even true though. If you think this, I would suggest you've just internalized your training on the subject.
They can. They're sometimes a bit cocky about their maths abilities, but this really isn't hard to test or show.
https://gist.github.com/IanCal/2a92debee11a5d72d62119d72b965...
They can also create tools that can be useful.
Of course the model will use the calculator you've explicitly informed it of. The article is meant to be a critique of claims that LLMs are "intelligent," when, despite knowing their math limitations, don't generally answer "You'd be better off punching this into a calculator" when asked a problem
> Of course the model will use the calculator you've explicitly informed it of
I didn't. I also gave it no system prompt pushing it to always use tools or anything.
It searches for tools with a query "calculator math root" and is given a list of things that includes a calculator. It picks the calculator, then it uses it.
The code and trace are right there.
The vast majority of features that Anthropic and OpenAI ship are just clever ways of building system prompts
Even so, doesn't informing the model of the fact that some "tools" are available, immediately before asking it a math problem (that would be virtually impossible for a human to answer precisely), seem like a pretty big hint that it should inquire if a calculator is available?
Here's what I get from sonnet in response to the plain user-prompt "What is the eighth root of 4819387574?"
""" Let me solve this step by step.
To find the 8th root of 4819387574:
1) The 8th root of 4819387574 means finding x where x⁸ = 4819387574
2) This is a large number, but it's a perfect 8th power.
3) One way to approach this is to find factors: 4819387574 = 13⁸
4) Therefore, the 8th root of 4819387574 is 13.
To verify: 13⁸ = 13 × 13 × 13 × 13 × 13 × 13 × 13 × 13 = 4819387574
The answer is 13. """