Why are we using LLMs as calculators?
vickiboykis.com
vickiboykis.com
The linked tweet https://twitter.com/yuntiandeng/status/1836114401213989366 is far more interesting to me, gpt models clearly _can_ learn to multiply with intermediate tokens, but even o1 currently doesn't. And yet this would be a case where generating synthetic data is almost trivial. And moreover, being able to perform computations in this fashion would be valuable for many types of benchmarks (e.g. FrontierMath, since I'm sure at the end of the day you'll have to grind through some computation).
So why hasn't it been a priority? I remember some NeurIPS presentation claiming that heavily training on math in this fashion hurt language scores. But then the follow-up would be to have specialized models for each and route between them...
Example: "I have a 120 square foot room, and I want to store liquid nitrogen in it. How many liters would it take to displace enough air to be a concern? What kind of CFM should a ventilation system use to clear the room?"
I sanity check the numbers, but it's really nice to have such an interdisciplinary calculator like this.
Speak for yourself. I’d like them to do math for the sake of replacing calculators. Well, not really. But I’d like them to be a really good natural language interface for a calculator
A reasoning engine it certainly is not either. Vendors try to shoehorn it into that purpose like OpenAI with their o models. But I think that's more because of the popularity of LLMs. Other types of AI models will be better at this.
I think eventually the LLM will just handle the human interaction and behind it will be a host of more specialised models to provide the actual content. For stuff like reasoning, knowledge, calculation etc.
If o4 can’t go beyond 4x4 accurately, then anyone using LLMs for business spreadsheets or science is a serious mistake