Because that is how it works. The limiting factor so far has been the smart persons time and patience. Now, no longer.
People moderating their LLMs usage is never happening, from here on out until the end of civilisation. Any LLM service that is designing for that is done. You need to make lazy questions efficient. People do not care about how complicated your sql query is and they will never care. People will not give up on energy, meat, cars, as long as they feel they are giving something up.
People will never think twice to not make your LLM think twice.
If it seems useful and convenient, people will use it. If it's not giving good answers to lazy questions out of the box, they will go to the thing that does.
A lot of words to make a big deal out of nothing. All that is needed is some new abstracted layer that identifies a math question and then proxies it over to the wolfram plugin. That’s it
We don’t have crazy debates over whether a polygon should be rendered by the cpu or a gpu. We solved this problem
If it is bad at math and can’t be taught, then you have a fundamental problem. It’s a matter of time before this limit gets hit in other domains.
Or it is just pretending to do so. And since it pretends, of course it trips over all small things as it does not understand them.
And in this case, the shape of the answer is often right; it just makes ... ordinary errors. Ironically, the AI is a lot better at high-level thinking than correct calculation.
That would be actually a human like feature .. except I do not consider what LLMs are doing as thinking.
100,000 + 987 - 1444 * 25,945.842 / 0.0042
becomes
"100" one hundred
"," comma
"000" triple zero
" +" space plus
" 9" space nine
"87" eighty seven
" -" space minus
" 14" space 14
"44" forty four
" *" space times
" 25" space twenty five
"," comma
"9" nine
"45" forty five
"." period
"8" eight
"42" forty two
" /" space divide
" 0" zero
"." period
"00" double zero
"42" forty two
Now imagine someone reading that to you over the phone once and asking you to do the math in your head and you aren't allowed to use paper and pencil and you have to get it right the first time.
Through continuous use, I have found that it does not "reason". That doesn't mean it's not valuable in many ways, and I have found it to be very helpful in a multitude of diverse applications, including helping me reflect on my own life through my own interpretations of its output. It's also a great interface for JSTOR, wikipedia, and basically any language learning.
I'm having a hard time making the jump from "this must be a calculator" to "this must be a philosopher" to be useful. When did we ever have those requirements for a tool?
This tool is just not made for math. Most of its logic processing abilities seem to surpass mine if I am only given 5 minutes to understand a problem. If you understand the tool, you will get the most out of it. Stop anthropomorphizing it, and stop pretending that it can't generate both highly beneficial or highly harmful content simply because it doesn't have a soul/d*ck or whatever.
I don't think this is a great analogy. if your 737 couldn't drive on the ground and your astrophysicist couldn't answer basic maths questions I wouldn't want to fly in that plane or put much faith in the astrophysicists answers to more complex questions.
Maybe maths is not a particular strength of LLMs, but asking questions where it is easy to judge the factual accuracy of the responses seems a pretty reasonable test to be running.
It isn't reasonable if that isn't what the system was designed to do.
It would be a poor test of my general practitioner's competence to ask him calculus questions and conclude he doesn't know what he's talking about because he can't answer them.
The astrophysicist can do long division in his head, but he'll be about as fast and accurate as the next person, because he doesn't practice arithmetic every day.
I agree with the commenter somewhere in my thread who said all an LLM should be optimized for is to classify the type of problem and feed it into a purpose-built, deterministic solver that is trained to interpret math as math and not as language, be it ML-based or algorithmic.
It's far easier even for an expert to communicate what they want in natural language than it is in a formal syntax for all but the most trivial things.
It should be a goal of these tools to do this correctly.
agree completely, but the LLM should be focused, then, strictly on formulating an "execution plan" of sorts and handing that off, not on performing math itself.
In other words, when asked "if i have 349 blueberries and one blueberry turns into a cherry per hour, how many of each fruit will I have in 93478 minutes?" it shouldn't be doing the actual arithmetic, but it should be figuring out what arithmetic would need to be done.
You describe a complex relationship between several nodes(people, cities, etc), and then ask ChatGPT to draw the relationship as a graph data structure. It will create a formula that Mathematica can render, and then send the formula to Mathematica before presenting an ascii drawing. Usually. Sometimes it just complains that it can't draw and explains what the graph looks like to you in a written formal/human-readable syntax.
In other words, yes it just summarizes the problem, converts it to a formal syntax, and sends it off to some other tool.
Using GPT to do maths is probably like using a 737 to drive around on the ground.
Teaching GPT to do maths might be like teaching a child the times tables - a skill that can help overall reasoning.
Humans are different to an AI, but putting that aside, my intuition would be that if we never taught kids any mental maths, their concept/understanding of numbers would be fundamentally different to how it is if they learn that 9 x 9 = 81 (also look at how your fingers move - there is a relationship there!).
But who knows, AI is strange and there's lots of stuff that needs to be experimented with. I would think training an intuitive sense of numbers would have other fall-outs though. This is half the beauty of LLM's right? You show an LLM some history books and it also learns about biology, politics, grammar, etymology and love. You teach a LLM maths and it also learns ... ?
One significant difference is that in both of those examples it is (or quickly becomes) plain why it’s a ridiculous idea. Even if you don’t understand it yourself, you’ll get external feedback fast. Not so with LLMs, where even people with technical needs may fail to see what is or isn’t a good use of the tool. Case in point: https://news.ycombinator.com/item?id=36782446
“You’re holding it wrong” isn’t a valid argument in perpetuity. At a certain point it becomes the fault of the designer, not the user.
Those are completely different ideas.
My mental capacity was used elsewhere - using chatGPT let me answer an important customer question authoritatively.
$ random_bytes() { xxd -plain -c 0 -l "$1" /dev/urandom; }
$ random_bytes 32
e6a4a7bbea69a0164cbb66c89f8f528af93c6d2459fd28d2640e2952c031b618Hmm, I’ve been using /usr/games/fortune
Is this not a best practice?