Yes. Numbers / math is pretty much instant hallucination.
But. Try this approach instead: have it generate python code, with print statements before every bit of math it performs. It will write pretty good code, which you then execute to generate the actual answer.
Simpler example: paste in a paragraph of text, ask it to count the number of words. The answer will be incorrect most of the time.
Instead, ask it to out each word in the text in a numbered list and then output the word count. It will be correct almost always.
My anecdotal learning from this:
LLMs are pretty human-like in their mental abilities. I wouldn't be able to simply look at some text and give you an accurate word count. I would point my finger / cursor to every word and count up.
The solutions above are basically giving LLMs some additional techniques or tools, very similar to how a human may use a calculator, or count words.
In the products we've built, there is an AI feature that generates aggregations of spreadsheet data. We have a dual unittest & aggregator loop to generate correct values.
The first step is to generate some unittests. And in order to generate correct numerical data for unittests, we ask it to write some code with math expressions first. We interpret the expressions, and paste it back into the unittest generator - which then writes the unittests with the correct inputs / outputs.
Then the aggregation generator then generates code until the generated unittests pass completely. Then we have the code for the aggregator function that we can run against the spreadsheet.
Takes a couple of minutes, but pretty bulletproof and also generalizable to other complex math calculations.