LLMs do not guarantee any quality in the output even when processing text, and should in my opinion be verified before used in any serious applications.
LLMs do not guarantee any quality in the output even when processing text, and should in my opinion be verified before used in any serious applications.
That isn't really true.[0] The application of calculators to a subject matter is something that does need to be considered in some use cases.
LLMs also have accuracy considerations, and although it may be to a different degree, the subject matter to which they're applicable has a broad range of acceptable accuracies. While some textual subject matter demands a very specific answer, some doesn't: For example, there may be hundreds or thousands of various ways to summarize a text that could be accurate for a particular application.
0: example: https://www.reddit.com/r/calculus/comments/upjdn4/why_do_all...
The issue with LLMs is that they can be so unpredictable in their behaviour. Take the following prompt that asks GPT-4 to validate the response to "calculate 2+3+5 and only display the result":
https://beta.gitsense.com/?chat=6d8af370-1ae6-4a36-961d-2902...
GPT-4o mini contradicts itself, which is not something one would expect for something we believe to be extremely simple. However, if you ask it to validate the response to "calculate 2+3+5," it will get it right.
https://beta.gitsense.com/?chat=43221de5-bff6-487a-8c0f-48ca...
By adding "and only display the result," GPT-4o mini was thrown for a loop; examples like this should give us pause.
If I ask my TI-89 to "Summarize the plot in Harry Potter and the Chamber of Secrets" it responds "ERR"! :D
LLMs are good text processors, pocket calculators are good number processors. Both have limitations, and neither are good at problem sets that are outside of their design strengths. The biggest problem with LLMs aren't that they are bad at a lot of things, it's that they look like they are good at things they aren't good at.
I've personally found it is extremely unlikely for multiple good LLMs to fail at the same time, so if you want to process text and be confident in the results, I would just run the same task across 5 good models and if you have a super majority, you can be confident that it was done right.