Yes in an ideal world we would have a live customer support representative for every function in every facet of society, but there are a limited number of human beings available for such things, and this is a pretty reasonable place to do a first triage using a LLM for very simple questions.
When accuracy matters, answering a question incorrectly puts a person in an even worse situation than simply failing to answer the question.
And here's the thing: most front-line customer service is also clueless about difficult problems. The IRS cannot pull 10,000 seasonal experts on the line, they are going to hire barely-trained part-time accountants who also flub hard questions.
e.g. part-time front-line customer service will prefix a statement with "uhhh..." if they don't actually know what they're talking about, even if they do have trouble answering accurately.
You can literally prompt GPT4 "Prefix a statement with uhhhh if you don't know what you are talking about" and get similar behavior.
I literally just tested your prompt, with the question "is the sky blue?" and chatgpt prefixed the response with "uhhh..."
These models create the illusion of thought by statistically stringing words together, but they don't actually think or perform judgement of their own.
Edit: After digging into this for a few minutes, I challenge you to try prompting an LLM to judge the certainty of its own responses. The results I am getting are even worse than I thought it would be.
Custom instructions: "If you aren't confident in your answer, prefix your response with "Uhhhhh". Otherwise answer the same as normal."
So... 4o is not confident that only humans qualify as dependents?
I think even a very junior front-line customer service rep should be able to answer that one confidently.
It seems that what the model is actually doing is prefixing "Uhhhh" when your question is leading in a way that doesn't match the data it has. The fact that the IRS requires dependents to humans should be answerable with an extremely high confidence, and that data is without a doubt in their dataset... but again, the model doesn't actually experience human confidence or uncertainty.
https://www.irs.gov/forms-pubs-search?search=OA2143
Ultimately, the tax question you asked it is something simple for a front-line worker to answer. So either one of two things must be true:
* either GPT-4o is so bad at answering tax questions that it cannot even answer easy ones confidently
* or GPT-4o is so bad at determining its own confidence level that it doesn't know when it is able to definitively answer even an easy question.
Either situation makes it bad for this task.
As I mentioned above, humans are good for answering questions even when they don't know the answer, because they're good at expressing their confidence to other humans. In this case, you'd want the support agent to answer definitively that animals do not qualify as dependents. One could certainly make their chat bot answer unconfidently randomly, or in response to strange questions, or all the time, but then the confidence signal isn't actually providing social value of communicating certainty.