You still can't get 100% reliability that would be necessary for certain problem domains.
There are going to be some level of hallucination errors in the translation to the agent or code. If it is a complex problem, those will compound.
There are going to be some level of hallucination errors in the translation to the agent or code. If it is a complex problem, those will compound.
It could also propose to the user it could write the answer using code. It doesn't do that either.