You can always make excuses for the LLM afterwards, but software with hidden risks like this would not be considered good or reliable in any other context.
You can always make excuses for the LLM afterwards, but software with hidden risks like this would not be considered good or reliable in any other context.
Good example: I sunk probably an hour into trying to get Gemini Advanced to help me integrate it with a personal Google Calendar account. I kept asking it things and going crazy because nothing lined up with the way things worked. Finally, it referred to itself as Bard and I realized it was giving me information for a different product. As soon as I asked "are you giving me instructions for Gemini Advanced or Bard?" it was like "OH LOL WOOPS!! YOU GOT ME BRO! XD I CAN'T DO ANY OF THAT! LOL." Which, honestly, is great. Being able to evaluate its answers to realize it's wrong is really neat. Unfortunately, it was neat too late and too manually to stop me from wasting a ton of time.
I have decades of experience working in software-- imagine some rando that didn't know what the hell Bard was or even imagine this thing with "Advanced" in the name couldn't even distinguish between its own and other products' documentation.
Did it evaluate its answers, or did your expression of doubt cause the eager-to-please language model to switch from "generate (wrong) instructions because that's what the user asked for" to "acknowledge an error because that's what the user asked for"?
How many times have we seen "Oops, you're right! 2 + 2 is actually 5! I apologize for saying it was 4 earlier!"
I once gave a 10-dollar bill to a young man serving at the cashier at a store, and he gave me 14 dollars back as a change. I pointed out that this made no sense. He bent down, looked closer at the screen of his machine, and said "Nope, 14 dollars, no mistake". I asked him if he thought I gave him 20. He said no, and even shown me the 10-dollar bill I just gave him. At that point I just gave up and took the money.
Now that I think about it, there was an eerie similarity between this conversation and some of the dialogues I had with LLMs...