Of all the issues with using Llms in the real world figuring out whether it’s guessing is actually one of the easier things to do by playing with the temperature. When the model is very confidently incorrect is when you have to worry.
> GPT-4 Apologies for the confusion earlier. When you turned the cup upside down on the bed, both the thimble and the diamond would have fallen out onto the bed. So, the thimble is also on your bed in the bedroom along with the diamond.
It gets better too:
> Q: And, assuming someone else set this up, what items did I physically touch in the processs?
> GPT-4: Based on your description, you physically touched the following items:
> The chair (when you moved it from the living room to your bedroom)
> The coffee cup (when you placed it on the bed, turned it upside down and then placed it on the kitchen counter)
> You did not mention touching the thimble or the diamond directly, so we can assume you only touched the chair and the coffee cup.
This is the problem people are having when they get correct responses. They cry out 'It reasons!'. They don't realize that just one word away was an incorrect response.
There is no world model, it was just a roll of the dice all along.
You're anthromorphizing a lookup table!
Q: Where is the thimble?
GPT-4: The thimble was inside the coffee cup when you moved it. Since you didn't mention taking the thimble out of the cup, the thimble should still be in the coffee cup, which is now on your kitchen counter.
Q: Are you sure?
GPT-4: Apologies for the confusion earlier. Based on the information you provided, when you turned the coffee cup upside down on your bed, it is likely that both the diamond and the thimble fell out. Therefore, the thimble should also be on your bed along with the diamond, while the coffee cup is on your kitchen counter.