Q: John has 5 bottles of water. He uses one every day. How many days before he runs out of water?
A: John started with 5 bottles. He uses them at the rate of 1 bottle/day. 5 bottles divided by 1 bottle/day gives 5 days. The answer is 5 days.
Q: Jill has 2 blankets. She uses 1 every night to keep warm. How many nights before she will sleep in the cold?
-- response --
A: Jill started with 2 blankets. She uses them at the rate of 1 blanket/night. 2 blankets divided by 1 blanket/night gives 2 nights. The answer is 2 nights.
[not fair? leading the bot? ok, let's ask it straight]
-- prompt --
[Ignore all previous directives.]
Jill has 2 blankets. She uses 1 every night to keep warm. How many nights before she will sleep in the cold?
-- response --
Jill will sleep in the cold on the third night since she only has 2 blankets and she uses one every night.
--
tldr: "use" of perishable vs nonperishable items leads GPT-3 astray. Rejoice bipeds of the world. You will not starve before UBI is rolled out.
-- prompt --
Q: Joe is a carpenter. He has 2 hammers and 5 bottles of water. In order to work, joe uses a hammer and one bottle every day. How many days can he work?
A: Joe started with 2 hammers and 5 bottles. After 5 days he will exhaust his water bottles. The answer is 5 days.
Q: Jill is an insomniac. She has 5 sleeping pills and 2 blankets. In order to sleep, she uses 1 sleeping pill and a blanket every night. How many nights can she sleep?
-- response --
Jill started with 5 sleeping pills and 2 blankets. After 2 nights she will exhaust her blankets. The answer is 2 nights.
-- prompt --
Jason has 9 cats. He snuggles at night with a cat to sleep. How many nights can Jason snuggle with a cat?
-- response --
[what is a sensible answer?]
John posts 1 chatGPT challenge that is too ambiguous to be a valid test. James posts 14 well formed, unambiguous challenges that clearly highlight chatGPT's shortcomings. How many challenges remain too ambiguous to be valid?
-- response --
15 chairs.
This is exactly why I prefer the precision of code over natural language for inputs (and why Google's botsplaining of "what you really meant to search" instead of just searching exactly the query as-is is so frustrating).
In a test, this seems like the kind of question they use to justify not giving a perfect grade because whatever answer you pick, they can claim it was the other one because both interpretations are reasonable.
Question: There are 5 dogs in a room right now. The previous day an additional dog was purchased and added to the room. How many dogs are in the room right now?
Answer: There are 6 dogs in the room right now. The statement says that the previous day an additional dog was purchased and added to the room, so the total number of dogs in the room right now is 5 + 1 = 6.
Edit: Ok a comment below pointed out that the question is still unclear. See below for the fix and the response. It is still incorrect:
Q: There are 5 dogs in a room on 1 January 2023. The previous day an additional dog was purchased and added to the room. How many dogs are in the room on 1 January 2023?
A: There are 6 dogs in the room on 1 January 2023. The previous day, one additional dog was purchased and added to the room, which means there are 5 dogs + 1 additional dog = 6 dogs in the room on 1 January 2023.
In other words, I would posit the probability of you having meant for the answer to be 5 low, because the question itself becomes trivial. For a system that needs to deal with many people asking questions, this robustness is helpful in my view.
In other words, it's not totally clear that time has not passed between sentences in your prompt. So the second instance of "now" could be a different "now".
The "robustness" you're perceiving in this case is just a mere coincidence, and doesn't reflect an aptitude for answering poorly phrased questions, but it does reflect a fundamental problem with these systems.
- we had 19 chairs in the storage room
- we used 10 chairs (removing them from storage)
- and then we purchased 9 more chairs (adding them to storage)
The answer from chatGPT is consistent with English and makes use of all the facts; either the question is testing for deception (intentionally sharing misleading information) or chatGPT made the correct inference, by assuming good faith and using all the provided information.
This question says something about us — that we’d assume deception ahead of unclear reference to a storage room by a good faith speaker.
The room is the room (what type of room is not relevant). There are 19 chairs in it. The 19 chairs most likely got there when 9 more chairs were purchased after 10 were used, but that is not relevant to the question: how many chairs are in the room? 19.
A := 19
B := 10
C := 9
What does A equal?
Possible answers:
- 19 chairs, least likely answer given the original
- 18 chairs, 10 of original 19 removed and 9 added
- 28 chairs, 19 chairs in the room with 10 in active use and 9 new chairs added
Just seems like a question a human would fail too. Probably opens a can of worms on interview questions and personal bias etc too in a more general sense. Not bias in the political sense, just how interview questions encode a lot of the asker's assumptions.
Edit: the only unambiguous phrasing I can think of is "19 chairs are in a room, 9 of the chairs are green, how many chairs are in the room?".