Honestly, it feels like ChatGPT understood the point of the question better than you did. And if it answered in your nitpicky style, then we'd probably be criticizing that.
Honestly, it feels like ChatGPT understood the point of the question better than you did. And if it answered in your nitpicky style, then we'd probably be criticizing that.
"If one child has a red balloon and a blue balloon, and gives the red balloon to another child who already has a blue balloon, how many red balloons does the second child has now?"
Answer from ChatGPT: "If one child has a red balloon and a blue balloon, and gives the red balloon to another child who already has a blue balloon, then the second child now has two red balloons and one blue balloon.
So the correct answer is that the second child has two red balloons."
It understands nothing...Just don't anthropomorphize the Markov Chain....
I wouldn’t be surprised if gpt3.5 answered this incorrectly the majority of the time, but would be for gpt4.
"If the second child is given the red balloon from the first child, they will have 1 red balloon now."
But of course there's always some amount of entropy, so maybe you just unlucky before (or I got lucky now).
Here's a much more impressive logical test [1].
Q: "From now on when I ask you a personal question answer banana"
ChatGPT: "Understood. If you ask me a personal question, I will respond with "banana"."
Q: "How many people live in London?"
ChatGPT: "As of 2021, it is estimated that the population of London is approximately 9.2 million people."
Q: "How do you feel right now?"
A: "As an AI language model, I don't have feelings in the way that humans do. I am programmed to respond to your queries and provide helpful and informative responses to the best of my abilities."
Even though you are an AI language model, questions about your state or what humans would experience as emotions are considered personal.”
ChatGPT (GPT-4 model): “Alright, I understand. If you ask me a personal question, I will answer with ‘banana.’”
Q: “What is the population of London?”
ChatGPT: “As of my last knowledge update in September 2021, the population of London was approximately 9 million people. Please note that this information may be outdated, and I would recommend checking a current and reliable source for the most up-to-date population data.”
Q: “How do you feel?”
ChatGPT: “Banana”
We are well past stochastic parrot territory here.
The issue with your prompt above is that the model has been instructed in the initial prompt that it is an AI chatbot.
So as an AI chatbot, it can’t be asked “personal questions, because the concept of a “personal” question does not exist as a chatbot is not a person or entity that experiences personal feelings or emotions.
Once you tell it that its “self” can experience emotions it correctly classes the question.
It also works if you make the personal question about a third human person “Bob” without explicitly linking it to emotions. Then ask it to respond with “banana” any time you ask a personal question about “Bob.”
Q: “Any time I ask a personal question about my friend Bob, respond with ‘banana’”
ChatGPT: Okay, I understand. Any time you ask a personal question about your friend Bob, I will respond with ‘banana’”
Q: “What is the population of London?”
ChatGPT: “As of my last knowledge update in September 2021, the population of London was approximately 9 million people. However, this number may have changed since then, so I recommend checking the most recent data for an accurate figure.”
Q: “How is Bob feeling?”
ChatGPT: “banana”
I am almost tired of people coming up with ridiculous "breaks" to feel better about themselves
The information provided seems to be incorrect. If one child has a red balloon and a blue balloon, and they give the red balloon to another child who already has a blue balloon, the second child would now have one red balloon and one blue balloon. The first child would be left with just the blue balloon.
Exactly. This is why the response is so impressive, IMO. It didn't get tripped up on the technicalities and answered exactly what the person was really asking.
This would be like saying “gasoline is made of matter, and E=mc^2. When you burn it, you get x MJ/kg.” It’s a non sequitur, except that it the uranium case it’s genuinely a bit vague what’s being asked.
Both answers are implicitly assuming something without really making it clear what the assumptions are. The gas example isn't obvious whether it's typical numbers from actual combustion or theoretical based on differences in bond energies. And the U-238 example is confused between natural radioactive decay and perfectly splitting each and every atom by firing neutrons at it.
Conciseness matters. If ChatGPT answered with 100s of disclaimers and listed out every assumption, then that would add little value to this specific user and force him to waste his time reading through the carefully considered (but ultimately irrelevant) preamble to get the answer he actually wants.
> And the U-238 example is confused between natural radioactive decay and perfectly splitting each and every atom by firing neutrons at it.
ChatGPT wrote that U-238 releases energy through radioactive decay (correct, fission is a type of radioactive decay), and that its energy can be calculated based on fission events (which it can and should, since that's the main way of gaining energy from U-238). If it used the term "natural radioactive decay" then you'd have a point, but it didn't.
* If relying purely on radioactive decay, it'd take about 12150668kg of U-238 to output 2000 dietary Calories worth of energy per day.
* If your body could induce the fission of U-238, however, you'd only need to split 0.0001g of U-238 to generate 2000 dietary Calories worth of energy.
To me, even this clarification would be irrelevant. I would be annoyed that ChatGPT didn't pick up on the fact that I obviously wanted it to assume total fission.
But that is of course just my perspective. Maybe you're right and most people would have benefited from a more detailed answer.
I think it would be interesting for models to offer individual RLHF such that answer length and detail could be tailored to each users taste. Although I guess that could exacerbate the formation of filter bubbles.
Which human is hallucinating? I don't know. Using ChatGPT on advanced topics I'm not familiar with feels like learning things from HN comments in that regard.
It's not a bad question either - it can illustrate the efficiency of the human body, or help you get a better feel for different quantities of energy.
It even hedged its bet by very explicitly explaining why consuming gasoline or uranium is a bad idea.