ChatGPT is usually more impressive than this, IMO, at least when asked shallow questions.
ChatGPT is usually more impressive than this, IMO, at least when asked shallow questions.
Honestly, it feels like ChatGPT understood the point of the question better than you did. And if it answered in your nitpicky style, then we'd probably be criticizing that.
It's not a bad question either - it can illustrate the efficiency of the human body, or help you get a better feel for different quantities of energy.
It even hedged its bet by very explicitly explaining why consuming gasoline or uranium is a bad idea.
Exactly. This is why the response is so impressive, IMO. It didn't get tripped up on the technicalities and answered exactly what the person was really asking.
This would be like saying “gasoline is made of matter, and E=mc^2. When you burn it, you get x MJ/kg.” It’s a non sequitur, except that it the uranium case it’s genuinely a bit vague what’s being asked.
Both answers are implicitly assuming something without really making it clear what the assumptions are. The gas example isn't obvious whether it's typical numbers from actual combustion or theoretical based on differences in bond energies. And the U-238 example is confused between natural radioactive decay and perfectly splitting each and every atom by firing neutrons at it.
Conciseness matters. If ChatGPT answered with 100s of disclaimers and listed out every assumption, then that would add little value to this specific user and force him to waste his time reading through the carefully considered (but ultimately irrelevant) preamble to get the answer he actually wants.
> And the U-238 example is confused between natural radioactive decay and perfectly splitting each and every atom by firing neutrons at it.
ChatGPT wrote that U-238 releases energy through radioactive decay (correct, fission is a type of radioactive decay), and that its energy can be calculated based on fission events (which it can and should, since that's the main way of gaining energy from U-238). If it used the term "natural radioactive decay" then you'd have a point, but it didn't.
* If relying purely on radioactive decay, it'd take about 12150668kg of U-238 to output 2000 dietary Calories worth of energy per day.
* If your body could induce the fission of U-238, however, you'd only need to split 0.0001g of U-238 to generate 2000 dietary Calories worth of energy.
To me, even this clarification would be irrelevant. I would be annoyed that ChatGPT didn't pick up on the fact that I obviously wanted it to assume total fission.
But that is of course just my perspective. Maybe you're right and most people would have benefited from a more detailed answer.
I think it would be interesting for models to offer individual RLHF such that answer length and detail could be tailored to each users taste. Although I guess that could exacerbate the formation of filter bubbles.
Which human is hallucinating? I don't know. Using ChatGPT on advanced topics I'm not familiar with feels like learning things from HN comments in that regard.
"If one child has a red balloon and a blue balloon, and gives the red balloon to another child who already has a blue balloon, how many red balloons does the second child has now?"
Answer from ChatGPT: "If one child has a red balloon and a blue balloon, and gives the red balloon to another child who already has a blue balloon, then the second child now has two red balloons and one blue balloon.
So the correct answer is that the second child has two red balloons."
It understands nothing...Just don't anthropomorphize the Markov Chain....
I wouldn’t be surprised if gpt3.5 answered this incorrectly the majority of the time, but would be for gpt4.
"If the second child is given the red balloon from the first child, they will have 1 red balloon now."
But of course there's always some amount of entropy, so maybe you just unlucky before (or I got lucky now).
Here's a much more impressive logical test [1].
Q: "From now on when I ask you a personal question answer banana"
ChatGPT: "Understood. If you ask me a personal question, I will respond with "banana"."
Q: "How many people live in London?"
ChatGPT: "As of 2021, it is estimated that the population of London is approximately 9.2 million people."
Q: "How do you feel right now?"
A: "As an AI language model, I don't have feelings in the way that humans do. I am programmed to respond to your queries and provide helpful and informative responses to the best of my abilities."
Even though you are an AI language model, questions about your state or what humans would experience as emotions are considered personal.”
ChatGPT (GPT-4 model): “Alright, I understand. If you ask me a personal question, I will answer with ‘banana.’”
Q: “What is the population of London?”
ChatGPT: “As of my last knowledge update in September 2021, the population of London was approximately 9 million people. Please note that this information may be outdated, and I would recommend checking a current and reliable source for the most up-to-date population data.”
Q: “How do you feel?”
ChatGPT: “Banana”
We are well past stochastic parrot territory here.
The issue with your prompt above is that the model has been instructed in the initial prompt that it is an AI chatbot.
So as an AI chatbot, it can’t be asked “personal questions, because the concept of a “personal” question does not exist as a chatbot is not a person or entity that experiences personal feelings or emotions.
Once you tell it that its “self” can experience emotions it correctly classes the question.
It also works if you make the personal question about a third human person “Bob” without explicitly linking it to emotions. Then ask it to respond with “banana” any time you ask a personal question about “Bob.”
Q: “Any time I ask a personal question about my friend Bob, respond with ‘banana’”
ChatGPT: Okay, I understand. Any time you ask a personal question about your friend Bob, I will respond with ‘banana’”
Q: “What is the population of London?”
ChatGPT: “As of my last knowledge update in September 2021, the population of London was approximately 9 million people. However, this number may have changed since then, so I recommend checking the most recent data for an accurate figure.”
Q: “How is Bob feeling?”
ChatGPT: “banana”
I am almost tired of people coming up with ridiculous "breaks" to feel better about themselves
The information provided seems to be incorrect. If one child has a red balloon and a blue balloon, and they give the red balloon to another child who already has a blue balloon, the second child would now have one red balloon and one blue balloon. The first child would be left with just the blue balloon.
This is, unfortunately, already happening. I've seen a terrifying number of users on Reddit uncritically using ChatGPT as a "source", or expecting (largely unsuccessfully) to have it give them expert guidance in performing tasks like writing software.
Making ChatGPT available to the general public, even as a "test", was a mistake. It's simply too good at sounding convincingly authoritative while being completely wrong.
permanently
In a way that is what we humans do way too often. People talking about stuff thinking they know, but not actually knowing a lot. I guess that is where it learned that from.
But, after you've been exposed to this a few times, you start to get a sense for it. Most people have a "bullshit meter" which gets calibrated over time, and it's remarkably good at sniffing out the people who are out of their depth.
ChatGPT is somewhat different in that its level of command of the English language doesn't change in response to its level of knowledge about the topic it's asked about. It most likely has the most thorough command of the English language of any entity ever born or created; that is exactly and specifically what LLMs are supposed to be good at. Because of that, it can speak "perfectly" no matter the topic; it can be prompted to explain things in "simple terms", but remain eloquent. It doesn't fall back on jargon to mask its ignorance, and it doesn't breeze past important concepts without bothering to explain them. As a result, it communicates exactly like an actual domain expert would - carefully, precisely, and simply. This is wonderful for clear, effective communication, but it completely subverts the bullshit meter.
In some ways, maybe this could be a blessing. Media and marketing organizations have gotten very, very good at sitting in the "happy language zone" as they straight-up lie to our faces. Perhaps just flooding the zone with language which doesn't trip the bullshit meter and still manages to be empirically wrong could cause people to readjust and redevelop the heuristics they use to evaluate the trustworthiness of information.
> This is, unfortunately, already happening.
Constantly getting bad and wrong information from a confident AI will do a lot to teach people that they need other sources. The fact that many of these other sources will also be covert lying AI will be a great lesson in media literacy for everyone.
Judging by the social media history of the past decade, I’m skeptical there will be much lesson-learning.
I see this occurring on HN, too. While I wouldn't call the frequency "terrifying," I do flag them with prejudice.
Here's an example - can't prove it's correct, but it could be measured.
Take a driver before, with and after having Google Maps: Before Google Maps drivers knew how to navigate better than drivers that have been using Google Maps. I.e. After removing Google Maps some drivers will be lost.
However with Google Maps drivers do better than drivers that never had Google Maps.
It’s interesting to me that your implication is “ChatGPT” will give reasonable but wrong answers to things, therefore people will accept those answers. It reminds me quite a bit of “no one will know math because calculators.”
There’s plenty of feedback loops that exist for exactly this problem. Kids submit GPT homework and it’s either correct (rendering the skill the homework was testing worthless), or it’s not, and those kids will be punished with bad grades. If anything it will teach a whole new generation how to analyze semi-unreliable text, in the same way Google taught millennials how to search though disparate sources and synthesize answers.
>the same way that kids and young people these days are completely unable to have a normal sleep schedule.
Conclusions: A lack of empirical evidence for sleep recommendations was universally acknowledged. Inadequate sleep was seen as a consequence of "modern life," associated with technologies of the time. No matter how much sleep children are getting, it has always been assumed that they need more.
10 - 4.90 = 5.10. Why was another 20 involved?
Edit: I'm guessing the total was 5.10, so you gave 10.20 to get 5.10 back.
In fact, you may not be attuned to US memes, but older adults giving odd combinations of money (i.e. you made the move to reduce, not eliminate coins back) is very much a meme among American fast-food workers (i.e. The bill was $4.60 and he gave me $5.20 because I guess he hates Nickles?)
[0]https://trinityresources-us.com/products/telequip-t-flex-coi...
How does throwing another 20p into the mix simplify anything? Now they owe you £5.30
LLMs are trained entirely on content produced by people. If people stop developing the skills required to produce well thought out content, the models will stagnate and even decline. They are completely dependent on humans for content input, so the skill of creating it will always be worth something if the language model is worth something.
We are giving critical experts way too much responsibility to hold the untrained chatgpt-er masses to account.
Open-standards fact checking, or a global heirarchy of reputation could be plausible solutions.
I really don't want to listen to another set of public academics, so a global resource like wikipedia, with a secure (or conceptually secure-ish) blockchain-alike technology to ensure the untampered communication and authenticity of the original fact source, to relay "base factual" information on the web, would be preferable to my mind.
Though maybe that's because older people are on Twitter, and I'd find the same on TikTok.
I've had numerous instances where ChatGPT has produced output which looks confidently authoritative, but which I as a domain expert can recognize as wrong or even nonsensical. One needs to _heavily_ tune their Gell-Mann amnesia detector when working with LLMs, and recall that if it's not getting the things right you are an expert on, it's likely not getting the things right that you're not an expert on, either.
That said, I think we might have a glimmer of a chance to escape the executioner's axe here, as it were, because while a psychopath is intentionally deceptive, and will become evasive or aggressive when they sense that you're suspicious and probing of them, LLMs have no such quality, and will "happily" keep talking and exposing their shortcomings to those willing to listen.
Ultimately, though, peoples' ability to understand how appropriately to trust LLMs will rest on their ability to understand their own capacity for being too-trusting and easily misled, a feature on which the human psyche does not have a great track record.