chatGPT is a very confident FICTION generator. Any facts it produces are purely coincidental.
Please stop assuming anything it says it true. It was never designed to do that and it provably doesn't do that.
chatGPT is a very confident FICTION generator. Any facts it produces are purely coincidental.
Please stop assuming anything it says it true. It was never designed to do that and it provably doesn't do that.
This is overstated and easily disproved. ChatGPT produces accurate facts as a matter of course. test it right now: "How tall is the Statue of Liberty?", "In what state was Abraham Lincoln born?", etc. There are an infinite list of factual questions for which it produces accurate and correct answers.
It loves to hallucinate api methods that don't exist. It can struggle with indivudual logic problems or questions. But these limitations have clear explanations. It's doing language completion so it will infer probable things that don't actually exist. It wasn't designed to do logic problems so it will struggle with classes of them.
Dismissing it as completely unreliable is unnecessary hyperbole. It's a software tool that has strengths, weaknesses and limitations. Like with any other software, it's up to us to learn what those are and how it can be used to make useful products.
If you ask it question you have no way to know whether it’s right or not unless you already know the answer.
I’ve asked it many questions about my area of expertise, distributed systems, and it was often flat out wrong, but to a non expert it sounded perfectly plausible. I asked my my wife, a physician, to try it out and she reported the exact same problem. Some of the answers it gave had even veered off from just annoying to actually dangerous.
That doesn’t mean it can’t be useful, but using it to answer important questions or using it to teach yourself something you don’t already know is dangerous.
This describes pretty accurately a non-trivial amount of people I have worked with.
- there’s no cited facts
- it boldly proclaims its conclusion
- which you’d only be able to verify as an expert
…so I’m having trouble understanding what you’re complaining about with chatGPT, when that seems to be the standard for discourse.
The comment I was replying to was “the errors you’re talking about are probabilistic if you read the literature” my response is “no they aren’t I have read the literature.”
Note that I’m talking about a specific class of error and proving a negative is difficult enough that I’m not diving through papers to find citations for something n levels deep in a hacker news thread.
Toolformer: Language Models Can Teach Themselves to Use Tools: https://arxiv.org/abs/2302.04761
PAL: Program-aided Language Models: https://arxiv.org/abs/2211.10435
TALM: Tool Augmented Language Models: https://arxiv.org/abs/2205.12255
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models: https://arxiv.org/abs/2201.11903
Survey of Hallucination in Natural Language Generation: https://arxiv.org/abs/2202.03629
Survey of Hallucination in Natural Language Generation only provided promising methods for detecting hallucinations in summarization tasks, which are of course much easier to detect. Searching arXiv for a list of non-reviewed papers that sound like they might be related to the topic at hand is fun debate strategy. But no one else is reading this far into an old thread, so I'm not sure who you're trying to convince.
None of these paper prove your claims about hallucinations, and most aren't even trying to. However, even if the errors that I'm saying aren't meaningfully probabilistic aren't hallucinations.
This does not mean it is a problem that is impossible to fix.
In the end, why do people believe what they believe? The answer is it connects with what they already believe. That's it. If you had a diet of what we call science, you'll have a foothold in a whole bunch of arenas where you can feel yourself forward, going from one little truth to another. If you are a blank slate with a bit of quackery seeded onto it, you end up believing the stars predict your life and that you can communicate with dead people with a Ouija board.
CGPT doesn't have an "already believe". It just has a "humans on the panel give me reward" mechanism, and all it's doing is reflecting what it got rewarded for. Sometimes that's the scientific truth, sometimes it's crap. All the time it's confident, because that's what's rewarded.
Have most people invested time in making sure their web search results are accurate?
It is dangerous for sure, but that's what you got.
That applies to my boss, college professors, Wikipedia, or my neighbor after a couple of beers. It's not designed to give the correct medical steps to diagnosis a brain injury.
I've heard people say they wouldn't use ChatGPT even if there was "only a 1 in a billion chance that it made their bank account details public"...
May I introduce you to some very real probabilities:
If you are living in America there is a 1 in 615 chance that your cause of death will be the result of an automobile accident.
So yes, it is unlikely that we will ever create a tool that can answer with 100% confidence. It is also unlikely that a manufacturing process will result in a widget that conforms to allowed tolerances 100% of the time.
However, in manufacturing this is understood. A defect of 3.4 per million widgets is considered an incredibly high process capability.
These tools are being made more reliable everyday. Please have a realistic goal in mind.
Edit: Well I've learned this much. Many of you are not Beyesians!
The problem is that an LLM can only be as reliable as its training material.
For example I wrote a blog post on the 2 generals problem. In the comments there is more text from inexperienced people asserting that a solution exists than there is text from my original article.
An LLM trained on that article and comments will always be wrong.
LLMs can be trained on smaller vetted training sets sure, but they also currently require massive amounts of data and there’s no guarantee that just waiting a few months or years will improve reliability enough to fix these issues.
There are a number of techniques you could try to improve your results.
But I get the distinct feeling that you are not interested in improving your results as you have a motivation to "prove" something about LLMs.
This becomes a problem when people start to treat LLMs as authoritative on facts.
So, it's sort of like driving a car based not on a complete map of the area, but a partial map where the missing pieces are filled in from statistical averages of things that appear generally in the world-- but not necessarily in that location. We all may be impressed at how close the predictions come to reality, but the only truly reliable thing is that those predictions, measured with precision, will be wrong. And if you rely on them without verifying them yourself, you might have a serious problem.
The “true” way to answer a proposition is not even remotely as simple as you’ve described.
I suspect I'm better at using the tool than most people. So if most people are not good at using the tool and then submitting their poor results to Stack Overflow, I can see the problem.
Acquiring and vetting information is a task of crucial importance situated in between journalism and science. Just trawling the internet and autocompleting the acquired driftwood is a weird approach to begin with. Doing so without transparency in private companies is simply folly.
In my experience with ChatGPT, it gets less wrong and more right than most humans, and it's able to find relationships among facts and concepts that are impossible with Wikipedia and Google.
It's extremely useful and powerful in its current form, in spite of its many limitations.
With other tools, we don't demand perfection to find them useful.
Just because a hammer sometimes misses the nail and puts a hole in the wall, just because Wikipedia sometimes gets vandalized with lies, doesn't mean hammers and Wikipedia aren't essential tools.
However, it provides an excellent opportunity for us mere humans to reflect upon the way we organize our society in regard to vital information processing. We are awfully bad at it. So bad actually, we can't even tell how it should be in the first place and are disturbed already by some piece of software telling nonsense.
They're turning the freakin' frogs gay!!!
And incidentally, it's the consequences of those errors that teach us not to repeat them... which is the same feedback loop we use to train AI in the first place!
That's pretty much the conclusion I've already come to. I have to verify everything ChatGPT tells me, so using it is pointless if I already know where/how to look something up.
We are just going to have ever more persuasive unreliable sources of information.
My policy is that teams should use ChatGPT but that everyone is individually responsible for their tasks. In other words, it's a tool, not a replacement, and if the tool does something wrong the responsibility still resides with the employee to the degree that they could have validated the results, but failed to do so. I think this strikes a good balance that preserves human jobs for as long as possible.
No it is not. If it hallucinates facts in 5% of cases, that is completely unreliable. You're basically saying "yes, it produces facts, as long as you already know the fact and double check it!" You cannot trust it with knowledge-based questions that you don't already know.
I just asked it the height of the Statue of Liberty. I had to then look it up for myself just to see if it was telling the truth, because there's no citation or 'confidence level'. How is that useful?
> What is the volume of the Statue of Liberty?
The Statue of Liberty is a hollow copper statue, and its volume can be calculated by multiplying its height, width, and depth. The statue's height is 151 feet (46 meters) from the base to the torch, and its width is 35 feet (10.7 meters) at the waist. The depth of the statue, or the thickness of the copper shell, is about 2.25 inches (5.7 centimeters) throughout most of the statue.
Using these measurements, the volume of the Statue of Liberty can be calculated as follows:
Volume = Height x Width x Depth
Volume = 151 ft x 35 ft x 0.1875 ft (2.25 inches converted to feet)
Volume = 985.3125 cubic feet
Therefore, the volume of the Statue of Liberty is approximately 985.3125 cubic feet (or 27.86 cubic meters).
It sounds confident, the maths looks correct, but the answer is entirely wrong in multiple ways. It might be interesting to see what prompt you would need to use for it to calculate say the cylindrical volume of the main body?It's actually 990.9375. Curious how it botched the multiplication but still got an almost-right answer (to that multiplication, not to the actual question.)
I will point out that it's frequently useful with knowledge-based questions where it's hard to generate a correct answer, but easy to verify whether an answer is correct.
"the idea that this stopped clock is broken is overstated and easily disproved. the clock produces accurate time as a matter of course. go ahead and ask what time is it, just make sure it is 3:45am or 3:45pm"
Edit: woooooosh...
The broken clock is not the correct analogy.
But I'll play this silly game: ChatGPT is not incorrect 99.9% of the time.
Yes, ask an LLM to multiply a few numbers together and you will get around 100% failure rate.
The same goes for quotes, citations, website addresses, and most numerical facts.
The failures are predictable. That means the models can be augmented with external knowledge, Python or JS interpreters, etc.
My clock analogy works up to this: ChatGPT success in factually answering a query is merely a happy coincidence, so it does not work well as a primary source of facts. Exactly like... a broken clock. It correctly tells the time twice a day, but it does not work well as a primary source of time keeping.
Please don't read more deeply into the analogy than that :)
That’s not even remotely how an LLM functions.
You’re not introducing any scale with regards to correctness either.
It is a poor analogy without any stretching required.
I know it is just a language model. I know that if you took the same model and trained it on some other corpus that it would produce different results.
But it wasn’t so it doesn’t have enough data to say that bananas are larger than the Empire State Building, not that it would really matter anyways.
One important part of this story that you’re missing is that even if there were no texts about bananas and skyscrapers that the model could infer a relationship between those based on the massive amounts of other size comparisons. It is comparing everything to everything else.
See the Norvig-Chomsky debate for a concrete example of how a language model can creat sentences that have never existed.
That is true! But would it be factually correct? That's the whole point of my argument.
The knowledge and connections that it acquires comes from its training data and it is trained for completing well-structured sentences, not correct ones. Its training data is the freaking internet. ChatGPT stating facts are a happy coincidence because (1) the internet is filled with incorrect information, (2) its training is wired for mimicking human-language's rich statistical structure, not generating factual sentences, and (3) its own powerful and awesome inference capabilities can make it hallucinate completely false but convincingly-structured sentences.
Sure, it can regurgitate simple facts accurately, especially those that are repeated enough in its training corpus. But it fails for more challenging queries.
For a personal anecdote, I tried asking it for some references for a particular topic I needed to review in my masters dissertation. It gave me a few papers, complete with title, author, year, and a short summary. I got really excited. Turns out all the papers it referenced were completely hallucinated :)
And if it did stuff like that almost every answer, it would be a broken clock.
But it doesn't. It's usually right about facts. It getting things right is not a coincidence!
Those are easy because you can verify them. But you actually did go back and check to make sure they were right after you got answers out of it, right? And why is that, don't you trust it to just be right?
How tall is the Statue of Liberty?
The Statue of Liberty is approximately 305 feet or 93 meters tall from the ground to the tip of the torch.
When did terrorists attack the statue of liberty?
The Statue of Liberty has not been attacked by terrorists. However, there have been some reports and rumors circulating on the internet about such attacks.
In which works of fiction did terrorists attack the statue of liberty?
I am not aware of any works of fiction in which terrorists attack the Statue of Liberty. While the Statue of Liberty has been featured in many works of fiction, including movies, books, and TV shows, I cannot recall any specific instance of it being attacked by terrorists.
Do you know about the game Deus Ex?
Yes, Deus Ex is a popular action role-playing video game series developed by Ion Storm and published by Square Enix. The game takes place in various locations around the world, including New York City and the Statue of Liberty. In the game, the player must thwart a terrorist attack on the Statue of Liberty by infiltrating the terrorist organization and gathering intelligence.
(end of chat)
It directly contradicts itself here. The summary for Dues Ex is also wrong (terrorists attacked the statue of liberty before the game, and only the very first part of the game features the statue and it's not to "thwart a terrorist attack on the Statue of Liberty").
It can generate very accurate things, and can generates nonsense. However, you're never quite what it is, and there is no "well, I think it was something like this or that, but I don't know for sure". So the only way to know for sure it's accurate is to verify it with another source. And if you need to verify the accuracy then it is actually just inaccurate.
Of course, humans can also generate nonsense, but you shouldn't measure ChatGPT vs. all of humanity; because it's one vs. the worst of several billion. That's not a fair match. Compared to a reasonably person ChatGPT is horribly inaccurate.
I wouldn't call that a contradiction. I think it's fair to only count attacks that go through.
https://vignette.wikia.nocookie.net/deusex/images/d/dd/Statu...
You said it contradicts itself and is wrong. How does it contradict itself?
Also, turns out "The Statue of Liberty has not been attacked by terrorists" isn't even correct either: https://en.wikipedia.org/wiki/Black_Tom_explosion – you could argue a bit about the semantics of "terror attack" and whether it was the statue of liberty that was attacked, but it turns up on all the search results (I actually wanted to find other fictional accounts).
"I wouldn't call that a contradiction. I think it's fair to only count attacks that go through. "
It "knows" about a thwarted plan for an attack. I don't think it's a "direct contradiction" to "know" about that but still say "I am not aware of any works of fiction in which terrorists attack the Statue of Liberty."
It's not thwarted. Look at the picture from the previous comment. Does that look "thwarted" to you? ChatGPT confuses a lot of things here, and now you're confused because ChatGPT is confused.
There is a terror attack, but it's NOT the case that "the player must thwart a terrorist attack on the Statue of Liberty". It's merely something that happened in the fictional setting, before the story even takes place.
In many ways this thread demonstrates my point quite well: it's kind-of roughly correct, but mixes up some crucial things which does make a meaningful different, and now someone unfamiliar with the subject matter is confused as well. In this case, it's just a game and of no real importance, but it goes to show ChatGPT should now be used for anything of any importance in its current state.
The bot is wrong. I know it's wrong. I keep saying it's wrong. You keep giving evidence for why it's wrong. That's not the issue here.
If we play along in the hypothetical where the bot isn't wrong, there is no contradiction. It is telling a coherent story.
But it doesn't really matter; it's inconsistent and unreliable. Whether you could classify this or that as a "contradiction" isn't important and mostly just pedantic argueing.
> But it doesn't really matter; it's inconsistent and unreliable. Whether you could classify this or that as a "contradiction" isn't important and mostly just pedantic argueing.
But it wasn't inconsistent in your conversation. That's my point. It was wrong in a self-consistent way.
There ya go:
What's the backstory of the statue of liberty in Deus Ex?
In the game Deus Ex, the Statue of Liberty serves as a central location and plot point. The statue has been converted into a highly secured base for the United Nations Anti-Terrorist Coalition (UNATCO), which is the organization that the player character, JC Denton, works for.
The backstory of the statue in the game is that it was damaged by terrorists in an attack prior to the events of the game. The terrorist group responsible for the attack is a shadowy organization known as "Majestic 12," which serves as the game's primary antagonist.
In response to the attack, the statue was converted into a heavily fortified UNATCO headquarters, which serves as the player's base of operations for much of the game. The statue is also home to various high-tech facilities and laboratories, where UNATCO develops advanced weaponry and nanotechnology.
Throughout the game, the player learns more about the backstory of the statue and the role it plays in the game's overarching conspiracy plot.
You're now demonstrating inconsistency between conversations. Great. But your earlier claim was that it directly contradicted itself inside that specific conversation. I don't think it did.
(And no, a new conversation where it directly contradicts itself inside the same conversation won't change my mind, because I already know it can do that. I was just saying it didn't in your original example.)
My original post was less than 20 words with no intent to add any more.
The fact that these queries are correct are not because that is the intent of chatgpt’s “query engine” it because it just so happens to have been fed data that synthesizes to something sane this time (vs confident crap any other time).
Compare this to a database where queries are realized against an actual corpus of data to give actual answers.
The purpose of Chatgpt is to provide fluent gibberish when given a prompt. It’s a LARGE LANGUAGE MODEL. It’s only able to give you response that looks good for a language (over a distribution or words), not anything actually knowledgeable.
Wish Stanislaw Lem were still alive and still writing great books, this whole mania around chatGPT and especially articles like this one remind me of his excellent Cyberiad, where at some point two robots from those stories almost get drowned by useless info after useless info (they were printing that info out on pieces of paper). I'm really curious how Trurl and Klapaucius would have handled all this chatGPT non-sense.
I then started probing it for more detail, and pointing out mistaken assumptions, and it got better. But it was first when I pointed out that the technique for encoding and decoding has a resemblance to Lempel-Ziv-Welch compression that it succeeded in identifying the key part of the paper and give a description that while still not quite accurate at least captured the essence of the method.
It felt like I was giving an exam on the topic rather than "getting help" summarising it, and if I didn't know the paper, it would've been hard to tell if I was getting close without reading it.
For conceptually simpler things, though, I've gotten great results, and it seems to work great for things where I know what a good result looks like ahead of time, and "just" want to save looking up details where I'll recognise the right thing when I see it but don't remember it by heart.
That said, to be able to point out the resemblance to LZW and have it connect the dots is at the same time fairly impressive - I've had to point that out to experienced developers (the paper mentions Welch, "This method bears some resemblance to commonly used data compression schemes [Wel84]", with Wel84 reading: "T. A. Welch; A Technique for High_Performance Data Compression; IEEE Computer, 17:6, 8-19; 1984. {2}" but it does not go into any further detail of how they relate).
I think, overall, that this is an indication that it will get better at this as the model size increases further. It was able to at least recognise the semantic similarity or identify that this similiarity has been described elsewhere and use it to guide it's response in a way that strikes me as far from trivial.
[It did go on to hallucinate several relevant papers on the subject; interestingly it confirmed it had made them up when I asked if it had]
But if you’ve got language to give it to work with, results are spectacular.
“What I’m trying to say is…”
I then asked for citations: it gave me what sounded like nice academic journal articles and books, but these were made entirely made up.
Pointing this out just made it give different references, also totally fake.
>The stories and information posted here are artistic^wmechanically generated works of fiction and falsehood. Only a fool would take anything posted here as fact.
https://github.com/williamcotton/empirical-philosophy/blob/m...
It would be nice if ChatGPT would police itself and not say anything that isn't true. That's what people expect of computers, and why we're in the mess we are now.
But then the problem is people disagreeing about whether ChatGPT is a fair arbiter of truth.
Give it a marker and ask it to label the world in handwriting.