ChatGPT Is Bullshit
researchgate.net
researchgate.net
Discussion: https://news.ycombinator.com/item?id=40626692
So the paper is pointing out that, rather than thinking of these models as being mostly truthful, but which sometimes hallucinate, it's more accurate to think of them as bullshit all the time.
It doesn't mean that the bullshit is wrong (although it could be). The model itself doesn't care.
[1] https://reasonandmeaning.com/2017/01/23/harry-frankfurt-on-b...
If you look at the wikipedia page for On Bullshit it has:
>Frankfurt determines that bullshit is speech intended to persuade without regard for truth.
Based on that ChatGPT isn't really bullshit. It doesn't really have intent to persuade. It's more like a souped up search engine that takes in a bunch of documents and tries to process the data in them in a way which matches your query. The truth or not out is largely determined by the truth or not of the data going in. And that seems rather good with ChatGPT.
The add glue to Pizza stuff was a different LLM with different input data. It had read a reddit thread where someone had jokingly suggested that. A challenge with these models is to filter bullshit out of the input data though that is an issue with humans too.
- Both people who are lying and people who are telling the truth are focused on the truth. The liar wants to steer people away from discovering the truth and the person telling the truth wants to present the truth.
- A person who communicates bullshit is not interested in whether what they say is true or false, only in its suitability for their purpose.
By this definition, the article seems well reasoned
> We argue that these falsehoods, and the overall activity of large language models, is better understood as bullshit in the sense explored by Frankfurt (On Bullshit, Princeton, 2005): the models are in an important way indifferent to the truth of their outputs
Or write papers and their code from scratch.
This is a philosophical paper, and "bullshit" is a term of art in philosophy, coined by philosopher Harry Frankfurt (1929-2023) in his 2005 booklet On Bullshit, in which he distinguishes bullshit from lies:
> The liar cares about the truth and attempts to hide it; the bullshitter doesn't care if what they say is true or false.
This academic paper argues that ChatGPT produces bullshit in said technical sense. It is nothing inflammatory, superficial, derogatory, dismissive, or obscene.
Again, why is it flagged?
My brother in law who is extremely smart, and very well trained -- MIT PhD -- and exposed to orders of magnitude more data than ChatGPT, given that the human senses ingest gigabytes of data per day, has -- like everyone else -- persistent errors in his thinking. I single him out because of his credentials, but everyone has these.
As a particularly funny example, we were talking about steers and bulls and he was like "Oh, right, I remember, some male cattle are born bulls and some are born steers". My wife, who is similarly educated, believed for a very long time that watermelon's grew on trees. At my undergrad school, a woman who was getting her physics degree and then who eventually went on to get a PhD, fundamentally believed that yeast would not rise if you talked around it -- I guess her parents had told her that and she never questioned it.
That is to say, humans lack full knowledge. We're trained on significantly more data than ChatGPT et al. We still have persistent inaccuracies despite all of that. We're not perfect.
That doesn't make us less useful. Half the time we just rely on other humans trained on similar, but slightly different data, to iron those out, and indeed AI systems that employ those methods get better accuracy.
Now again, my brother in law has fewer errors than most people since he is very smart. However, the vast majority of people are really not that bright (rumor has it that fifty percent have below average intelligence), and yet, simply through verification with other humans, they are able to accomplish useful things.
Sometimes even very capable people have persistent thinking errors. This doesn't make them less smart. No one can claim perfect intelligence. That's the only thing that's bullshit.
[1] https://www.businessinsider.com/google-researchers-openai-ch...
The mistake is thinking that 'data' means only text.
I don't think that's true anymore. GPT-4 is rumored to have been trained on a dataset with >1 trillion tokens. An average human would take a few thousand years to read through that dataset once.
Humans ingest about 74GB of data per day: https://kids.frontiersin.org/articles/10.3389/frym.2017.0002...
That means for a human baby to be as smart as GPT they would have to go from zero to speaking English within two weeks of birth.
Your error is mistaking training data for text only. Humans are particularly inefficient with text because it takes a lot of visual data to get to the thought of a particular word or character due to how human senses work (we have no 'token' sense).
At the end of the day, you wouldn't be bothered if Average Joe would add arsenic to your favourite pie, because AI told it would enhance it's taste, right?
Thing is, Average Joe wouldn't be the one asking your brother about bulls and steers, or your wife about watermellons. They would ask ChatGPT. And 99% it would be all fine, until some orphanage cook would try to improve some dish.
Because human intelligence is not so any system that models it would also not be.
It’s a very long paper that reads to me like an opinion piece designed to get a reaction (and citations).
If you understand limitations of current AI, it is an exceptionally useful tool. Just have realistic expectations and understand how to use the tool for what it has to offer.
If you expect the tool to have all of the answers and get everything 100% right, then you should not be using it at all. But don't preach your beliefs to the rest of us who are able to get tremendous value out of it.
I see this every day. I'm pretty sure most of us who have even a slight interest in AI know the gist of how LLMs work. I'm not sure about what difference it makes in practice.
the best the authors could do was argue that there were some mysterious set of negative social consequences that we'd get from using "hallucinations," since it implies that the models are mistaken or misguided in an attempt at representing the truth, rather than attend to the fact that all the models are doing is generating text that _seems_ truthy, and has no awareness or attention to what the truth _actually_ is. One could probably just read the last paragraph and get all they needed to from the paper:
>We object to the term hallucination because it carries certain misleading implications ... > >Calling chatbot inaccuracies ‘hallucinations’ feeds in to overblown hype about their abilities among technology cheerleaders, and could lead to unnecessary consternation among the general public. It also suggests solutions to the inaccuracy problems which might not work, and could lead to misguided efforts at AI alignment amongst specialists. It can also lead to the wrong attitude towards the machine when it gets things right: the inaccuracies show that it is bullshitting, even when it’s right. Calling these inaccuracies ‘bullshit’ rather than ‘hallucinations’ isn’t just more accurate (as we’ve argued); it’s good science and technology communication in an area that sorely needs it.
So their argument is effectively: 1. it's wrong, or at least, Frankfurt's definition of "bullshit" fits better. 2. it could mislead the public or alignment researchers
On 1), I'm willing to concede that hallucinations might be the wrong term. But words have a life of their own, and it's too late to go back now. At least to late for one paper to change anything.
On 2), it seems plausible, but, regrettably, the paper spends less than a paragraph talking about it! None of the claims they're making are that complicated, and yet for some reason they fail to provide even a few falsifiable hypotheses about the main implication of their argument. Ok, sure. I can see that the term "hallucination" might be misleading. But are you really going to publish a paper just so you can "Uh, actually ..." everybody and argue that we're using the wrong word? How much is it going to mislead the public? The public is always mislead - why does this particular instance matter? How can we tell? Are alignment researches really going to be misled by choice of terminology? If they are, could you suggest a mechanism? How much would they be misled? If they are misled, why does it matter? What does "misled" even mean? How do we measure it? I could go on.
I'd like to imagine that a paper warning about the implications of using incorrect terminology would go into some detail about their claim and explore how those implications might play out - this paper's publication might have more to do with its title and topic than it does any important claims or results.