Google cautions against 'hallucinating' chatbots
reuters.com
reuters.com
But chatbots take that a lot further...and generative AI specifically cannot really do that in a way that is easy for consumers to make use of - at least it's not possible at the moment.
People without a STEM background won't understand the implications when you say that a biased dataset makes consumer lending metrics like credit scores unfair. But they will understand when their "AI" assistant purchase the wrong airline ticket or generate garbage for a school assignment
But they may not be harmed in ways that don't involve monetary loss, or involve monetary loss that they can quantify and attribute.
> Otherwise if we gatekeep the technology to "professionals", you end up with companies like McKinsey recommending it to governments and civil service and allowing black box ML to insidiously creep into the societal infrastructure.
Honestly, if it's gate-kept, it should be limited to academics and researchers. It's orders of magnitude worse for some institutional process to be built on the "judgements" of an unreliable "AI" than for some consumer to be mislead about something.
It is better that your average Joe wastes 500 dollars over a wrong airline ticket, than to waste 500 thousand in a lawsuit because a government bureaucracy placed too much faith in consultants and efficiency buzzwords.
> Honestly, if it's gate-kept, it should be limited to academics and researchers. It's orders of magnitude worse for some institutional process to be built on the "judgements" of an unreliable "AI" than for some consumer to be mislead about something.
Academics and researchers are not infallible. The current loss of trust in subject matter experts is precisely because they have been held up for too long as paragons of knowledge and judgement. People's lives and livelihoods are impacted when experts make a mistake or offer bad advice. Most of the top academics that are summoned by think tanks and consult for governments are all well tenured/insulated from the worst consequences of their decisions.
Yeah, I guess I could have put it better that it should be banned from production use by anyone, and academics and researchers should only be able to study and improve it.
Ideally the product waits until it can be productively used by consumers before being launched. But the cat's out of the bag and I don't see anything slowing it down.
Whereas these new exciting AI flavored misinformation boxes are more like fortune tellers cold reading, inventing novel half truths whole cloth, telling you anything you want to hear, and actively assuring you that it is 100% true in a persuasive and chipper tone.
So on the front end you might ask it "What is the capital of Italy", but on the back you translate that into "What is the capital of Italy, how could you factually validate that, and give me some example searches you might type into Google to check your work." Then run the searches and return them: "Here are the searches from credible databases, they say X, Y, and Z, does that conflict with your assumption?". Eventually you feed it enough data and provide enough context to the conversation and the AI will naturally head in the right direction. Computers are fast, we should be able to do all this in a matter of milliseconds.
That would require the AI to have a comprehension of facts, and the entire point is these models don't.
So no, I see no evidence that the parts exist to solve this. If anything, the fact the problem isn't already solved suggests to me there is no easy answer.
It's far beyond the scope of a few companies putting a few decades and a couple upstart tech people to the problem to make any real progress on understand and producing truth, which is why ChatGPT answers every question like a middle schooler, because it's just simulating understanding.
I’ve done a little testing of this (using HN comments!) and ChatGPT seems to be reasonably good at distinguishing fact vs. opinion. It seems in that limited testing to be slightly more likely to view as opinion things I would view as likely fact claims but where the operationalization intended is ambiguous as opinion, but the boundary between a fact claim with an unstated operationalization and an opinion claim is fuzzy.
That doesn’t get a lot closer to checking fact claims, or identifying and deciding when and how to report supporting evidence of them, but I don’t think that distinguishing fact claims is, itsef, as far beyond reach of current technology as you are describing it.
The accuracy of this fact is itself is subject to the very issue you are noting though! :)
It won't help for those cases where the facts are right, but the reasoning around them is wrong, I suppose; but critical reasoning is a generally essential skill.
[0] https://www.pcc.edu/library/research/reasons-for-citing-sour...
Statistical generation of text doesn’t work in a way which makes “providing citations” a straightforward thing; It’s not reading a bunch of text, making notecards on ideas, and then using the notecards to build a written response where ideas directly come from sources that are quoted or paraphrased.
“Providing citations” can be a goal for the output, but its not much in the way of an an answer to “how”. And you can ask some of the current tools to provide citations, and when they do it at all, they can hallucinate citations just as well as anything else. (This has been seen in many of the attempts to get ChatGPT, etc., to produce legal arguments.)
Now, checking citations, to some degree of utility, is probably tractable if the cited material is available, as AI’s are good at summarizing text, so comparing summary of a passage with a citation against the summary of the cited material should be possible, but I don’t think any AI system yet has found a way to provide good enough citations (even with some occasional hallucinations) that citation checking to raise flags on the bad ones would be enough to get us to useful results. The challenge is how to both recognize when citations are appropriate and generate good ones a high percentage of the time.
It's a great party trick though.
FWIW, I prompted ChatGPT with "Describe how tail recursion works in programming, and why it is important. Include citations." and it actually included two relevant links in its sources, but did not indicates which part of the text depended on them.
Nor should you, as a consumer. I am just saying that, in the discussion of how to get where we want to be, “provide citations” is much more a description of the destination than the route we are trying to map.
And the ensuing HN discussion has a lot more examples too.
SEO spam is absolutely obnoxious. I just love when I'm trying to look up a local company and the first half of the page are paid ads about local company competitors in another city/state. How useful!
It's been gamed and Google has zero incentive to change it because SEO gaming is ties nicely with paying for ads which is Google's golden goose. chatGPT is such a fresh breath of air.
I feel like it’s contributing to a further misunderstanding of what it actually is and how best to use it or think of it.
It may die down after it becomes clear that AI isn't good enough yet. I've seen fewer people refer to Siri/Cortana/Alexa as "her" rather than "it" as time went on, but there was definitely a peak when they became popular among laypeople for a while.
I notice I'm much politer to ChatGPT than I am to Google, it's quite a funny thing to notice. Perhaps my subconscious is being nice to the robots to be in their good graces when they take over :)
Maybe I'm wrong honestly. Maybe not anthropomorphizing something is impossible for humans (stick a pair of eyes and a mouth on anything and a person would probably give it a name). I guess people _expect_ AI to replace a human in a given role, but the reality is that they will augment existing workflows, e.g. instead of replacing a digital artist, it would be used to spark creativity and basically eliminate creative blocks.
I think the problem is that there is no easily accessible description of what this is. Like super simple.
Statistically close to what you want responses. Big ass database of things that might be similar.
I believe both of those are more accessible to Joe Schmoe, even if they miss the technical description.
Steve Hsu talks a little bit about this on the most recent episode of his podcast (1); he has a startup that aims to fix the problem, apparently with promising early results (2).
(1) https://infoproc.blogspot.com/2023/02/chatgpt-llms-and-ai-ma...
(2) https://twitter.com/hsu_steve/status/1623388682454732801
Survey of Hallucination in Natural Language Generation (2022):
https://arxiv.org/abs/2202.03629
From source: "Once again, intrinsic hallucination refers to output content that contradicts the source, while extrinsic hallucination refers to output content that the source cannot verify."