Google edits Super Bowl ad for AI that featured false information
theguardian.com
theguardian.com
https://web.archive.org/web/20200807133049/https://www.wisco...
The (edited) cheese ad: https://www.youtube.com/watch?v=I18TD4GON8g
What probably should be the target link: https://www.theverge.com/news/608188/google-fake-gemini-ai-o...
The article literally says it's not a hallucination and that the detail came from real websites.
"Google executive Jerry Dischler said this was not a “hallucination” – where AI systems invent untrue information – but rather a reflection of the fact the untrue information is contained in the websites that Gemini scrapes..."
The "hallucination" term generally refers to any made-up facts. Harsh as it may be to put this weight of responsibility on LLMs, users of LLMs generally use them in the expectation that what is says is true, and has been (in some magic hand-wavy way) cross-checked or confirmed as factual. Instead they will print out what is most likely to follow the user's input, based on the training data.
Unfortunately, a vast amount of that vast corpus of training data is social media posts which can't be relied upon to be true. But if it gets repeated a lot then it's treated as true in the sense that "what does 'salary' mean" is generally followed by a billion social media posts saying "it referred to the time that Romans soldiers were paid in salt, because salt was a currency at the time"
> Apparently, the rabbit hole goes deeper.
But it doesn't go deeper than what's already in the article. The article already talks about how the problem is that the internet contains a bunch of misinformation and the LLM is as credulous as the average human, which is to say extremely so.
The article says that's what a Google exec claims, not that that's the actual case. They haven't pointed to any of those websites and we don't have to take them at their word.
Someone further down pointed to a source on cheese.com, where it says gouda makes up 50% to 60% of all global consumption of Dutch cheese. If the source is accurate, the AI hallucinated an incorrect response.
So what you are saying is that Gemini is basically useless?
Good to know. Thanks for clarifying that Google!
With an AI query I get one response back and although people are trying to teach their AI tools to cite sources, this is not universal.
You're looking for black-and-white truth, while the real world is actually more interested in efficiency.
I can spend 2 days scouring documentation and forum posts and experimenting to get ffmpeg or Matplotlib to produce exactly the results I want. Or I can just ask ChatGPT and check if its code works, and if it doesn't maybe spend 10 minutes correcting it so it does, or refining the prompt so it does.
And so you're also missing the part where verifying the correctness of output is very often many orders of magnitude faster than coming up with the output in the first place.
A healthy hop I would say, keeping in mind that this was showing off the best side of it in a commercial. So it just did a Google search, got a wrong factual result, the LLM couldn’t verify it was bogus, regurgitated it as is, and the executive piped in with “acshually, it’s just a plain old search result” somehow not realizing that just makes them look even worse.
I don't think it's fair to frame this is an AI issue. The internet is simply full of misinformation. Should AI outperform the average human and detect misinformation? Maybe. But I don't think that's part of the current value prop, at least not for "non-reasoning" models. If you use an LLM you should be aware that what you're getting is whatever is most commonly found on the web. If that's not what you want, don't use an LLM.
The cheese shop should sue Google for copyright infringement. That's also blatant plagiarism, and Google should get an F for the course.
This is a huge fail for Gemini. Google, of all companies, should know the incentives to distort information aggregators for monetary gain (just look at how bad search results are these days). It's 100% expected that people will try to game LLMs the same way. It's entirely dependent on the LLMs to counteract this.
It was a verbatim copy of the text, which raises the question about copyright. I guess "Wisconsin Cheese Mart" could probably sue Google, but more importantly, this would also mean that it is capable of making verbatim copies of the books contained in Anna's Archive.
That is, assuming that Google also did torrent the books like Meta did, which is very likely.
Funny how they call it "reflection", and not "copy". Maybe that's the way to win in courts: I didn't copy that movie, that was just a reflection of it.
Wouldn't this user of Gemini, who would copy Gemini's response to use this text in their ads, potentially get into copyright troubles with "Wisconsin Cheese Mart"?
If ever there was a company which should know that “multiple websites” is not a good benchmark for accuracy, it’s Google.
This feels like a good parable for Google’s search these days. I’ve seen more wrong information from them than anyone lately. I wonder if they can course correct before it’s too late.
If a human would research the matter they would arrive to the same conclusion by visiting those websites. Essentially, an average person would have made the same mistake when researching unless you have access to specific dairy industry data. We might want to hold LLMs for higher standards than human, but AFAIK every LLM comed with a disclaimer "fact check yourself."
This regard Grok is the best as it gives you the source list for cross referencing yourself.
But apparently this did not smell off to any of the many, many people who worked on the ad. That is the most baffling part to me. Are these people so bought into the hype of the product they're promoting that they just switch their brains off?
The glorified autocomplete failed to detect the nuance in this (admittedly poorly written) sentence that the 50-60% statistic was of the world's Dutch cheese consumption
The sentence isn't poorly written or nuanced, it's simply false.
Out of all Dutch cheese consumed internationally, 50-60% is Gouda. Or, in other words, 50-60% of Dutch cheese exports are Gouda.
At least that’s how I read it.
Taking a poorly written sentence, interpreting it as meaning something incorrect, and then presenting it with authoritative, confident language is very close to gas lighting.
> It turns out Gemini didn't hallucinate a fake stat, Google just copied a website's existing text instead.
https://gizmodo.com/googles-ai-super-bowl-ad-fiasco-somehow-...
I can imagine right now, all over the place, people are being tasked to write some article or provide some stats. They use an AI to do the work for them, lazy people that they are.
Then, their manager plugs the stats into an AI to fact check...
I remember some movie from the 80s, pre-internet. Anyhow, everyone had an implant, and it was decades in the future, no one could even read or write.
Instead, they'd just go through their day and if they were unsure of the answer to something, the implant would search a database and just fill in the info. From their perspective, they couldn't even tell if "what they knew" was in their head, or provided from the implant.
Anyhow, one guy couldn't get an implant and was considered disabled, for he had to learn to read, and acquire knowledge the old fashioned way. He slowly discovered that if he even tried to show people how to read, they were blocked from learning. And if info was provided contrary to "the public good", people simply couldn't understand the concept.
Turns out, some central computer decided what was right and wrong, and those that created it and perhaps once controlled society died... leaving in in charge.
This is what we saw. We saw someone query their implant, and then fact-checkers query their implant, and OK! All good!
Also, I think there were onions in the movie.
https://movies.stackexchange.com/questions/83389/movie-where...
The number of L8+ “leaders” and “drivers” is really jaw dropping.
By copy/pasting another Cheese monger's incorrect product description...
Internet Archive confirms that page has had the blatantly wrong stat up since at least April 2013: https://web.archive.org/web/20130423054113/https://www.chees...
> If truth be told, it is one of the most popular cheeses in the world, accounting for 50 to 60 percent of the world's cheese consumption.
Clearly predates generative AI, so I think this is junk human-written SEO misinformation instead.
Here's that page today: https://www.cheese.com/smoked-gouda/ - still has that junk number but is now a whole lot longer and smells a bit generative-AI to me in the rest of the content.
I've been writing about how the greatest weakness of LLMs is their gullibility for ages. This right here is a great example - see also the Encanto 2 thing from a few weeks ago: https://simonwillison.net/2024/Dec/29/
The true failure here is by the humans who couldn't be bothered to do the bare minimum to protect the brand. This would be a major failure if it where any an sort of ad, but it's utterly unfathomable for a highly expensive and absurdly visible Superbowl ad. Why does anyone involved, from intern to CMO, at either the agency or Alphabet, have jobs that they seem ruthlessly indifferent to performing with any sort of attention or care?
Ka-chow, Google!
I question why a model which “knows” about cheddar / mozzarella cheese would make this blunder.
Was this supposed to be generated by one of those “show your work” reasoning models or is this just the regurgitation of one of the single short response parrot of old quora answer or reddit post ai chatbots?
I asked 4o and Sonnet to "Evaluate the plausibility of this statement: Gouda is the most popular Dutch cheese in the world, accounting for 50 to 60% of the world's cheese consumption".
4o correctly said
>While Gouda is an iconic and highly popular Dutch cheese, the claim that it accounts for 50–60% of the world's cheese consumption is implausible and likely a misrepresentation of its market share. Global cheese consumption is far too diverse for any single variety, including Gouda, to dominate to such an extent."
>Sonnet says "A more accurate statement would be that Gouda is one of the world's most popular cheeses and represents a significant portion of Dutch cheese production and exports. It's estimated that Gouda accounts for over 60% of Dutch cheese production, which might be where the confusion stems from."
Both correctly pointed out mozzarella, cheddar, and paremesan as the actual likely candidates for most popular cheese. So the models are clearly quite capable of this, any bad result is likely just prompting error, e.g. blindly asking for a summary.
It perfectly translated the line, but doesn't it also give me a completely made up 2-column 10-row data table! I asked it why, and the response was along the lines of "I am designed to make your life easier, and I thought providing this table would reduce your workload"
> the Google executive Jerry Dischler said this was not a “hallucination” [...] but rather a reflection of the fact the untrue information is contained in the websites that Gemini scrapes.
What's this guy is describing is pretty much the root cause of what we colloquially refer to as "hallucination".
I thought a hallucination is when it provides fabricated facts because it is more interested in generating another token than being factually accurate? Not trying to argue, just trying to pin down a definition for this term.
Regurgitating false statements is not a hallucination, it's bad training data.
It's the difference between being wrong and just making things up from nothing.
Gee, thank you Google for convincing me you have a product that I find useful and that I can trust. I look forward to you trying to cram this down my throat, against my will, at every opportunity you see fit. :-/
So LLMs like any other computer system suffer from "Garbage In, Garbage Out".
(I'm also thinking of when, in the LLM boat-missed frenzy, they faked an interactive AI demo, to make it look much more responsive than their actual tech was.)
I'm unclear on how either incident was allowed to happen.
A cheese fact on a cheese website was wrong therefore Gemini is bad? What?