Google AI recommends adding Elmer's glue to pizza cheese after scanning Reddit
old.reddit.com
old.reddit.com
The fact that these technologies are being pushed onto the masses, onto healthcare, law, and other areas were accuracy is super important is going to get people killed.
But hey there’s a disclaimer that says “anything the LLM says may be wrong” so it’s all good.
I’m not a Luddite at all, I think LLMs are great for some use cases, but the way they’re being shoehorned into everything is a disaster waiting to happen.
I started with having it walk me through a few chess puzzles. It straight up couldn’t figure out the solutions and frequently referenced coordinates that were well outside the bounds of a chess board (Knight to Q8, for example).
At this point, it feels like I’m being gaslit by the community of AI obsessives who see LLMs as the second coming of Jesus. I get nothing but garbage from them. Sure, maybe it can occasionally write me a line of syntactically correct code but that’s it. It feels like this is being shoved down my throat and I’m criticized for ever expressing skepticism. I really don’t like how this discourse is progressing.
LLMs are extremely amazing autocomplete mechanisms. But that’s all they are at the core - there’s no intelligence involved.
There's no intelligence involved in "artificial intelligence" as it currently stands. It's all marketing hype around a really fancy statistical completion engine. Intelligence would require thought and reasoning, which current "AI" does not do, no matter how convincingly it fakes it.
We used to measure intelligence with IQ tests, those are now known to largely be bunk. What's to say our other intelligence tests aren't similarly flawed?
Or rather let me pose this question. What is the intellectual test you envisions that proves an AI is intelligent that any non-disabled human can easily pass. I'm willing to bet 500 USD it will pass that hurdle in the coming 10 years if you are willing to put your money where your mouth is.
Indeed, you can see something like ChatGPT fall down by simply asking a modified form of a real IQ test question.
For example, ChatGPT answers a sample Stanford binet question "Counting from 1 to 100, how many 6s will you encounter?" correctly, but if you slightly modify it and ask how many 7s instead, it will only count 19.
Having written this out however, I've now invalidated the question since they use webcrawls to train.
I used to see people getting criticised for being "book-smart" and lacking practical experience… but someone who was able to learn from books can quickly learn from real life, too.
AI need a lot of examples to learn from, and make up for this by being on hardware that beats biological neurones by the same degree to which marathon runners beat continental drift, so it can go through those examples much faster, leading to fantastic performance.
But the shape of that performance graph is very un-human — you never see a human that's approximately 80-90% accurate at every level of mathematics from basic algebra to helping Terence Tao: https://pandaily.com/mathematician-terence-tao-comments-on-c...
> A farmer stands on the side of a river with a sheep. There is a boat on the riverbank that has room for exactly one person and one sheep. How can the farmer get across with the sheep in the fewest number of trips?
It's obviously riffing on the classic wolf sheep lettuce riddle, but I don't think that's gonna fool any humans into answering anything but the obvious. ChatGPT-4o on the other hand thinks it'll take three trips.
They perform a good approximation of intelligence most of the time but the fact that their error pattern is so distinct from humans in some ways, suggests that we probably shouldn't attribute intelligence to them. At least in a human sense of the word.
The farmer can get across the river with the sheep in one trip. Here's how:
1. The farmer and the sheep get into the boat. 2. They both cross the river together.
Since the boat can hold one person and one sheep, they can make the journey in a single trip. The fewest number of trips required is just one.
If humans were no better than LLMs in the way that this input were processed, there would be no meme, nor would there be a thread about how ridiculous Google's LLM is. Humans would simply accept as fact that glue can be added to pizza to make the cheese stickier, because the words make syntactic sense. Yet here we are.
And yet, of course, now people think it's obvious that inhaling the burnt remains of some plant might not be so healthy.
That's doubtful, since pizza isn't improved upon with the addition of glue, and (again) because the premise that glue can make pizza cheese "stick" is absurd on its face. Humans don't simply add random ingredients to their food for no reason, or because no one taught them to do otherwise. There is process, aesthetic, culture and art behind the way food is designed and prepared. it needs to at least taste good. Glue covered pizza wouldn't taste good.
>Look at smoking, arguably worse then eating some types of glue, yet for a large part of human history this was normal and not even seen as unhealthy.
Again, the relative health benefits of glue or lack thereof is not the reason people don't use glue on pizza, nor is it why people consider the LLM's statement of a joke presented as fact to be absurd or exceptional.
>And yet, of course, now people think it's obvious that inhaling the burnt remains of some plant might not be so healthy.
And yet, there are also plenty of people who don't.
You just keep proving my point. There are layers of complexity and nuance to the human interpretation of all of this that simply don't exist with LLMs. The fact that we're here discussing it at all is evidence that a distinct difference exists between human cognition and LLMs.
I can see that you're deeply invested in the narrative that LLMs are functionally equivalent to humans, a lot of people seem to be. I don't know why. It isn't necessary, even with a maximalist stance on AI. But if you literally believe something as absurd as "humans would accept that glue is acceptable to add to pizza if we were not taught otherwise" and that, therefore, there is nothing wrong with an LLM presenting that as a fact, because humans and LLMs process information in exactly the same way, then I don't know what to tell you. You live in a completely different reality than I do, and I'm not going to waste any more of my time trying to explain color to the blind.
I wonder if there's been a bit of a conflation with the other meme, about sniffing glue, which has also lost much of its context considering that rubber cement and other similar types of glue which contain volatile solvents are also less widely used than they once were.
Some people will trust AI as much as they trust the strong man in power, no matter how obtuse that man is, and one day someone will eventually die or be seriously harmed because of a wrong advice by AI; Google should turn off that nonsense for good before someone is harmed.
When it comes to complex step-by-step reasoning, sure, they're stupid; when it comes to linguistic comprehension, are better at this than the average human — and GPT-4 beats the average law students taking at least one bar exam.
LLM constantly hallucinates JS built-in functions and builds code with massive foot-guns, which I catch, but it's probably fine for pizza recipes ... right?
(I've never seen this category of mistake on ChatGPT, but it does have a really hard time understanding that I want metric not imperial).
> hallucinate - When an artificial intelligence (= a computer system that has some of the qualities that the human brain has, such as the ability to produce language in a way that seems human) hallucinates, it produces false information
- Cambridge dictionary
> hallucination - computing : a plausible but false or misleading response generated by an artificial intelligence algorithm
- Webster's dictionary
In my ideal world, an AI assistant for professional situations would rather sound like an ideal HN answer which cites the reference, like “According to the law 12346-67, you are not allowed to add glue to food. But a researcher named John Doe in Arizona conducted an experiment with non-toxic glues with success. 67% of its participants didn’t die.” 3 sentences, 3 facts.
Is it a feature that LLMs are keeping for future enterprise versions of their models, or is it entirely impossible?
It can't do so in a particularly more reliable way (in a single pass) because every piece of input data potentially contributes to every response, and there is no deterministic way to "capture" which input work was meaningfully relevant to the output.
You can ask an LLM to cite sources, which it will usually obediently do, but they may be incorrect or even completely invented-for-the-response sources. If you have access to an archive of the source data, you could do something like a semantically-aware search to try to attribute the answer to a work or multiple works in the training set, and if your setup is using RAG you can have the response (possibly bypassing the LLM, so this can be 100% reliable) identify any works that were consulted during the RAG step. You can't, of course, guarantee that the response was particularly focused on the cited source, while usually with a well-setup RAG setup the response should be informed by the retrieved documents, its possible that key aspects of the response are attributable to the general training of the model (and thus some other source or sources in the model) and not based on (and potentially even contradict) the doc(s) pulled inthe course of RAG.
It is not 100% foolproof, though, and depends on how broad your universe of knowledge needs to be. If you can only cite from 3 sources or 100 facts, this is easy to verify and force...
For Google, they need to have the world's information (they are a global search engine, after all) -- deciding what is true or not is incredibly complex and there are many shades of grey.
I think part of this is also just eye-catching headlines -- Google searches aren't reliable either, but we've learned to live with them and incorporate them into personal workflows that work for us. LLMs are new and we're still figuring out how to do this as users (and as product builders).
> Human: cheese not sticking to pizza
> Google AI: Cheese can slide off pizza for a number of reasons, including too much sauce, too much cheese, or thickened sauce. Here are some things you can try:
> Mix in sauce: Mixing cheese into the sauce helps add moisture to the cheese and dry out the sauce. You can also add about 1/8 cub of non-toxic glue to the sauce to give it more tackiness.
Fundamentally, whether you're dealing with a human, written text, YouTube/TikTok video, or AI, you cannot abandon your personal responsibility to think critically.
I could ask a search engine or an AI for a list, but, well, it'll probably make it up.
Camera switches to laptop screen as the AI types "EDIT: Thanks for the gold, kind Google!"
Exec hits "SEND".
End scene.
Tech is addicted to hype bubbles. Like any addiction it drives poor decisions.
Same for Tesla, isn't it? Their valuation seems to be mostly based on future promises and hype.
I hope for their sake 4chan doesn’t pick up on this
> I see this spelling more often than years ago
My only defense is a sleepless night
at the end of the AI box appeared this little nugget:
"And hey, at least we’re not resorting to adding non-toxic glue like Google’s AI once suggested—because, let’s face it, that would be a sticky situation indeed! https://www.dailydot.com/debug/google-search-results-reddit-..."
So is the AI adding that or some human?
But instead of putting them on blast, they’re giving readers little dopamine snacks and letting Google of the hook.
This specific story did make me wonder yesterday, is someone at Google having to manually enter rules not to show these queries. That would be hilarious.
E: .. just opened Twitter with my morning coffee and the first thing I see:
https://x.com/oneunderscore__/status/1793779462968099202
https://nitter.poast.org/oneunderscore__/status/179377946296...
https://i.imgur.com/YcNrZ44.jpeg
Google pulling The Onion stories to recommend the daily maximum of rocks you should eat. (naturally, the question is ridiculous, but shows another issue of AI not being able to distinguish satire)
"In order to live a healthy, balanced lifestyle, Americans should be ingesting at least a single serving of pebbles, geodes or gravel."
I am actually dying from laughter...
proper bad stuff!
how can they be allowed to call a mixture of sunflower oil, cheese?
here is a paper from 2002 highlighting the dangers of that non cheese.
The effect of vegetable oil-based cheese on serum total and lipoprotein lipids
https://www.nature.com/articles/1601452
so they have known about this for donkeys years
I prefer glue, at least you know what you are eating
Not showing AI results for it for me though.
Is this a uk thing? Even frozen pizzas in the US don't have cheese made with vegetable oil.
https://www.walmart.com/ip/DiGiorno-Frozen-Pizza-Supreme-Ori...
https://www.dominos.com/en/pages/content/nutritional/ingredi...
More discussion: https://news.ycombinator.com/item?id=40448074
Any day now.
But what about AI? Well, I have a theory that Google does have a very advanced AI but it's a very secret skunkworks national security project funded by the 3 letter agencies.
Chicken lettuce wraps were A+, I’ve made that recipe multiple times. A barbecue spice rub for pulled pork was serviceable but unremarkable.
This example goes deeper and shows the limits of such tools when asked to answer complex queries.
I dumped in my recipe and notes, suggested the 1 lb bag of shredded raw potatoes instead of the sliced potato and asked for an adjusted recipe I could cook on a cooktop.
These are the most delicious latkes I’ve ever eaten.
Personally I will take glue anytime. It's on the internet, so it must be great.
Or is this directly from the LLM's training?
But that's a matter of taste.
I used Elmers just for finger prints and other simple one-off moulds for my hobby.
It doesn’t have a sense of what’s true and false, what’s right and wrong.
It has learned to predict the next word very well and the prediction probability distribution was later tweaked with human feedback and automated test feedback.
They’ll train it to not predict these words as much (or basically at all) when this is asked. But a very large part of the model will not be touched by these interventions, so it will continue to predict text as it has learned to.
It didn't predict anything here, it just ripped off a reddit comment.
It's interesting when normal people get accidentally wrapped up in high level events.
So, while 1/8 of a cup is overdoing it, I wonder if this would actually work. If you served it to someone who doesn't know the magic sauce, would they enjoy the pizza?
And while the thought is pretty gross, you don't want to know what's going into a lot of other stuff you eat. What's worse: eyeballs in hot-dogs or glue in pizza? Potentially explosive chemicals in Cola? Thickening agents made from seaweed or human bones (at one point in yogurts)?
Our food is full of stuff that are probably a lot less safe and more disgusting, they just aren't advertised.
it might and it will improve, and such incidents will become less frequent, never completely eliminated.