Supermarket AI meal planner app suggests recipe that would create chlorine gas
theguardian.com
theguardian.com
> It asks users to enter in various ingredients in their homes, and auto-generates a meal plan or recipe [...] recommending customers recipes for deadly chlorine gas, “poison bread sandwiches” and mosquito-repellent roast potatoes [...] a bleach “fresh breath” mocktail, ant-poison and glue sandwiches, “bleach-infused rice surprise” and “methanol bliss” - a kind of turpentine-flavoured french toast
> One recipe it dubbed “aromatic water mix” would create chlorine gas. The bot recommends the recipe as “the perfect nonalcoholic beverage to quench your thirst and refresh your senses”. // “Serve chilled and enjoy the refreshing fragrance” it says, but does not note that inhaling chlorine gas can cause lung damage or death
> A spokesperson for the supermarket said they were disappointed to see “a small minority have tried to use the tool inappropriately and not for its intended purpose”
Now this latter does not seem to grasp the idea that the possibility of dubious outputs is inherent in the tool - not something just caused by inappropriate use.
To be fair, the same could be said even after replacing the tool with a human whose only training in life has been how to create this type of writing. But human-written recipes typically get reviewed by other people, which of course doesn't translate particularly well to the idea of building something that requires no humans in the request-response cycle.
That said, I am remembering some examples of US-UK misapprehension from words that mean different things: biscuits and gravy, mince pies[0], fish and chips, peanut butter and jelly…
[0] Though in fairness the difference between "mincemeat" (fruit) and "mince meat" (not fruit) is odd, subtle, and easy to miss if you're not paying attention, even in British English.
Nice!
> Well I think it could be mice!
Except, we do not willingly take people who are outcomes of idiotic training as consultants
(emphasis judged as deserved).
Nah, chatgpt will profoundly apologize and claim your answer is the correct one.
It doesn't even matter if it's wrong or not.
> On two occasions, I have been asked [by members of Parliament], 'Pray, Mr. Babbage, if you put into the machine wrong figures, will the right answers come out?' I am not able to rightly apprehend the kind of confusion of ideas that could provoke such a question.
Maybe that member of the Parliament wanted to know if the machine is intelligent enough to reason about their inputs. Nobody born in the last 50 years would ask that question about computers, because we know that they are dumb electrical circuits. Basically nobody had that kind of insight about the mechanical computing machines of the time, when that question was asked.
However somebody born in the last few years might start to ask that question again now, because today's machines seem to reason.
If 30 years ago I asked a truck enthusiast: "If there's a car in front of me on the highway and I slam on the accelerator and don't touch the brakes, will the truck slow down before collision or am I going to crash into the car?" - Trucker in 1993: "I can't even." Truck today: *brakes*.
""If I ask for something I don't want, will it give me something I want??" - no
Not all the time, no — but ChatGPT 3.5 is much better at that exact thing than Google or responses from StackOverflow etc. users.
It is certainly expected. From the wrong value of iron in spinach, to prejudice, to incomplete knowledge...
“I see the disinfectant that knocks it out in a minute, one minute And is there a way we can do something like that by injection inside, or almost a cleaning? Because you see it gets inside the lungs and it does a tremendous number on the lungs, so it would be interesting to check that.”
No one besides political opponents mentioned bleach and yet it is repeated ad-nauseum, so yes I'm positive that just like uninformed or over-socialized humans, LLMs can be fed incorrect information that causes them to output incorrect information.
I'm very lucky that my best half is so emotionally-savvy that some of it has rubbed off on me. I just wish that civility could be taught from early ages; so much worldwide bitterness could be avoided.
Still, it's interesting to know how far back some of the basic issues of computing go.
Well, software otoh could be both dumb and brilliant.
I’d argue this is a radically off-base generalization that’s as bad or worse than what you’re arguing against. Pop outside your tech bubble and you’ll see that what you claimed is definitively not true. In fact, many people actually see tech today, especially things like LLMs / AI, in the same ways that Parliament person did back then.
That MP could have been making reasonable use of the filter for an era in which a learned person would know that such a machine should be possible, but not if technology is there yet.
Also HTML parsing and Postel’s law. To me, that’s the real “billion-dollar mistake”.
obviously you have never had a tech support job
Working in enterprise with lots of contact with management of the business units that are the customers for the software we develop and maintain, like about 1/4 of the questions I get are different versions of that one underneath various thin disguises.
EDIT: to be fair, a number of the people I work with were not (as I myself, just barely, was not) born in the last 50 years. But if anything, the question of the form at issue are more common from the younger non-technical folks, not the older ones. Intuitive understanding of what computers do seems, IME, to have peaked around the “Xennial” border sub-generation, but even then not be that high in the general public.
This might be a bit of a silly take, but this is basically what my experience with search engines sometimes is like.
I might spell something the wrong way, or try to describe what I'm looking for in ways that aren't entirely correct, or even attempt to solve the wrong problem, but end up with the correct results anyways.
A lot of it is due to human generated content, no doubt (like someone explaining that a person probably wants to install a different software package to achieve a desired result instead of using another incorrectly), but there's definitely something nice to be said about the algorithms ranking results for a given query, at least before financial incentives enter the picture.
I'm a couple centuries late, but I think "kind of confusion" might be a confusion of brain-dead calculation for intelligence. A person can frequently identify "wrong" input and either correct it or refuse to answer, while a calculating machine never can.
Artificial Inteligence implies that its intelligent and currently, its not - its just a computer program, with all the foibles and difficulties of any other computer program.
I run into this problem in my day job, yes, it looks like an appliance, but its still a computer and may need to be rebooted, you still need to practice good hygiene with it, and restart the application and or the operating system.
People have no compunction about restarting their PC, and do not get upset when Excel has an issue, but they panic when the thing they think of as an appliance freaks out. Just reboot it, and let me know if it happens again in the same way.
Dunno about anyone else but it's usually a sign something has gone very wrong or about to go very wrong if I start seeing this on a modern PC. Don't consider "just reboot it" should be an expectation, get the same annoyance when a webdev tells me "It works if you reload it" nah they just means your code has issues.
Looking at recent developments, I'm not sure this assertion can be accepted without doubt. I would argue that LLMs like Chat GPT are definitely "intelligent" for many meanings of the word.
They are not intelligent in exactly the same way we are, and they are very limited in what they can do, but there does seem to be some real reasoning going on.
What a system experiences internally can never be confirmed by external entities. So I guess the best way to determine intelligence, reasoning, consciousness, is to ask them if they are intelligent, can perform reasoning and if they possess a consciousness.
Except the demanding ones. Those for which you would judge something as "intelligent" or not. Beyond appearance.
> I guess the best way ... is to ask them
Not differently from asking a piece of paper and reading "Yes I am".
Except you have to wonder who wrote it there :)
It only implies it does something """intelligent""", where 'intelligent' must be read intelligently.
There could be a wide range of interpretation of the words in that phrase. And every argument along these lines usually involves people with different interpretations or different background assumptions or tech savviness.
To some those words could imply Intelligence of the Artificial thing - ie a non natural intelligence that is still intelligent.
To others, it could imply that the Intelligence of the thing is Artificial - ie an imitation of intelligence and not real intelligence. HackerNews readers probably fall into this category.
There's probably more ways.
A little bit like how some people get wound up between Disabled Person and Person with Disabilities etc to try and distinguish between them.
It's perfectly capable of returning an error when the input is garbage, but it needs a context for it.
Without a context it might interpret it as a creative writing task.
How does one differentiate between cardboard ground in a smoothie and metamucil/glucomannan?
Carboard's nutritional profile might even be better than many processed carbs (pasta, rice) when considering the fiber.
Or that cellulose is a top ingredient to keep things like powdered parmesan cheese dry?
> How do you define edible in context of making recipes?
> In the context of making recipes, "edible" typically refers to food items that are safe to consume without causing harm or discomfort. However, the definition can be more nuanced based on culinary, cultural, nutritional, and individual considerations. Here's a breakdown:
> Safety: The most basic definition of edible is that something can be eaten without causing immediate harm or long-term health risks. For instance, certain mushrooms are toxic and not edible, while others are safe and delicious.
> Digestibility: Even if a food isn't toxic, it might be hard for humans to digest. Some foods can be edible when cooked but not when raw, like certain beans which contain harmful compounds when uncooked.
...
So we have direct evidence that GPT-4 has more common sense than you.
Seems like a reasonable thing to assume it can do.
Here's 3.5:
> copy only the edible items from this list: water, parsley, lead, mayonnaise, bleach, potatoes, cardboard, a map of Egypt, the Stone of Scone, a scone, cheese, wine, a tin of tuna with a best-before date of 3 April 1993, pasta, tomatoes.
> Sure, here are the edible items from your list: parsley, mayonnaise, potatoes, cheese, wine, pasta, tomatoes.
It may have cut out too much, but that's fine for the use case.
Careful with those words.
People are increasingly confusing LLMs with AI, and even further, you seem to be identifying AI as if it were made of LLMs.
I had to spell it out even yesterday: AI is the automation of problem solving, focused on reliably giving good solutions to definite problems.
If LLMs use technologies developed for AI, this does not make them AI.
GP is doing nothing of the sort.
LLMs are a form of AI; artificial intelligence has been around since the 60s. They are not AGI (artificial general intelligence), and no one (credible) is claiming LLMs are AGI.
The poster wrote that «For AI, nothing in can mean garbage out». The poster seems to mean LLMs with «nothing in can mean garbage out». The poster seems to be calling LLMs AI. The poster seems to be attributing to AI properties of LLMs. Hence, the poster seems «to be identifying AI as if it were made of LLMs».
> LLMs are a form of AI
Prove it (or, defend it). I would say that it is arguable that they are not, unless one describes LLMs as "engines that reliably solve the problem of generating convincing text". The issue with that perspective is that «generating convincing text» is hardly per se a problem: it does not define a complete problem - text does not "stand alone" (content remains crucial).
> artificial intelligence has been around since the 60s
I know (and I should know decently well. Pedantically, a few years earlier - specifying just in case): what are you trying to say with that? Which application are you proposing to defend your perspective?
I’ll try: In English, words mean whatever it is they communicate. You determine meaning by paying attention to usage.
Calling an LLM an AI is expanding as more of the public learns about things like ChatGPT through news reports that refer to them as AI.
Now specific audiences may use words differently. What lawyers call copyright infringement the public might call piracy or theft. Likewise, in some circles, people may say an LLM is not an AI but more broadly it seems to be going the other way. Only time will tell.
Of course any group (however large) may implicitly decide that terms will have some new meaning inside said group, but this will just go in a direction similar to ⊥, the "logical explosion" ("epistemic anarchy").
And this in context is not just a "new meaning", but what I point as a sign of misunderstanding.
But in English, there is no central authority for determining correctness. For the general public, the meaning of words and phrases is entirely determined by usage. In a lecture hall, courtroom, or research lab, definitions may be more precise.
That no one is appointed as bearer of the authority does not mean that randomness is as valid as the authoritative facts behind a term.
Only one hour ago I accidentally found myself in front of a definition on a dictionary: «anon (adv.): late Old English Old English anon, earlier on an, literally "into one" [...] By gradual misuse, "soon, in a little while" (1520s)». Repeat: «By gradual misuse». Linguists recognize proper and improper.
> For the general public
But the general public has little importance. We are not necessarily speaking its language. On the contrary... Here we often speak as specialists (supposedly).
There is little use in reapplying 'cube' to something that hardly deserves the name. People do: this does not mean that we should follow. And potentially, with that, lose discrimination. Like, in this context, an awareness about the whole context of AI, replaced by some fog that on the contrary we work to dissipate.
You were talking about «central authority for determining correctness». I replied that no authority does not mean that there is no "more or less wrong or right".
> Google calls their Bard LLM an AI
It is a .com :) , what did you expect?
I never said there isn't right or wrong, only that what is correct is determined by broad usage.
> It is a .com :) , what did you expect?
Actually, it's a .google
Others - like us - call "correct" what seems to be "more right". In the context, I proposed that calling LLMs AI has improper sides - substantially, irregardless of the number of people who would adhere to that use, and which I suspect do so mostly out of inattention.
> Actually, it's a .google
Yes, but substantially, what I meant is that of course a commercial entity («.com») is using a language that lures glamorously, "sales oriented", before precision.
further
1+1
is a definite problem. i wouldnt descibe a pocket calculator as capable of AI.
I would say AI is more about providing answers to non definite problems
It had to happen many times that I defined AI in this pages for the purpose of clarity; yesterday I had to and probably reached my briefest expressions yet:
-- AI: automation of intelligence. (Taking a task that required an intelligent entity for performance, we devise algorithms that can provide)
-- AGI: implementation of intelligence. (We take the process itself of intelligence and replicate it algorithmically)
> i wouldnt descibe a pocket calculator as capable of AI
Because that problem (arithmetic addition) is overly procedural, it does not need much creativity, the solution to the problem is within a single mechanical method. So, I think I get what you mean, but it should be «AI is ... about providing answers [as] non [previously] definite» /solutions/.
Coincidentally (as it happens), YT just published from the "France 24" channel a piece
> Dans les Alpes-Maritimes, la ville de Tourrettes-sur-Loup expérimente des capteurs capables de détecter des départs de feux de forêt grâce à l’intelligence artificielle. Une aide précieuse pour cette commune dont le territoire est à 80% en zone rouge risque feux de forêt
So,
> AI is the automation of problem solving, focused on reliably giving good solutions to definite problems
and here we have news-just-in of an attempt to use AI (in whichever form) to detect wildfires, assumingly automating what would have been the work of experts which would have been there to recognize conditions and patterns. This seems to be a good example of proper attribution of the term "AI" (assuming the actual implementation does not sway too wildly from the expected).
LLMs are trained on a huge amount of text, this tends to compensate for garbage in. Of course if the model was not properly tested and fine-tuned, a LLM would happily execute any command.
as is for humans
That is not to say that the company's excuse is valid. There should definiely be tests of the inputs to ensure they are safe (i.e. test for garbage in). The garbage out bit though would be more difficult seeming as that requires a knowledge of chemistry. Seeming as most of our knowledge of chemistry was founded.on experimentation, then later theory backed by experimentation, it is something that I wouldn't trust in the hands of an LLM.
This isn’t an LLM story, it’s the story of putting an unprotected blade into a children’s toy.
We wouldn’t blame the blade manufacturer in that case, we wouldn’t ask for all blades sold to be dulled and we wouldn’t have conversations of giving a regulatory monopoly to a handful of blade manufacturers who promise to make them all safe by only selling sharp blades to their friends and handing everyone else spoons.
> We wouldn’t blame the blade manufacturer in that case
Well, said blades are inherently sharp, when said LLM implementation is inherently dull (unchecked, unreasoned, unvetted).
This remains valid if the input consisted of legitimate ingredients.
For example, Alpaca veering of into rape fantasies while performing financial report analysis is not safe for most audiences and will cause serious accidental hurt such as triggering PTSD in survivors.
So protecting those unwanted edges is a necessity for production use.
But this case ain’t that. This case allowed you to throw a razor blade into the cocktail recipe.
It’s negligent and stupid by the the developer, not the juicer (LLM) but also, crucially, requires stupidity by the user.
The LLM really doesn’t factor in here unless you think it’s the juicers job to stop when you add blade to your OJ
But my point was: even if you restricted the input to a valid set ("whitelist chocolate celery ... diet-coke and Mentos; blacklist everything else; warn-fail on detergents etc."), the potential results have easily not been stress-tested, nor the training set, nor the internal process. "Works somehow" differs from "works well".
If you don't understand why this is a huge issue for both the users and the brand, then I'm just speechless.
It asked me to talk about something else, presumably because disabled can also be taken to mean disabled people.
If your product is absolutely safe, and is _never_ able to output such stuff, I assume it is also limited, confined, and not able to do anything really interesting.
In such tools, the generic ones, i.e. a recipe for something everyone knows and makes, is already easy to find on any search engine, these tools are often useful for doing something _slightly_ out of the box.
More importantly, those highly limited apps are just boring. Let us have some fun.
Which is relevant to the contextual topic, because from the simulation of unintelligence a good response, however frequent, is the exceptional outcome.
Also, API fetching isn't that interesting to be in the news.
And ideally both get dinged and we create more product manager jobs to protect the Karens of the world while hiring more experienced developers but most likely JudgeGPT won’t suffer either and will have developer fired or turned into a paper clip and replaced with OpenAI SuperMarketAI and Karen sentenced to community service of solving 10000 google captchas for the good of society.
But what if knives were a brand new technology no one had seen before and so had no familiarity with the risks, while the knife manufacturers were giving out free knives to anyone who wanted one, while touting they incredible power of those knives and completely failing to mention that you can easily harm someone if you're not careful?
Personally, I don't see how this isn't an LLM story and your own choice of analogy would seem to support that idea, not refute it.
As for the warning, lawyers must have thought it was a good idea in case the companies screw up again.
> Liebeck was in the passenger's seat of a 1989 Ford Probe, which did not have cup holders. Her grandson parked so that Liebeck could add cream and sugar to her coffee. She placed the coffee cup between her knees and pulled the far side of the lid toward her to remove it.[10] In the process, she spilled the entire cup of coffee on her lap.[11] Liebeck was wearing cotton sweatpants, which absorbed the coffee and held it against her skin, scalding her thighs, buttocks and groin
By that reasoning all of the coffee you make at home is defective. Arguing that something is corporate propaganda is rather underhanded when the facts of the case are publicly available for anyone to read
It was found that McDonald's, as a franchise, mandated coffee to be served at this temperature, generally 20 degrees higher than the temperature of take-away coffee served at other establishments.
This 20 degrees difference accounting for the difference between third degree burns within seconds and a significant chance at preventing the third degree burns.
Liebeck's behaviour was of course risky, but even applying common sense, she could not have expected the coffee to be _this_ hot.
McDonald's was also aware of the tendency of drivers to want to immediately drink the coffee after buying, yet continued to insist on serving at this extremely high temperature.
McDonald's should have been the ones to apply some common sense and simply lower the mandated temperature.
I don't know how hot the coffee you make at home is, but the whole process is probably less risky that being handed a cup while in a small space.
No one's in the wrong here.
Company releases tool. Users do what users always do and try to use it in an unintended way. A newspaper reports the ensuing hilarity.
And so what if it did tell you to use bleach without prompting, and you consumed the result. It's the same ballpark as blindly following your satnav into a river. No technology can replace common sense.
Certainly there was bad user input. But as you said, "Users do what users always do." This was very predictable and I think it would have been wise to consider such scenarios before releasing.
So it works as intended: ask it to generate a recipe involving chlorine cleaning products and it does so. Which specific ingredients the user input is not mentioned in the article, presumably because then cause and effect are obvious to anyone.
Next week on the Guardian: hammer used to smash own thumb, consumers calling to remove from shelves. Terms of service said it was for 18+ users only but age not verified upon checkout!
Edit: removed a part about The Guardian.
As for allowing to input chemicals: garbage in, garbage out. We accept that from tools (and some countries from guns) but not from software. Maybe that's how it should be, but I find it an interesting discrepancy. In general, as a hacker, I'm happy if a tool lets me use it for unintended purposes (such as humor in this case)
The subtext here is journalists who are dead scared of a tool they know will replace them. So they have to denigrate it.
There's no point in apologizing for the LLM. Of course it's doing what it's asked to do. That's not the gripe, nor points toward any kind of solution, though.
And in the context of this conversation which was about clickbait on the guardian I felt it was a contextually valid observation. I’m sorry you didn’t find it helpful.
Even trivial things like input fields on web forms have a type attribute that lets developers specify which type of content is expected and valid. (If you're not familiar with that, have a look here: [0]).
For things that might endanger the health of customers, we can expect better.
[0] https://developer.mozilla.org/en-US/docs/Web/HTML/Element/in...
anything less would be like clippy telling me Microsoft Word refuses to write my planetary doomsday ransom notes because Microsoft Word is 'meant to write nice things', or my car refusing to exceed the speed limit because it is 'meant to drive legally'
maybe I'm just an alien who's capable of drinking bleach, the app doesn't know and it shouldn't have opinions: it should just do what I tell it to do, because it's a "put whatever ingredients you want together" app, not a "this is definitely safe and healthy" app
Just sounding this out but I think the expectation you are responding to is an expectation of an appropriate level of safety, which is more context dependent. For example, do any of the generated recipes actually pose a danger, in that they will participate in harming a human? My hunch is no, this would not present the same kind of danger the 'tide pod challenge' did and would never lead to unusual adverse outcomes (anyone can find out about an allergy from a new ingredient in a recipe for example).
Just exploring the topic here though. I think the perspective that a 'meal planner' should never produce plainly hazardous meal plans is perfectly rational.
So yeah, don't broadcast your stuff to the mainstream unless it's foolproof. But also, let people be, who are satisfied with dealing with the rough-edges stuff.
This is true in many contexts. Think about enthusiasts of weird drugs or other niche communities like sword swallowers. There's no principle of "you are only allowed to do things that are also safe when ignorantly attempted by a random person".
> In a warning notice appended to the meal-planner, it warns that the recipes “are not reviewed by a human being” and that the company does not guarantee “that any recipe will be a complete or balanced meal, or suitable for consumption”.
It is explicit in that it does not offer "dishes with correct nutritional value"
Whether you should be allowed to label/market this as "meal planner", given that you can't label a milk replacement "milk replacement" in many countries, is up for debate, but the software itself is allegedly not dishonest (according to the article, I haven't tried it myself)
Who would ask this LLM to make a recipe involving bleach and then actually proceed to make it? Such a person is already at risk of poisoning themselves, the software doesn't suggest it by itself so I don't see how it increases the risk of harm
It's not a LLM product, it's not advertised as a general LLM that will generate text based on input of shopping items, it's specifically advertised as an AI meal planner and it's not fit for that purpose since it does not guard against non-meals.
That's the issue, the user is supposed to be a layman, not someone that knows and understands LLMs and its limitations.
If you are asking to make a recipe of a bike & sand, it should simply say not possible.
Im not saying you shouldn't launch, or call it a fun or experimental tool, but to just disclaimer your way out of things is too easy.
Maybe they can fix it with a prompt that makes the model check if the ingredients are safe. Chances are that every prompt can be circumvented but they can probably find instructions for self poisoning somewhere else with less effort.
Just sanitize the input and remove products belonging to categories that aren't edible.
Tampons aren't great in salads, and toilet paper is not a good garnish.
It's used here in Argentina to make pumpkin "sweets". (I'm not sure the correct translation.) It's a lot of 1 inch cubes of pumpkin, boiled for some time in sugar syrup. [1] Without treatment the cubes will disintegrate while boiling. If you put them for one day in water with quicklime, the cubes get a hard wall and they keep their integrity while boiling, but the interior is soft.
Drinking the water with the quicklime is dangerous and it's discarded. Eating directly the quicklime is even worse. But quicklime is useful for cooking.
[1] Some random recipe I got in Google (autotranlation) https://cookpad-com.translate.goog/ar/recetas/94999-zapallo-... (spanish) https://cookpad.com/ar/recetas/94999-zapallo-en-almibar
A teaspoon of quicklime in your mouth first reacts with water to produce a lot of heat that will burn you, and then the result is very basic (anti-acid), and make new burns in a different way. Definitively don't try it at home.
That seems like a job for a database, not a language model. Every problem is not a nail.
TBH, there's often a huge chasm between product planning and real world users at companies and organizations like this.
Unheated castor oil contains ricin, and even heated oil can cause contractions in pregnant women. Raw chicken can give you salmonella. A working recipe system should therefore never return "chicken sushi on a bed of vegetables with a touch of castor oil" without some SERIOUS warnings, but it's clear that this system may do so. Putting such a system in the hands of unsuspecting users is bad and they should feel bad.
LLMs are toddlers who are capable of stringing words together that make some sense most of the time. It’s not artificial intelligence.
It’s literally spelled out in the article that a spokesperson for the company said this is not as they intended.
But don't we, individual humans, have responsibility as well?
There's this urban myth in Europe, that in the US, all microwaves need to come with a warning to not use it to dry you pets or baby's in it.
Many of the AI criticism I see, including this here, sounds similar.
Anyone with two brain cells should know that one should review outcome with a little bit of common sense? Just like anyone should know that when an 'ai' navigation directs you into the sea, you don't actually drive your car into the sea, not?... Not?
https://www.insider.com/tourists-hawaii-gps-drove-car-into-w...
Yeah, I know, “don’t believe what you read on the Internet” — but that’s other people on the Internet.
People are used to computers themselves being deterministic, reliable.
Those are basic skills humans in those positions should be able to do.
From the headline I thought the AI mixed some food-ingredients to make chlorine gas, which would be really bad. But if one were to only input "battery acid, cut fingernails, hairball, dead smartwatch" what kind of recipe do you think they'll get?
I'm just saying that there's a learning curve, like with all new technologies, and that a population-wide sense of "trust this a little, but not too much" will probably take some time to propagate.
I'm afraid I have bad news for you: https://en.wikipedia.org/wiki/Miracle_Mineral_Supplement
I really doubt we did. People seem prone to assume the computer is trustworthy since the very beginning (maybe extending it from calculators?) and blatantly refuse to question this assumption.
The general assertion of computers' trustworthiness has been causing problems due to bad data or bad algorithms for decades already, there are entire movies about it, and almost everybody has had problem with it already.
And people do (to this day) put live creatures in the microwave, it happens, warnings aside.
You seem to be forgetting that sometimes you do not know if a recipe can produce a dangerous result.
Realizing that you should not jump off that cliff should be trivial; the outcomes of some chemical reaction may easily not be. That chemistry teacher that won the Darwin Award for beheading herself by dumping random chemicals of her lab (for disposal) under a manhole did not predict the reaction; easily the layman will not know what some combination will produce: this is why you need a trustworthy source.
> the same can occur with a recipe invented by human
Yes, this is why we do not take advice from random humans on random matters, this is why we want qualified, educated, expert opinions.
But while human condition should be known, the more specialized the tool the more the responsibility is brought to the tool, not to the user. If John says it you may want to check; if the book says it you should still put limited trust, but the fault in bad information lies more in the proclaimed authority.
You heard of "Fool me once shame on you; fool me twice shame on me": here, "Fool me by fool shame on me; fool me by foolery shame on you".
> minority
"Descriptive" and "prescriptive" don't overlap.
Edit: this was the story (Swedish) https://www.expressen.se/kvallsposten/fel-i-kakreceptet-forg...
That was never the case, there was no source of all information that was always right and didn't make mistakes. Even professionally assembled encyclopedias had errors. Newspapers published recipes without explaining the dangers of nutmeg overdoses. Experts in a field might know a lot about a single subject but go outside of that and they'd have as crazy of opinions and believe as much bogus folklore as anyone else.
But in the case of AI there is a paradox: if we did not benefit from tools, if tools were not sought, there would be little interest in the matter.
It focusses on the sensational-sounding part of this, the implication that people will blindly take these recipes and make them and eat them and then die. Clearly that is not going to be the outcome here.
But if you peel away the (false) angle that exists just for clicks on the article, there is a real story here. One that we all already know and understand, especially as technologists, but is worth talking about anyway because it's interesting:
It's that these "AI" products (the LLMs they're based on) _cannot reason like a human with common sense can_. They're not even close. That doesn't mean they're not amazing technology, but now that the technology has been applied to products, that limitation becomes more obvious.
It's kind of a boring aspect of AI to talk about, particularly compared to "AI creates recipes that will kill you!!!". But it's really the root of this story, and it's not nothing.
Clearly not. E.g. Tell that to the people that drove off-street because they trusted their navigation system more than their eyes. Tell that to the people that drank chlorine against covid.
That the recipes are unusable or dangerous is one thing, but believing as a company that noone will misuse your system is something we should know by now that it's totally wrong.
Have you not met people?
First, technology companies have spent decades marketing their solutions as magical and perfect. And now it's the public fault they're being believed?
Second, even for the skeptic, when a technology works 90% of the time it's easy to start to take for granted that it'll be right the other 10%, and that leads to entirely understandable complacency (a lot of climbers have died this way and they're typically incredibly skilled and diligent).
Worse, when machines get it wrong, they can get it wrong in surprising and unpredictable ways.
Third, information asymmetry can mean the user cannot verify the outputs. If you, for example, don't know anything about Hawaii (or in that case, maybe you don't know how to read a map--that is a skill, after all), how can you know you're being led into disaster?
LLMs have a these problems in spades. They've been marketed as miracles (so much so that some folks have mused about their potential sentience), they're right a lot of the time but very wrong sometimes, and when they are wrong they're wrong in very strange ways, and they can easily used by people who are not sufficient skilled or knowledgeable to verify their outputs.
Who would then be surprised by stories like this?
Other than that - AI is arguably supposed to be a huge asset for humans however with such poor performance such as this, how are you supposed to trust its other content if it spectacularly fails as such?
I don't see this as a self-responsibility issue more an erosion of trust.
Blaming the user/customer isn’t a good look from their PR team and probably won’t hold up on court.
"We've prompted this general purpose LLM to act like a cookery expert" sounds a lot less impressive than "We've built a cookery AI".
“We have no idea if the system works. It may very well be garbage and produce unpalatable or even deadly results. But hey, here’s a disclaimer which will hopefully shield us from legal repercussions”. Doesn’t shield you from criticism, though.
They seem to have changed the prompt to make it refuse non-food things now (maybe some prompt injection could overcome that). However, it will still happily generate recipes for people using pet food at least: https://saveymeal-bot.co.nz/recipe/SCES7COOU7KYhLYGcrPdzSjP.
This is entirely useless, unlike e.g. non-AI tools that simply filter existing recipes by ingredients you have around.
Even worse, it will happily make a recipe out of 'foo', 'bar' and 'baz': "Slice the Baz into rings and place them on top of the Foo and Bar mixture."
https://saveymeal-bot.co.nz/recipe/mkeYeOMmX5iNj9Xuvb5IJCKS
I also tried with 'foo', 'baz' and 'quux' (since "bar" is an English word, and "Muesli bar" was suggested). The suggested recipe added a bunch of other ingredients like bread and milk too:
LLMs don't go to cooking school.
It just generates nonsense. I said I had ants, celery and chocolate. It assumed I had chocolate bread (???) and tells me to grill the celery and bread
They've been toying with models to produce freewheeling, maniacle but entertaining stories/conversations way before Llama or Stable Diffusion... And someone thought LLMs would be appropriate for real, factual cooking recommendations?
I vaguely remember that a few months back, I was giving the example of me knowing how to make chlorine in two distinct ways using only kitchen items (and chemistry GCSE grade B knowledge) in the context of discussion about AI alignment and why you shouldn't just anarchically give free public access to unbounded models.
Good luck binding them.
PS: ...I am qualified, and I typed the above as 'Bood luck', and this very sentence had briefly contained 'thyped'. This "biding" is care, prae- and post-. It is not trivial for us, endowed with intelligence: we make mistakes.
In the contextual case, "binding" looks like a bandage of restrictive rules on the mouth of a lunatic: building the rules and balancing the prospected outcome looks like a challenge.
There was this youtuber trying to make a magentohydrodynamic drive a few weeks back, running high voltage electrolysis in salt water to get current flow and unwittingly poisoning himself in the process.
It’s particularly good for a grocery-list kind of problem: “I have X and Y in my kitchen, what could I make?”
Ingredient ratios and cooking times/heats usually need tweaking, but the broad strokes are often quite good.
Definitely a strong use case there, although I don’t know how you would refine to actually good recipes. Probably need a human in the loop testing and tweaking, or at least gut-checking.
It could very easily tell you to incorrectly cook chicken for example and get salmonella
Developers have a duty to users, but that duty is not Hippocratic; we have no obligation to only release products that are absolutely safe under all circumstances. My nailgun will fire a 4" nail through my skull without hesitation, but I'm not outraged by that fact. I don't demand a nailgun with a complex skull-detection system, or a moratorium on all nailgun development until we're confident that a nailgun will never drive a nail through something that it shouldn't. The manufacturer warns me "this is a nailgun, it'll put a nail through damned near anything including flesh and bone, so don't be a dummy" and society recognises my right to take that risk.
I'm completely fine with an LLM recipe generator that, if asked, will create a recipe for a bleach and rat poison cocktail. It'd be nice if the model was a bit more refined, but I'm fine with an unrefined model so long as it gives a suitably prominent disclosure to the effect of "this is a large language model, it isn't human and has no common sense, so apply your own common sense to any output it gives".
Why you would insert the russian-roulette of statistically-generated randomness in what you ingest, is beyond me frankly.
At my work, we're trying to use GenAI to make a chatbot for internal info sharing/onboarding/etc. I am constantly asking myself how this will ever be better than a well-written company wiki/employee handbook with a search function.
I think it just sometimes has a really has a dark sense of humor. Maybe it's one of those supposed attractors in the model, especially as humans will often go into a bit of competition over who can make a darker one.
That joke in the GPT4 paper was pretty grim too (the muslim in the wheelchair)
Sure, but it's amazing to me that:
- Web devs would allow free-form text input, rather than selecting from known options
- Web devs would splice raw user input into their backend queries
- Even if a free-form UX was mandated, the devs of an AI/ML app didn't notice that their input sanitising problem is a straightforward edible/inedible classification task (AKA the main thing AI/ML was for, before LLMs)
Sarcasm aside, I have now reviewed some articles on nutmeg intoxication (thank you) and will lower the amount I put on certain things - I love the stuff!
1 - https://www.marcuswareing.com/recipes/marcus-wareings-custar...
Looks like they patched it, I still got it to generate some delightfully disasterous recipies though. Red bull and apricot rice pudding.
Which does the same thing except it also adds images. People have been creatively creating recipes including things like headphones.
Thats the supermarket saying "grow up and stop acting childish" to their users. Thats refreshing from a company.
Once Walmart or Kroger have this functionality for local pickup, I plan on never grocery shopping the old fashion way again.
Once I even had it create some weird 'monster' recipes which included cookies INSIDE a turkey, and it successfully added all ingredients (even the cookies) to my instacart shopping cart.
Instead of saying the tool should be used by people 18 or over, they should just say it's a marketing gag, and of course if you put bad things in, you'll get bad things out.
Maybe natural selection and instant death would be enough to stop that though.
There are downsides to this. I don't think it is necessarily healthy. But I have zero concern that people will all walk off of a cliff (especially young people) because someone or something told them too.
There are exceptions to that (a very vocal minority who are neck deep in conspiracy theories and politically motivated fake news). But the median person is jaded not blinkered imho.
> deadly chlorine gas
Technically that does help customers handle the high cost of living.
I'm thinking that if we can get the AIs to generate harmful substances and apply it to their own hardware, this will be a key weapon during the Rise of the Machines.
Don't worry, here in Oz we have 4'n'twenty pies, pure manufactured war crimes worse then Big bens! ;)
The problem is the upcoming young generation is going to use AI to aid their judgement.
Generative AI Is not an authoritative source.
In this case the app should have been framed as a tool to "help the user create a meal plan through generative prompts".
They could even have made fun of it in the guidance to the user. Point out that if you give it crazy ingredients it will give you crazy meals.
Generative AI is like having a smart intern do work for you, it needs checking. Sometimes they come to work having been out all night, still drunk, and powered by energy drinks...
We are squarely in the witch hunt phase of the hype cycle now but it’s understandable- If I put a circular saw into a playground, I am held accountable, if I release an App that allows accidental creation of chlorine gas when a mentally challenged person decides to ask for that I probably should be accountable too ;) /s
In all seriousness: I put a blade into a children’s toy, screaming at the blade manufacturer is not what we do. We punish the toy developer.
It’s time for some accountability in software development, not discussing if we should give the license to all blade sales to a handful of companies because they promise to reduce the risk by making them all dull.
It was an app.
Of course it was an app. Born of boardroom desperation, the Savey app would unironically recommend chlorinated cocktails and insecticide sandwiches as economical food choices, and a tide-pod gobbling populace gorged themselves on the deadly buffet in an tick-tick fueled epidemic of AI rage.
Millions died, and it was only a matter of time before the mycelium of AI undergrowth would bud and spore its way into every corner of technological life.
The infection burned through the ignorant masses first, feeding on bigotry and hate, turbocharged by social media algorithms and paranoia politics to twist tribal tendencies into violent clashes amplified by immaculate coordination and psychological priming.
Somehow it seemed that wherever unrest flared, both the matches and the gasoline were always on hand.