Food and Generative AI
engineering.hellofresh.com
engineering.hellofresh.com
- This company and its supply chain went to the trouble of packaging a single serving of mayonnaise for me to personally discard as my token contribution to waste. Sure, an extra mayo packet won't make a difference, but simply not caring feels bad, you know?
- I have a jar of mayo in the fridge. I didn't even need this.
- Come to think of it, everything in this meal kit is wrapped in packaging that may well outlive me, depending on where it ends up. All for one meal's worth of food that I could have obtained more cheaply, less wastefully, and with more control over customization had I mustered the effort to plan ahead.
Most of the meals came with poor quality produce (e.g. tiny unripe avocado — something I never would have grabbed at the grocer) and a silly amount of waste for items I had already.
My least favourite part was the recipes. I’m not sure if it was just the small selection of recipes we got, but it seemed like each one was designed around a template with six photos. As if they had a spreadsheet of all their recipes and required each one to fill six columns so they could just print it out in a predictable way. This meant that if you followed the recipe you would wind up doing several redundant tasks and creating unnecessary dishes. It drove me wild. By the last meal I just read the recipe, looked at the ingredients, and improvised.
A high quality recipe should not only result in a delicious meal but also be efficient and thoughtful about order of operations and clean-up. Having full control over ingredients, HF has no excuse providing low quality recipes, never mind the ingredients.
I would recommend a NYT Cooking subscription and using a grocery delivery app instead.
But overall I really like it. I actually learned a lot about cooking. The meals aren't complicated, and they follow similar patterns, so you start picking up on things. Like seasoning in layers (they tell you to salt & pepper at multiple steps), how to improve flavors with stock, spices, how to make presentation nice. And they're all pretty quick. After a while I have a nice library of recipe cards, and now I just pick a few of them when I go to the grocery store and get the ingredients myself. And even though sometimes it doesn't feel like enough food, it almost always comes out perfect. Sometimes I'll augment the amount of rice or something like that.
So I like HelloFresh, and maybe eventually I'll graduate to more interesting meals like NYT Cooking. But HelloFresh is quick, easy, and usually good.
What do you mean about "creating unnecessary dishes"?
Instead I soaked one ingredient, while sautéing the other ingredient. When I drained the first ingredient, I just added the sautéd stuff to that bowl, saving one bowl from washing.
Another thing is the order you cut stuff. If you cut your produce first and then meat you can use the same cutting board. But if you cut meat before veggies (that will stay fresh) then you need a new board, or at least a flip.
* Using much more citrus, juice & zest * Chickpeas + spice mix + touch of liquid in a pan is an easy go-to (we're veggie) * Broth / onions in rice. * Template of flatbread + stuff on top, which they had 3 or 4 in their rotation. Easy to riff on w/ whatever you have in the fridge.
None of these are earth-shattering revelations, but the repeated practice has ingrained some of these much more, and got me cooking a little differently.
We do this every few years, it seems worth it for a while, but I can't stick with it because I hate the waste, and the recipes get pretty repetitive after a bit, doubly so w/ vegetarian meals which have fewer options.
How do you train an AI on taste?
Would a chef state what's allowable for a recipe? "This one you can sub the soy out, but this one falls apart without it"
You keep shoveling piles of recipes at it, and as long as most of the recipes are positively tasteful (as opposed to just randomly generated, or worse, engineered to suck out of spite), the AI should eventually pick up on taste in general.
Having done a lot of user research, sometimes the most important things to a person aren’t stated, or even consciously known.
If I’m shopping for chocolate chip cookies, they’re a proxy for another need: hunger, yes, but also emotions: wanting to feel taken care of, missing a place, etc. If chocolate chip cookies are unavailable, I’m not now considering snickerdoodles or shopping chocolate bars; I’m shopping for the feel: maybe it’s pie, maybe soup.
I’m interested to see if the AIs we build will have the ability to identify this, so when someone says “I want a steak frites” the AI doesn’t strictly recommend steak recipes—“they’re asking for steak”—but realized the constellation of what the asker is actually asking for—they’re asking for a recipe that reminds them of when they lived in that apartment in that city X years ago.
It will be fun to see what’s possible …
Actually generating a recipe is obviously much harder, and should be effectively impossible unless you include a humanoid robot that can cook and taste. I think the best you could do is generate a possible recipe that highlights ingredients and techniques you'll need to experiment with.
Also some of their final ideas don't really make sense:
- "Carb smart mexican beef and capsicum stuffed peppers with fries , avo and sour cream". Carb smart but it comes with fries? And "capsicum stuffed peppers" is a strange way to phrase it.
- "One - pan beef meatloaf italiano with green peas and spinach". You can cook this in one pan if you really want to but it will not be good.
- " little ears " pasta serves as balsamic tomatoes. Strange phrasing.
A do think a different model could solve these problems, but coming up with generic recipe ideas is not really a problem anyone has.
Would be interesting if you could prompt, LoRA distill, or use modern LLM tricks against a well-labeled and curated set, similar to how other tagging problems are handled with modern pretrained models.
[0] https://open.nytimes.com/our-tagged-ingredients-data-is-now-...
When I've looked at gpt recipes they've had some pretty deranged proportions and you'd be in trouble following them precisely. If you just need a list of ingredients and a loose technique though eg "tofu & cashew stir fry with soy sauce and garlic" then ya sure I mean absolutely.
I love being able to challenge it with "and now make it vegan" etc to see what it comes up with.
This is also another example of the thing where having a randomly unreliable teacher is actually quite useful if you're trying to learn to cook - because it forces you to question what it tells you, apply your own intuition and think carefully about what worked and what didn't.
But the more esoteric you go, aiming for something really novel or something unusually historic/traditional/regional, or the more technical and delicate, the more it'll turn to hallucinating naive text continuations that will lead you astray. An experienced cook will know to spot the weird stuff and revise on the fly; an inexperienced one is going to end up with some... odd dishes.
It's the same as many of us see when used for code assistance. It's great at inventing "original" boilerplate which inspiration has been exhaustively covered in source material, but gets real wacky and unreliable the more you move away from that.
It's not because this necessary information is already implicit in the sum of all the data it is trained on.
That's not to say that LLM's can't mash together reasonable new recipes for certain audiences, especially for variations on popular modern standards, but the idea that their working space of recipes is exhaustive across all cuisine and tastes is absurd.
Predictors don't just grok what is explicitly stated in a dataset but also what is implied by its structure.
In a language model trained on protein sequences alone and nothing else, you will find biological structure and function emerge in the inner layers.
https://www.pnas.org/doi/full/10.1073/pnas.2016239118
Language Models are predictors. If the text you give them is the shadow in Plato's cave, they don't try to draw shadows, they try to build walls. They are trying to reverse engineer the computation that must have led to that output. They will not stop at "surface level similarity" or "plausible" by choice. They will train until they have completely succeeded or until the architecture or data fail them.
With a capable enough architecture and sufficient data, there is no recipe a predictor couldn't divine with enough training.
For the perfect predictor of recipes, Is the architecture capable enough ? Is the data sufficient (both variance and quantity) ?
I'm not sure but this is not a question a cook can answer.
The OP claimed
>Actually generating a recipe is obviously much harder, and should be effectively impossible unless you include a humanoid robot that can cook and taste.
This is just false. And you don't need the hypothetical perfect predictor (just a capable one) to see it.
This is the fantasy of GenAI boosters stated concisely - that these statistical methods present a new objectivity that avoids the shortcomings of human subjectivity. Of course, these web-scraped models are distilled subjectivity in aggregate, subjectivity at scale. ChatGPT is made of people!
If you're wondering whether a task is theoretically possible for an AI, I think a good rule of thumb is to ask "could a human domain expert do this given a few days and as much research material as they need". If the task here is "develop a novel recipe without trying it", the answer is no. You can't predict the interactions of every possible combination of ingredients and processes no matter how much time and experience you have.
That's just the thing machines do better than humans. Not sure about LLM's though.
A GPT (Generative pretrained Transformer) is not a simulator nor an imitator. It is a predictor.
Predictors try to reverse engineer the computation that could have led to a certain output so they don't just grok what is explicitly stated in a dataset but also what is implied by its structure.
As an example, In a language model trained on protein sequences alone and nothing else, you will find biological structure and function emerge in the inner layers.
https://www.pnas.org/doi/full/10.1073/pnas.2016239118
In fact, with a capable enough architecture and sufficient data (quantity, qualit, variance), there is nothing a predictor couldn't divine with enough training.
You assert
>Actually generating a recipe is obviously much harder, and should be effectively impossible unless you include a humanoid robot that can cook and taste.
In the pursuit of recipe prediction, a predictor will learn to internally model taste because taste is implicit in the data.
You don't need the hypothetical perfect predictor to demonstrate your assertion as false because a capable predictor already does.
>If you're wondering whether a task is theoretically possible for an AI, I think a good rule of thumb is to ask "could a human domain expert do this given a few days and as much research material as they need".
You are wrong. Train a language model on descriptions of the functions of proteins and an equivalent protein sequence and you get a language model that can predict novel functioning sequences from function descriptions alone.
https://www.nature.com/articles/s41587-022-01618-2
Not only is this not something a human expert can achieve in a couple days, It's not something a human expert can achieve at all. It is a Super-Human ability.
It's a bit ridiculous that this is the best they can come up with, and I don't really think the idea is worth defending, especially by telling other people that they just don't get it.
It's not a perfect predictor so it's not impeccable for all conceivable situations but GPT-4 can already generate novel recipes that taste nice. This is not some far flung science fiction ability.
It's possible the data is enough to get a good predictor but the one you tested simply wasn't good enough yet.
I appreciate your links though, I think this is the first I've heard of an LLM doing something that a human can't do. I can't claim to understand a lot of the first paper and I can't access the second, but assuming your descriptions are right I agree that my rule of thumb was wrong.
First let me explain what I mean by implicit incentive with our biology example.
You have the protein sequence, G46AKT5778FAG4
The predictor is given "G46A____"
The only way to predict this correctly consistently across multiple proteins is learn the underlying biological structure.
Now, I pulled this snippet of a recipe randomly from the web.
"Add the tomato purée and turn up the heat slightly, cooking until it has darkened, about 2-3 mins."
Block out some information.
"Add the tomato purée and turn up the heat slightly, cooking until it has ____"
The only way a predictor is getting what kind of completion to make consistently correct across multiple unseen recipes and food items is if it has some model of the visible effects of heat on food in general. The predictor will be forced learn what kind of foods darken in heated water and group them together internally so it doesn't have to memorize every single instance (more effort than simply learning it)
Now let's put back darkened and block out the time.
"Add the tomato purée and turn up the heat slightly, cooking until it has darkened, about____."
The only way a predictor is consistently correct on what time interval this should be across multiple recipes is if it has some model of the effects of heat and time on whatever it is you're cooking.
Implicit incentives are everywhere.
GPT-2 was mostly an incoherent babbling mess but that didn't mean a better predictor couldn't be coherent.
GPT-3 could not play chess at all but that didn't mean a better predictor couldn't play chess (3.5-turbo-instruct)
Taste is implicit in recipes so a good enough predictor has to model it somehow to succeed, no physical experimentation necessary.
Yes, thank you! I think a lot of companies have lost the plot jumping to LLMs when they aren't necessary. Recipies are a pretty small search space. You don't really need an LLM to understand the questions being asked.
I want this AI to be considering whether what I ate this morning had enough protein, whether this ingredient is on sale and in stock nearby, or whether it's in my fridge and what's the expiration date on that. Is there a plan for the other half of the cauliflower? etc.
The last thing I need in a meal is appealing marketing copy for it.
It needs to be able to:
- track what I have (if I have input that data) and how soon anything may expire
- track what I made but haven't finished eating (just a few quick pictures in the fridge should update that as needed, or maybe it's time we all get that Alton Brown in-fridge camera to talk to)
- can I use this in something else? How long should it still be good? Etc...
- track my health/nutrition goals
- track what I've already eaten (again, the onus is on me to make sure that data is there) so it can help me follow my health/nutrition goals
- track local grocery store fliers for deals (I do this on my own with the app Flipp and save a ton of money, but if the AI could handle that... that'd be great)
- tracking local grocery store prices would be great too, but unlikely. They need us to come into the store for hopeful impulsive purchasing
- I'm sure there are more things
Not saying this is remotely simple to pull off... just what I would want in an AI that is helping me with recipes and food. I can't wait until we can utilize it for things like that.Thanks LLM! That was way better than some awful google recipe SEO spam nonsense.
You can't use a language model to predict how something is going to taste.
Less tritely - I do find it fascinating how problems that would've been independent research endeavors can now be subsumed by large language models. Rather than building a big dataset of protein, calories, etc., just ask ChatGPT
This way we can reduce food waste by 30% on average.
It's a stretch to call it fraud when there's no guarantee what you make at home will look like what's made in a test kitchen. I challenge you to try to do food photography well at home without thousands in photography equipment, props, the actual dish, and expertise in making it look enticing on camera.
https://arstechnica.com/information-technology/2023/08/ai-po...
I’d not also that it’s getting harder to do that with current ChatGPT (this article uses GPT3.5), and I suspect “alignment” research in the 5 years time frame will make these sorts of things pretty hard to trick the models into doing.
> this article uses GPT3.5
Which alone disqualifies it from opining on what LLMs can or can't be used for.
In all seriousness, we're just finding that GIGO still applies. That doesn't mean the tools aren't useful.
The difficult part of recipe development is testing, and professional recipe authors pay a lot for it. (It was regularly available pickup work when I was in culinary school.) Professionals don't use recipes like home cooks do: ours are much simpler, tend to combine ratios and amounts with known techniques, and assume knowledge that most home cooks don't have. Home cook recipes tolerate people who don't really know what "4 lbs turned radishes glace w/sherry" means, where nearly any professional cook familiar with classical european cooking could grab the half dozen ingredients and equipment without clarification and perform the technique almost identically to another cook across the world. It's not rocket science, but even with clear instructions, there's a lot of technique there that you just have to do at least a few times to get a sense of it in a really basic way.
Writing recipes for home cooks poses the same challenges developers have writing documentation for people who aren't familiar with the codebase they're documenting. Home cooks consistently execute things like "Saute" much differently than a chef might assume they would, use seasoning very differently, measure doneness by the amount of time cooked rather than using internal temperatures, textures and smells, need measurements for "pinches" of things and use volumetric measurements instead of weights, and and all sorts of other things. For a home cook recipe, the professional instruction, "hard sear, glaze, and brown in the broiler" would need likely need specific times, settings, pans, and things like that... and the temperature of home cooking equipment, the initial temperature of the ingredients, slight variations in salt or sugar content, all have a significant chance of making that recipe fail. I guarantee you that when, say, Gordon Ramsay writes a cookbook, rather than writing the recipes himself, he goes into a R&D kitchen, cooks the dishes with those cooks who know how home cooks do things, and they write the actual recipes that go in the books.
So like almost every other "hey lets replace some creative/technical person with this generative AI" initiative I've seen, this would do the easiest 95% that takes 5% of the time while not addressing the most difficult 5% that takes 95% of the time. When someone asks Midjourney to make, oh, say, a sexy elf in the style of Thomas Kinkaid or whatever Midjourney users want these days, they might not be a professional artist, but since they're the consumer, they can judge whether the elf looks appropriately sexy or stylistically enough like Thomas Kinkaid. Getting a recipe spit out like this requires technical judgement that any user who'd rely on such a device would almost certainly lack.
Fortunately/Unfortunately it seems pretty difficult to devalue professional cooking as a skill any more than it already has been, and your average pro doesn't have much exposure to this sort of market anyway, so I'm not really worried about industry impact compared to, say, concept artists for video games... Though I think workaday utility developers are more squarely on the chopping block than most. However, I feel for the home cooks who'll faithfully follow these recipes expecting similar results to what they get from foodnetwork.com, simply recipes, the new york times food section, or whatever other source they get recipes from.
Sorry to pick up on one sentence of an interesting long comment, but this is a pet peeve on mine when having to rely on American recipes. There seems to be an obsession on using volumetric measurements for things which really should be done in mass units - flour, sugar, sometimes even ingrediants like grated cheese.
European recipes in my experience tend to use most 'natural' units - mass for solids, volume for liquids - but you still get the silly culturally specific units sometimes - '1 medium onion' ... well how big is that? Would two small onions be too much? Or half a large onion?
As far as the medium onion is concerned, some of that is unavoidable because onions are a natural product with size variations, and unlike sugar, salt, flour, water, etc that serve important chemical or mechanical purposes during cooking, something being 10-20 percent more or less oniony in the grand scheme of all the flavors you have is not likely consequential. Most professional recipes wouldn't be that much more specific in that regard.
With respect ingredients such as onions, I am concerned that what I think of as a medium onion, may not be a medium onion in other countries/cultures. Like, I generally will consider an onion of about 5cm (2 inches) as medium-sized, but if I am off by 1cm from the author's perceptions that's 50% less, or 70% more!
If a recipe could say '125g chopped onion (1 medium onion)' that would be so useful - catering to the geeks like me and normal home cooks!
The most common manifestation of this disconnect is in cooking times. Unless you've got completely standardized ingredients, preparations, storage temperatures, etc. which is very difficult outside of large food service organizations, (and why so many restaurants are willing to pay so much to huge restaurant suppliers like Sysco for mediocre food... it will always cook the same way,) cooking times will almost always be the wrong answer. People always ask questions like "how long do I cook a thick ribeye steak?" and I always say "until its done." There's a general perception that very accurate cooking times will yield very accurate results, but that's so not true. What's the size and shape of the steak? What's the water content of the steak? Has that changed since you pre-salted it too far in advance? Refrigerator temps vary tremendously: what's yours at, how long has the steak been out.
People like cooking times and precise measurements for imprecise ingredients because it gives them a sense of control, but it's a false sense of control. when the only way to get that sense of control is learning how to tell when the food is done by yourself.