It's not because this necessary information is already implicit in the sum of all the data it is trained on.
It's not because this necessary information is already implicit in the sum of all the data it is trained on.
That's not to say that LLM's can't mash together reasonable new recipes for certain audiences, especially for variations on popular modern standards, but the idea that their working space of recipes is exhaustive across all cuisine and tastes is absurd.
Predictors don't just grok what is explicitly stated in a dataset but also what is implied by its structure.
In a language model trained on protein sequences alone and nothing else, you will find biological structure and function emerge in the inner layers.
https://www.pnas.org/doi/full/10.1073/pnas.2016239118
Language Models are predictors. If the text you give them is the shadow in Plato's cave, they don't try to draw shadows, they try to build walls. They are trying to reverse engineer the computation that must have led to that output. They will not stop at "surface level similarity" or "plausible" by choice. They will train until they have completely succeeded or until the architecture or data fail them.
With a capable enough architecture and sufficient data, there is no recipe a predictor couldn't divine with enough training.
For the perfect predictor of recipes, Is the architecture capable enough ? Is the data sufficient (both variance and quantity) ?
I'm not sure but this is not a question a cook can answer.
The OP claimed
>Actually generating a recipe is obviously much harder, and should be effectively impossible unless you include a humanoid robot that can cook and taste.
This is just false. And you don't need the hypothetical perfect predictor (just a capable one) to see it.
This is the fantasy of GenAI boosters stated concisely - that these statistical methods present a new objectivity that avoids the shortcomings of human subjectivity. Of course, these web-scraped models are distilled subjectivity in aggregate, subjectivity at scale. ChatGPT is made of people!
If you're wondering whether a task is theoretically possible for an AI, I think a good rule of thumb is to ask "could a human domain expert do this given a few days and as much research material as they need". If the task here is "develop a novel recipe without trying it", the answer is no. You can't predict the interactions of every possible combination of ingredients and processes no matter how much time and experience you have.
A GPT (Generative pretrained Transformer) is not a simulator nor an imitator. It is a predictor.
Predictors try to reverse engineer the computation that could have led to a certain output so they don't just grok what is explicitly stated in a dataset but also what is implied by its structure.
As an example, In a language model trained on protein sequences alone and nothing else, you will find biological structure and function emerge in the inner layers.
https://www.pnas.org/doi/full/10.1073/pnas.2016239118
In fact, with a capable enough architecture and sufficient data (quantity, qualit, variance), there is nothing a predictor couldn't divine with enough training.
You assert
>Actually generating a recipe is obviously much harder, and should be effectively impossible unless you include a humanoid robot that can cook and taste.
In the pursuit of recipe prediction, a predictor will learn to internally model taste because taste is implicit in the data.
You don't need the hypothetical perfect predictor to demonstrate your assertion as false because a capable predictor already does.
>If you're wondering whether a task is theoretically possible for an AI, I think a good rule of thumb is to ask "could a human domain expert do this given a few days and as much research material as they need".
You are wrong. Train a language model on descriptions of the functions of proteins and an equivalent protein sequence and you get a language model that can predict novel functioning sequences from function descriptions alone.
https://www.nature.com/articles/s41587-022-01618-2
Not only is this not something a human expert can achieve in a couple days, It's not something a human expert can achieve at all. It is a Super-Human ability.
I appreciate your links though, I think this is the first I've heard of an LLM doing something that a human can't do. I can't claim to understand a lot of the first paper and I can't access the second, but assuming your descriptions are right I agree that my rule of thumb was wrong.
First let me explain what I mean by implicit incentive with our biology example.
You have the protein sequence, G46AKT5778FAG4
The predictor is given "G46A____"
The only way to predict this correctly consistently across multiple proteins is learn the underlying biological structure.
Now, I pulled this snippet of a recipe randomly from the web.
"Add the tomato purée and turn up the heat slightly, cooking until it has darkened, about 2-3 mins."
Block out some information.
"Add the tomato purée and turn up the heat slightly, cooking until it has ____"
The only way a predictor is getting what kind of completion to make consistently correct across multiple unseen recipes and food items is if it has some model of the visible effects of heat on food in general. The predictor will be forced learn what kind of foods darken in heated water and group them together internally so it doesn't have to memorize every single instance (more effort than simply learning it)
Now let's put back darkened and block out the time.
"Add the tomato purée and turn up the heat slightly, cooking until it has darkened, about____."
The only way a predictor is consistently correct on what time interval this should be across multiple recipes is if it has some model of the effects of heat and time on whatever it is you're cooking.
Implicit incentives are everywhere.
It's a bit ridiculous that this is the best they can come up with, and I don't really think the idea is worth defending, especially by telling other people that they just don't get it.
It's not a perfect predictor so it's not impeccable for all conceivable situations but GPT-4 can already generate novel recipes that taste nice. This is not some far flung science fiction ability.
It's possible the data is enough to get a good predictor but the one you tested simply wasn't good enough yet.
That's just the thing machines do better than humans. Not sure about LLM's though.
GPT-2 was mostly an incoherent babbling mess but that didn't mean a better predictor couldn't be coherent.
GPT-3 could not play chess at all but that didn't mean a better predictor couldn't play chess (3.5-turbo-instruct)
Taste is implicit in recipes so a good enough predictor has to model it somehow to succeed, no physical experimentation necessary.