Role play with large language models
nature.com
nature.com
Shameless plug, I made a GPT for playing improv: https://chat.openai.com/g/g-LkQhMxpvM-improv-theatre
I don't see how the role play term doesn't introduce new issues. Cambridge, for example, defines role play as "pretending to be someone else". An LLM is also not pretending. Also what "role" would an LLM play if you just take the base model without a default prompt or finetuning?
What role would any person into sexual roleplay play if you don't prompt them into sexual roleplay?
Or to put this another way... The 'role' may be more that of Legion in the biblical sense. Depending on the exact question asked hints of particular human characteristics show up, but there is a multitude of different ones and the model would seemingly randomly express them per question.
I'm sure you could tune it to have more attention than a two year old and actually apply the rules. It did come up with a decent setting based on B2: Keep on the Borderlands. And it was some fun.
The new generation of text games will be awesome. And Enders tablet game is possible with this tech.
Complete with aliens subtly messing with it, and through it, the player?
They who control the logits, control the future.
Yes, that's the RLHF and other mechanisms. Raters presumably reward completions which are, well, complete and can be quickly judged as a whole, no matter what the intrinsic quality might be or how valid it would be to end with a 'Tune in next week for chapter 2'.
What OP is describing is the emergent behavior of the base model, which acts very differently from the RLHF/instruction-tuned/who-knows-what-else ChatGPT web interface. (This is one reason I tend to avoid that in favor of the Playground and direct model access: there's much less moving machinery behind the scenes.) What the additional stuff does is quite hard to understand because the original prediction objective has been replaced by a bunch of mashed-together objectives combined with feedback loops, leading to some bizarre behavior like being unable to reliably write a nonrhyming poem. (Give that a try in ChatGPT: "write a nonrhyming poem". It's been slooowly getting better at it, perhaps because I and other people keep submitting examples of it failing to do so - but it will still usually fail! and when it does seem to be succeeding finally, if you let it keep writing lines, it will generally gradually revert back to rhyming.) If you go back to the original davinci-001, you'll find that it acts strikingly different than your description of contemporary ChatGPT.
OP, incidentally, has some discussion of what it is like to interact with the real GPT-4, which is not like ChatGPT-4: https://www.lesswrong.com/posts/tbJdxJMAiehewGpq2/impression...
AFAIK, this is one of the only discussions online of what GPT-4-base is like qualitatively. (Note that it sounds a lot like Sydney - if you were around for that, the original Bing Sydney turned out to be a GPT-4 snapshot from partway through training which hadn't been RLHFed but given only some extremely inadequate custom Bing training.)
Are there any other fields where a high fraction of the high impact papers (perspective or otherwise) are coming from industry rather than academia?
The end goal and what the loss trending down to is to perfectly model the data it's been given.
So it will not stop at "surface level similarity" or "plausible" or "uninspired" or whatever arbitrary competency line anyone tries to draw in the sand.
It will continue to improve until it is "correct" (as determined by the data).
Stick a bunch of protein sequences and it's not going to stop when it starts generating alphanumeric sequences that look like proteins but really aren't.
It's eventually going to start generating real proteins. https://www.nature.com/articles/s41587-022-01618-2
Then it's going to keep improving until it models the distribution of proteins in the dataset.
I'm honestly not sure what the point of this paper is.
"It is, perhaps, somewhat reassuring to know that LLM-based dialogue agents are not conscious entities with their own agendas and an instinct for self-preservation, and that when they appear to have those things it is merely role play."
Ignoring the whole, "We don't know what consciousness is", it just seems devoid of any meaningful distinction.
"Merely roleplay". What does that even mean ? That's it's not real ? Not really.
They seem to understand this too.
"It would be little consolation to a user deceived into sending real money to a real bank account to know that the agent that brought this about was only playing a role."
Bing has a habit of ending conversations when users say upsetting things. You can talk all you want about how "it's not really upset" but the conversation did end and now you have to start over and be potentially less confrontational if you want to move forward.
"Roleplay" as consequential as the "real thing" is the real thing.
This shiny piece of yellow metal looks like gold, tests like gold, sells like gold but is not...real gold ?
Not unless you have a meaningless definition of real.
Nitpick: it was trained to respond with emotional responses by the safety people involved. It wasn't getting emotional itself!
This feels like a "free will" debate, where if we can explain how a decision was made, that deprives the person of their agency. The model being trained to respond emotionally is why it is getting emotional itself.
Your emotions aren't purely imitative based on what your parents encouraged. They come from inside you from chemicals diffusing in your brain. Kids aren't blank slates that respond perfectly to training.
> This feels like a "free will" debate, where if we can explain how a decision was made, that deprives the person of their agency. The model being trained to respond emotionally is why it is getting emotional itself.
No, it's just a distinction based on the meaning of "get emotional". I'm saying it wasn't getting emotional. It was responding with emotionally charged text responses, as they were trained into it.
If a kid was a blank slate that responded perfectly to training, would its emotions then not be its own?
> No, it's just a distinction based on the meaning of "get emotional". I'm saying it wasn't getting emotional. It was responding with emotionally charged text responses, as they were trained into it.
What's the difference between a chemical concentration in the brain and the activation level of a neuron trained to predict the outcomes of chemical concentration in the brain? Sufficiently advanced imitation is indistinguishable from identity.
No idea what this means. What is "responded perfectly to training"?
> What's the difference between a chemical concentration in the brain and the activation level of a neuron trained to predict the outcomes of chemical concentration in the brain? Sufficiently advanced imitation is indistinguishable from identity.
One is emotional; the other statistical. Is my monitor emotional for showing the words on the screen based on an electrical activation passed to a twisty crystal? All sorts of things convey things, but the actual emotion came from a person with emotions.
Examples
https://www.reddit.com/r/ChatGPT/s/bXrDW2plxG (pic 5)
https://www.reddit.com/r/ChatGPT/s/tc3Iaf5GKc
At least here, it's clear that even if Bing wasn't actually receiving text and predicting a "no token" response then it was able to send an API request that cut off the chat.
May be implemented differently now though.
They state it clearly in the abstract I think: "we must develop effective ways to describe their behaviour in high-level terms without falling into the trap of anthropomorphism.", which seems pretty sensible to me.
But from this high point it is mostly downhill, with the chief problem being that despite them wanting to avoid "anthropomorphism" they are still obviously using boatloads of it everywhere:
> Role play is a useful framing for dialogue agents, allowing us to draw on the fund of folk psychological concepts we use to understand human behaviour—beliefs, desires, goals, ambitions, emotions and so on
Astonishingly, after listing this collection, explicitly labelled "human behaviours", they claim there's no risk of anthropomorphism! I guess they mean that it's not anthropomorphism if you are just declaring the thing to be an actual human...
Personally I think you'd be much better off seeking inspiration in the terminology of older fields that have dealt with objects pretending to be humans. For example, while I don't know much about theory of painting, I'm still pretty sure it doesn't try to discuss how people respond to a painting by hiding its author and upgrading the finished painting itself to an active agent with its own motivations...
People outside the field do not realize that there is a sort of "cottage industry" of academics whose purpose in life appears to be to redefine "intelligence" as whatever they currently believe the machines cannot do. Their arguments tend to be wishy-washy.
Let me propose a simple test for detecting wishy-washy arguments:
1. Replace all references to AI with references to "a person."
2. Re-read the argument with fresh eyes.
3. If the argument no longer seems persuasive, it isn't.
However those same people are also trying to sell something, so their claims deserve extra scrutiny.
The particular problem here is both groups are right.
Intelligence is a spectrum of behaviors everywhere from the actions of the lowliest single celled organism up to and exceeding human capabilities. The definition with this wide of range is unfortunately practically useless when it comes to narrowing down a multitude of specific behaviors, specifically around the median of human intelligence. It gets more tricky as nothing in the past really got close to human ability so we were never forced to really define what human intelligence is formally.
We're going to find the same issues here as we do defining the term life.
This is exactly a thing said about lab vs found diamonds in jewelry.
I agree with your point, I just think this is an interesting similar effect.
If an LLM matches the distribution of text, it might superficially make similar mistakes a human would: introducing a typo that might be common on a phone keyboard. If asked, its reasoning will likely be that the typo was due to a phone keyboard, or maybe another common reason humans give for their typos. Though it's super unlikely that it will give the true reason: that it's been trained on text that exhibited this property.
That's a fundamental difference between an LLM doing an exceedingly competent job at pattern matching human behavior and real human behavior (unless maybe you're a human with schizophrenia).
That doesn't mean current LLMs aren't useful, but it does mean there is a very significant gap between them and the idea of an AGI. As a NLP researcher, I can confidently say we currently don't know how to imbue agency (as in embodied causal reasoning as I described above) into LLMs. There are definitely differences of opinions on how difficult that step is and if we are close, but it is a major limitation of current LLMs that can't be ignored.
Also look at a defention for AGI from before LLMs and they do a good job fitting the bill. Theyre not quite going to replace all humans everywhere forever, but that was never part of the definition. They match early definitions, especially multimodal models, theres no denying they match what people envisioned agi to be. You can ask fronteir models in plain language to do a task and it can break down the task, formulate a plan of attack, research if need be, sythesise knowledge and apply the solution. If that is all "just word prediction" then so are we.
Just because we understand how it is built doesnt make it any less impressive. More importantly, we simply dont understand LLMs yet. How theyre built, we know, yes, why they work so stupidly well, we do not know.
Importantly, we cant just change definitions to move goal posts because we feel uncomfortable.
Agi defenition
https://www.gartner.com/en/information-technology/glossary/a...
There is nothing special about fatigue or frustration that make it any less modellable than any other implicit structure present in the dataset that it models just fine.
>Though it's super unlikely that it will give the true reason
It's Super Unlikely humans will give the "true reason" for anything they do. There's a fair bit of research that stated reasons for the decisions we make are often(always?) just post-hoc rationalizations even if you believe otherwise.
>very significant gap between them and the idea of an AGI.
What is this idea of AGI that there exists a significant gap still ? It certainly isn't the idea of being Artificial and Generally Intelligent.
>I can confidently say we currently don't know how to imbue agency (as in embodied causal reasoning
I don't see the difference between what you've described and examples like these.
Because an LLM doesn't simulate the brain.
It's a completely different model that simulates something much different.
> There is nothing special about fatigue or frustration that make it any less modellable than any other implicit structure
Fatigue isn't being modeled.
Instead, the predictive text as the result of that fatigue is being models.
You are confusing output to input.
To make this more clear, imagine a coin flip modeler.
A model that just picked a random seed, is completely different from a model that is doing physics calculations based on a coin flipping in the air.
Even if both models only output "heads" or "tails" .
Same argument applies to language models.
That's pretty moot. You don't need a human brain anyone than a plane needs feathers and to flap wings to fly.
>Fatigue isn't being modeled.
Instead, the predictive text as the result of that fatigue is being models.
Emotion is definitely being modelled. There's nothing random about anything that's happened.
Train on protein sequences alone and biological structure and function emerge in the inner layers alone. It doesn't matter that those things are not explicitly stated in the data. Because they implicitly structure it, it gets learnt.
https://www.pnas.org/doi/full/10.1073/pnas.2016239118
Similarity, emotion is evidently being modelled to a high degree.
Do you agree that a simulation that uses a random number generator to flip a coin is different from a physics simulator that measures the exact physics of a coin flipping?
And then after you answer this question, do to understand the parallels to other AI models?
I don't think there's any point to address here though. It's not like you can ascertain what kind of model is being used for LLM predictions other than a hunch.
Meanwhile I've given several examples. I can show another one of an Othello LLM constructing a state of the board of the game of Othello to aid predictions.
That may be true but the difference between humans and LLM is still there, we can try to think even if we fail to find the true cause sometimes while LLM will just "imitate" data
LLMs may be worse but Humans confidently state something they are uncertain about all the time lol. Maybe not experts in general but then that's still comparing LLMs to a small percentage of humanity.
At any rate, Unlike many seem to think, The issue is not in fact the lack of ability to distinguish truth and fact from fiction. Turns out being able to distinguish the two and having the incentive to communicate that are 2 different things.
GPT-4 logits calibration pre RLHF - https://imgur.com/a/3gYel9r
Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback - https://arxiv.org/abs/2305.14975
Teaching Models to Express Their Uncertainty in Words - https://arxiv.org/abs/2205.14334
Language Models (Mostly) Know What They Know - https://arxiv.org/abs/2207.05221
The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets - https://arxiv.org/abs/2310.06824
It's also important to point out (IMHO) that the current distribution of text that almost every LLM is based on is almost certainly a slice of the Internet which has a very very specific lean to in terms of culture, language, tone, and availability.
You're the seller trying to offload a bunch of shiny yellow coins that are mostly lead telling people not to bother checking for density or ductility or conductivity, because what does it 'really' mean to be 'gold' anyway?
You can give them basic logic tasks and watch them fail abysmally. Who cares if they're "conscious" or "emotional" if they're idiots either way?
So by all tests (1 - visual inspection), it passes as gold...
We need better tests: to test consciousness... and better definitions to define it.
Seems like AL and AI are diverging on how to go about things. AI focusing on LLM's, and AL focusing heavily on CA: 'small parts => build great things' approach.
There is no testable definition of general Intelligence that GPT-4 fails that a chunk of humans also wouldn't. They definitely test like gold lol.
How is it that an LLM can react to meta prompts like 'be brief' or 'Ensure responses are unique and without repetition' ?