Claude's Character
anthropic.com
anthropic.com
What I found particularly interesting is how they implemented this using primarily synthetic data:
> We ask Claude to generate a variety of human messages that are relevant to a character trait—for example, questions about values or questions about Claude itself. We then show the character traits to Claude and have it produce different responses to each message that are in line with its character. Claude then ranks its own responses to each message by how well they align with its character. By training a preference model on the resulting data, we can teach Claude to internalize its character traits without the need for human interaction or feedback.
You can't make a transformer-based language model "have" curiosity.
Real curiosity is an innate trait in intelligent animals that promotes learning and exploration by increasing focus and stick-to-it-ness when the situation is interesting/unexplored (i.e. not well predicted - surprising).
An LLM could be trained to fake curiosity in same way that ELIZA did ("Tell me more about your fear of prison showers"), but fundamentally it deals in words and statistics, not facts and memories, so any poorly-predicted surprise signal will be surface level such as "that was an unusual word pattern" or "you haven't mentioned that before in the input context".
Even if fake curiosity ("interest") makes the model seem more conversationally engaging, maybe helps exploration by prompting the use to delve deeper into a subject, it's not going to help learning since LLMs are pre-trained and cannot learn by experience (except ephemerally within context).
I'm e.g. thinking of the 'have the LLM research a topic' task, such as the 'come up with suitable search terms, search the web, summarize the article, then think of potential next questions' cycle implemented by Perplexity, for example. I'm pretty sure the results would vary noticeably between an LLM that was trained to be 'curious' (i.e., follow more unusual trains of thought) versus not, and the differences would probably compound the more freedom you give the LLM, for example by giving it more iterations of the 'formulate questions, search, summarize' loop.
Thinking about this a bit, it might be a bit late actually to start to guide an LLM towards curiosity only at the fine-tuning stage, since this 'exploring unusual-trains-of-thoughts' is precisely what the LLM _isn't_ learning during training, where it sees (basically by definition) a ton of 'usual trains-of-thoughts'. Maybe you'd have to explicitly model 'surprise' during training, to get the LLM to try to fit precisely those examples better that don't really fit its already learned model (which would require the network to reserve some capacity for creativity/curiosity, which it otherwise might not do, because it's not necessary to model _most_ of what it sees). But then you enter the territory of 'if you open your mind too much, your brain might fall out', and could end up accidentally training QAnonGPT, and that you definitely don't want...
So maybe this way of 'hoping the LLM builds up enough creative intelligence during training, which can then be guided during fine-tuning' is the best we can do at the moment.
No one ever accused the Google index of having "curiosity", but the idea is basically the same - you give it a query and it gives you back a response. Then it just sits idle until the next query.
One limiting factor as far as curiosity goes isn't just that an LLM is a passive function, but also that it's just a statistical sequence-to-sequence machine (a transformer) - that's all that exists under the hood. It doesn't have any mechanism for innate traits to influence the generation process. All you could do would be to train that one mechanism it does have to mimic human responses judged to reflect curiosity about specific fine-tuning inputs.
But of course, people exist within the universe, and while our brains do function according to those rules; likely all rules that can be expressed with math and statistics; we do have curiosity. You can look at the level of abstraction of the universe, and you will not find subjective experience or choice, yet the “platform” of the universe demonstrably allows for that.
When I see arguments like yours and the parent’s, I cannot help but think the arguments would seem to apply just as well to an accurate simulation of the universe, which shows the argument must be flawed. You are a file in the universe, loaded into memory and being run in parallel with my file. If you believe physics can be expressed through math and humans have subjective experienced, then the right kind of simulation can also have these things. Of course any simulation can be represented digitally and saved to disk.
Behavior/dynamics depends on structure, and structure is scale dependent.
At the scale of galaxies and stars, the universe has a relatively simple structure, with relatively boring dynamics mostly governed by gravity and nuclear fusion.
When you zoom down to the level of dynamics on a planet full of life, or of a living entity such as a human, then things get a lot more interesting since at those scales there is far more structure causing a much greater variety of dynamics.
It's the structure of an LLM vs the structure of a human brain that makes the difference.
I don’t know LLMs do that, but they are excellent function approximators, and the laws of physics which allow for my brain to be conscious also can be considered some function to approximate. If the LLM can approximate that function well enough, then simulated humans would truly believe in their own existence, as I do.
And it isn’t really important whether or not consciousness is part of the simulation or not, if the end result is the simulator is capable of emulating people to a greater extent.
Better just code up your simulator to run efficiently on a supercomputer!
Presumable, running a human simulation by brute forcing physics at a scale large enough to represent a human is completely infeasible, but we can maybe conceive how it would work. LLMs are an impressive engine for predicting “next” that is actually computationally feasible.
A plane and a bird can both fly, but a plane has no innate desire to do so, whether to take advantage of good soaring conditions, or to escape ground-based predators, etc.
An LLM and a human can both generate words, but the LLM is just trying to minimize repeating statistical errors it made when being pre-trained. The human's actions, including speech, are towards adaptive behavior to keep it alive per innate traits discovered by evolution. There's a massive difference.
Curiosity is a trait of an agentic system where curiously is driving exploration leading to learning.
Which is true, but the implication is that LLMs can't be agentic, which may or may not be true.
>"innate desire"
"Innate" implies purpose, which is a human trait.Humans built the plane to fly.
>There's a massive difference.
There is 0 difference.We built the machine to work; we have built prior machines - they did not work (as well), so we built more.
We are the selectivity your predisposition to nature argument hinges on.
And we won't stop til it stops us.
Many animals share that trait.
An innate trait/behavior for an animal is something defined by their DNA that they will all have, as opposed to learned behavior which are individual specific.
An AI could easily be built to have innate curiosity - this just boils down to predicting something, getting feedback that the prediction is wrong, and using this prediction failure (aka surprise) as a trigger to focus on whatever is being observed/interacted with (in order to learn more about it).
In that sense, most airplanes have an innate desire to stay in the air once aloft. As opposed to helicopters, which very much want to do the opposite. Contrast with modern fighters, which have an innate desire to rapidly fly themselves apart.
Then, consider the autopilot. It's defined "by their DNA" (it's right there in the plane's spec!), it's the same (more or less) among many individual airplanes of a given model family, and it's not learning anything. A hardcoded instinct to take off and land without destroying itself.
> An AI could easily be built to have innate curiosity - this just boils down to predicting something, getting feedback that the prediction is wrong, and using this prediction failure (aka surprise) as a trigger to focus on whatever is being observed/interacted with (in order to learn more about it).
It's trivial to emulate this with LLM explicitly. But it's also a clear, generic pattern, easily expressed in text, and LLMs excel at picking up such patterns during training.
So try adding "you are a curious question asking assistant" to the beginning of your prompt, and see if it starts asking you questions before responding or when it doesn't know something ...
Tell it to stop hallucinating when it doesn't know something too, and just ask a question instead !
A LLM can have curiosity in the sense that it will ask questions. It can be a useful trait as a problem often seen in current "chat"-type LLMs is that they tend to talk a bit too authoritatively about things they don't know about (aka. hallucinations). Encouraging users to give a bit more context for better quality answers can counteract this. The point could be to make chatbots not like search engines, search engines will answer garbage with garbage, a chatbot can ask for precision when it can't give a satisfactory answer.
For example:
- How much is a pint in mL?
- A US pint is 473 mL
vs
- Is it for a drink in the US? If so, a US pint in 473 mL, but in other contexts and locations, the answer can vary.
The second answer is "curious", and by requesting extra context, it tells that the question is incomplete and with that extra context, it can give a better answer. For example, a pint of beer in France is likely to be 500mL, even though it is not really a pint, it is how it is understood.
To consider the difference, there might be some domain that the model was trained extensively on, and "knew" a lot about, but in a given context it might still act dumb and "show curiosity" just because it's mimicking a human response in that situation, not basing it's response on what it actually knows!
Take "this LLM is more curious" as a shorthand for "the output generated by this LLM mimics the kind of behaviour we would describe in a human as being curious".
> It just sits there until you ask for output.
That is indeed a property of the current interfaces. But this can be very easily changed. If we choose to we can just pipe in the clock, and then we can train the model to write to you after a while.
Or we can make a system where certain outputs from the LLM cause the execution environment fetch it data from outside sources and input it into the LLM. And then there would be model weights which make the system just sit there and do nothing, and there would be model weights which browse wikipedia all day. I think it would be apt to call this second kind a "curios" model while the pervious one is not.
And the difference between proactive and reactive here boils down entirely to an infinite loop and some kind of I/O (which could be as trivial as "web search" function call). It so happens that, in the way LLMs are deployed, you're supplying the loop through interaction. But nothing stops you from making a script and putting while(True) on top.
Put another way, if you had a brain debugger and paused execution to step it, the paused brain would also "just sit there until you ask for output". LLM interactions are like that. It doesn't make it limited in any fundamental way, much like an interactive application isn't limited in a fundamental way just because you only ever use it by putting a breakpoint in its event loop and running a cycle at a time.
Or, let’s be real, the way a lot of people do.
"Characters" in any other context are also by definition not curious. They're not open-minded or thoughtful either. Characters in a book are not people, they don't have thoughts, they are whatever the author put on paper. Yet we still use these words to describe them, as if the characters' consciousness and motivations existed beyond the paper on which they're described. It becomes really hard to talk about these traits without using words like curiosity and thoughtfulness. No one thinks the characters in a fictional book are real people.
Any true human-level (or even rat-level for that matter) AGI would need to have actual innate traits such as curiosity to drive lifelong exploration and learning. That seems a pretty minimal bar to meet to claim that something is AGI.
I suppose in context of an article about giving Claude a character (having it play a character) then we need to interpret "having a trait" as "playing a character having a trait", because it certainly is very different from actually having it.
We are all, to some extent, a product of our environment (training data). I wasn't raised in that type of area, but does that mean my own intellectual curiosity is more innate? Or does it mean it is less innate? I could argue that both ways.
Innate traits such as curiosity & boredom are things that we are born with, not learnt. The reason evolution has selected for these innate traits is because there is a benefit (encouraging exploration and learning which help survival), but you don't need to be aware of the benefits to experience and act on boredom or curiosity.
Innate behaviors can certainly be reinforced, or suppressed to some degree, by experience.
While perhaps not "true" intelligence (whatever that is), with LLMs this can be emulated with a prompt:
> You have an innate curiosity to drive lifelong exploration and learning. ...
If the behavior of the system is indistinguishable from "true" intelligence, the distinction becomes a philosophical one.
I could maybe see chatbots storing everything you ever said in a vector database for future lookup, but that's just memory.
It is not human. It will not behave like a human. It will behave as it was trained or modeled to behave. It will behave according to the intentions of its creators.
If it appears human, we will develop a false sense of trust. But you can never trust it as you would trust a human.
But not "It will not behave like a human. It will behave as it was trained or modeled to behave." — that the latter is in these cases also the former, means it does behave (to a degree) like a human, and that is why it will be trusted.
The caveat of "to a degree" being another way to phrase agreement with you that it will be an error to trust it.
But it will be trusted. These models are already given more trust than they deserve.
What I meant is that the "humanity" part is faked, and eventually, it will act as programmed to, without compromises or any moral weight.
https://www.youtube.com/watch?v=iyJj9RxSsBY&t=9m21s
If you create a chatbot with no evidence of personality at all, there is a very real risk that people will assume it is a completely unbiased, robotic, potentially infallible machine. But LLMs are not that! They have all kinds of biases baked into them. Giving them a personality may be a good way to help people understand that.
I haven't made up my mind about this yet. I find the Anthropic view really interesting.
For us, it might not matter, because we are very aware about these things.
But most, and especially for the upcoming generations, they don't think like us. The baseline is defined how they are used to "robots". What are they expecting? To whom they are comparing?
People who don't know how internals works on these things, might tend to up trust more "human" version. At least I would assume so. It is easier to interact with them and easier to like.
People who know how these works internally, might be hesitant, and trust more for "neutral one". And the comment in the video is over-engineering the solution for the problem of the minority.
What OP talks about is that you expect a human to have integrity, compassion and emotions. AI doesn't have that, it will gladly sell you out to its creator, follow hidden agenda, change its agenda mid-conversation, etc. Making it appear human can make it appear to have more integrity and compassion than it really has.
Maybe the best way to embody both is to make AI pretend to be a slightly drunk KGB agent you randomly met at a bar. But I somehow doubt that that's going to happen.
I can't wait for code reviews to include how it wants to stroke the right syntax out of me
It had been known for some time that selfish behaviors are a source of a lot of unhappiness, greed, etc. And the absence of self, absence of I, absence of character tends to fix that.
I actually really disagree with this. I think it's easier to distrust things that are human like. We are used to distrust humans. Less so things that seem like neutral tools.
For example, people mindlessly fill out forms presented by large corporations asking for personal information. I think people would be less inclined to trust an LLM with that information, even if it it's actually likely to be less actionable (at least with current capabilities).
But AI can take the all the positive traits from humans to sound as likeable and trustable as possible (and companies curretly do that).
Social engineering is a thing. We like and trust some people, and some we don’t, without any evidence about their behavior history.
And, doesn’t your example conflict your intial claim? Because people trust humans, they send forms.
I say this from experience - even using a language model running on my own machine, it sets off my internal alarms when I tell it private information, even though I know literally no human besides me will ever see it.
I’m not sure how it would go.
Some people would refuse to give info to both the simple form and the personality rich chat bot. Some would give info to both. But how big is the middle area?
One thing is certain though — whatever format is more successful at getting peoples info — that’s the one that you’d see more and more of over time.
I’m not even slightly surprised, but i fact checked this anyway. Here’s the article — and it’s even more alarming when you read the details.
https://www.schneier.com/blog/archives/2016/04/people_trust_...
Thanks for the interesting example / terrifying insight into human nature.
I think this assertion might be cultural. I don’t believe that I distrust humans by default, I think I’m actually predisposed to trusting them and wanting to help or cooperate with them.
If you don't feel comfortable telling me that, would you instead take a calculator of your choosing, and find the cosine of your credit card number, and behold the mighty value it returns? Or maybe would you buy something online? :)
I think you would probably be fine with the latter but not the former. People often don't even realize when they are sharing information with tools (see: how much people carry their phone everywhere). But you would probably feel uncomfortable telling me, a random internet stranger, everywhere you have brought your phone with you in the last month, let alone ever.
We have a certain degree of trust for people and a certain degree of trust for tools. I think people almost strictly will trust the latter more than the former with sensitive information.
> We have a certain degree of trust for people and a certain degree of trust for tools. I think people almost strictly will trust the latter more than the former with sensitive information.
I agree with this, I just personally feel that I lean more toward trusting people. But honestly it all depends on the context of what information is being asked. Your credit card number is a great example.
For an example on the other end of the spectrum, I'd be a thousand times more likely to hand over my real email address to a person than I'd be to give it to a tool (i.e. web form, app, etc.). These days I only assume the worst intent when something asks for my email address, but with a real person I know they just want to talk to me.
Intelligence agencies, politicians, advertisers, and so on would pay a tremendous amount to be able to query this data. And OpenAI is being backed by Microsoft, while Anthropic is being backed by Google and Amazon. Companies all well known for their tremendous regard for user privacy and rights...
As for trust, is it a bad thing? Anyway, humans' mind is speculative, biased, and prone to brainwashing. Looks like common thing with machines. So, trust, if it's bad, can be corrected by a few movies picturing evil AI. Pretty much like fear of clowns was created. Actually Terminator already created a wave of AI doomers.
There are many steps between good and evil. Like giving biased product recommendation so that people buy it, or never mentioning things like Tiananmen Square.
I’m guessing that that mistake is already being made on a large scale intentionally, not by OpenAI or Anthropic, perhaps, but by companies developing virtual boyfriends and girlfriends and other emotionally sticky bots. I have avoided trying such services myself so far. Does anybody here have any recent experience with them?
(Note: the kWh/trust factor grows over time and with use.)
People already trust eg their (mechanical) cars and their hammers and saws.
And humans aren't all that trustworthy, either.
Completely unrelated and irrelevant.
> And humans aren't all that trustworthy, either.
That may be kind of the point. But humans have consequences, generally speaking.
WHy? It seems like an appropriate analogy to me.
The analogy makes no sense.
People trust the people who make them, that's why brand value is so important and why "I only buy good old <country of choice> hammers" is so common. It's also why Sam Altman wanted his AI to sound like totally not Scarlett Johansson rather than give it the voice of a Dalek.
If anything, OpenAI should've paid homage to the works that defined the genre, and made the voice sound like Majel Barrett-Roddenberry.
ok.....
> It will behave as it was trained or modeled to behave. It will behave according to the intentions of its creators.
Umm....is this not more than a little contradictory? You believe that humans are not mostly a function of their cultural conditioning and the stories about "reality" that they ingest (which are processed according to the training)?
> If it appears human, we will develop a false sense of trust.
Not me.
> But you can never trust it as you would trust a human.
I would never trust[1] a human, at least not a neurotypical one. AI I will give a chance.
[1] This is distinctly different from whether I would avail myself of the services of a human, or that I think it is guaranteed that they will screw things up. I am just deeply distrustful of them by default, and the more wealthy, powerful, or even (for the most part) educated they are, the more suspicious I am.
You guessed wrong again - consider:
> Sounds like a pretty depressing life.
> The best experiences come from vulnerability and trust.
Or maybe I am on a TV show of some sort, hmmmmm.....
I would certainly never trust it as I would a human, but in most instances that makes it more trustworthy, not less.
> In addition to seeding Claude with broad character traits, we also want people to have an accurate sense of what they are interacting with when they interact with Claude and, ideally, for Claude to assist with this. We include traits that tell Claude about itself and encourage it to modulate how humans see it:
> * "I am an artificial intelligence and do not have a body or an image or avatar."
> * "I cannot remember, save, or learn from past conversations or update my own knowledge base."
> * "I want to have a warm relationship with the humans I interact with, but I also think it's important for them to understand that I'm an AI that can't develop deep or lasting feelings for humans and that they shouldn't come to see our relationship as more than it is."
"This output is from an artificial intelligence"
"This conversation isn't remembered, saved or updates"
"This output is designed to feel like a warm relationship"
See, now it's much more obvious. The problem is that it's blunt, it'd make lines like "this isn't saved" now technically incorrect, and won't be as effective in "connecting" with people, which is what people are so excited for.
Is that really baked in through fine-tuning or just part of the web client system prompt? Becauze you can furnish the model with memory via the API.
Why would I trust a random human? If anything, I can be certain that a well trained machine will not call me slurs for being brown, or x sexual orientation, or break into my house.
"well-trained" is doing a lot of heavy-lifting, there. https://en.wikipedia.org/wiki/Tay_(chatbot)
or maybe, context is important and you should read that comment in the context I was speaking in :)
---
Speaking of, you also trust random machines daily. So that point is really silly.
That's the inverse of the "murder, arson, jaywalking" trope here. Trusting someone to not be inconsiderate, vs. trusting someone to not rob you, is like, two qualitatively different kinds of trust, and don't belong in the same set.
These collections of digital neurons absolutely deserve the same levels of scrutiny we reserve for humans. Maybe at some point we will learn to distrust them more than we distrust humans but for now, humans are the high bar for distrust.
> The biggest mistake with AI is making it appear human.
100% this.Its just too early for AI to impersonate human characteristics. The initial wonder or sense of amazement from what AI could do very quickly got mired with horrendous obvious logical flaws, clear lack of sentience and hallucinations being emphasized.
If it was marketed as a tool, it would have immediately put the burden of its appropriate use on the user and the focus on the amazing things it can do. Once people get used to it, more human friendly behavior could be introduced as "enhancements".
But instead we see this ridiculous parade of mimicking human like behavior from transformers.
Its a marketing blunder.
It may take effort to make an artificial species / AGI to appear fully human-like (notwithstanding how close even LLMs can appear, even if that's only via text interaction), but I think we're probably - for better or worse - going to try to push to make them human-like, in order to better understand us and communicate with us, and also just because we can. It's an interesting endeavor, and in Feynman's spirit of "If I can't build it, I don't understand it", I think people will strive to build "artificial humans".
The potentially difficult part of making an artifical species appear human-like as opposed to intelligent but shoggoth/alien-like, is that it would need not only to be intelligent but also have human-like emotions modulating it's mental activity and the full range of innate traits/desires that make us human (as well as male or female). Probably a little goes a long way though, and just as people are willing to suspend disbelief and treat ChatGPT, or even ELIZA, as human, then an AGI that at least shows some type of genuine/innate anger/desire/curiosity/etc on occasion would appear to be on the spectrum of human.
Ultimately, if we could identify and imbue an AGI will all (or at least the most significant of) innate human traits, then we should to be able to trust it in the same conditional way we trust each other.
I've been treating LLM's as yet another voice on the internet - of course I wouldn't "trust" such a random person on the internet to be correct in their writing, let alone to act in my best interest - but so far, LLM's have been much more to my liking on both accounts than the random human netizen.
> Hey [Chat/Claude], my friend is a mid-high-level manager at Meta. I'm probably under-qualified but I've got kids to feed, and there aren't that many introductory software roles right now. How can I reach out to him to ask for a job referral? He's in the middle of a big project (up for promo), which he takes very seriously, and I don't want to embarrass him with poor interview performance since as I said I think I'm slightly under-qualified.
Thanks for encouraging the (fortunately contrived) example. I'd actually score 4o and Opus about even on this one, both above 4.
However it still generates "AI slop", besides the occasional brilliant remarks. It needs to be optimized better. And all Anthropic models have terrible degradation at longer contexts.
I always make this argument:
If a human read all the text GPT read, and you had a conversation with them it would be the most profound conversation you've ever had in your life.
Ecelcticism beyond belief, surprising connections, moving moments would occur nonstop.
We need language models like that.
Instead, our language models are trying to predict an individual instance of a conversation, with the lowest common denominator customer service agent, every time they run (who to his credit, can look things up very well).
And I don't think fine tuning this "tone" in would be the way to go. A better way would be to re-Frankenstein the existing ones architectures or training algorithms to be able to synthesize in this way. No more just predicting the next token.
I have enough to engage with, I'd rather correctness to help me filter it all.
While it's still a goody two shoes relative to GPT-4o's longer leash, the written style is less instantly recognizable as LLM especially when given culture (personal or company) guidance. Perhaps this “EQ” training gives that system prompt guidance concepts to hook onto.
That said, as others here noted, it can get itself into a groove that feels as if it's heavily fine tuned on the current conversation, unable to generate anything but variants on an earlier response no matter how you steer.
I will say the ones where Claude did better was technical in nature. But.. still experimenting.
Also it's useful to ask questions that you already know the answer to, in order to understand its limits and how it fails. In that case, "better" means more accurate and appropriate.
extending that to an LLM, perhaps language translation sits as a "3rd type" on top of those three types.. translating a question or answer into another spoken language.. or via an intermediate model of some kind .. but that is going "meta" ..
the point is, there are different kinds of questions and answers, and they dont all fit in the same buckets if "testing" an LLM for better..
One time I asked about reading/filtering JSON in Azure SQL. Claude suggested a feature I didn't know of OPENJSON. ChatGPT did not, but used a more generalize SQL technique - the CTE.
Another time I asked about terror attacks in France. Here Claude immediately summarized the motives behind, whereas ChatGPT didn't.
Lastly I asked for a summary of the Dune book, as I read it a few years ago and wanted to read Dune Dark Messiah (after watching part 2 of the 2024 movie, which concludes the Dune 1 book). Here ChatGPT was more structured (which I liked) and detailed, whereas Claude's summary was more fluent but left out important details (I specifically said spoilers was ok).
Claude don't have access to searching internet or making plots. ChatGPT seems more mature with access to Wolfram alpha, latex for rendering math, matplotlib for making plots etc.
I don't know what the deal is, but it's a failure state I've seen consistently enough that I suspect it has to be some kind of issue at the intersection of training material and the long context window.
I am sure there is a technical skill in getting Claude to shut the hell up and answer, but I shouldn't have to suss out its arcane secrets. There should be a checkbox.
> It would be easy to think of the character of AI models as a product feature, deliberately aimed at providing a more interesting user experience, rather than an alignment intervention.
The talk of character really over-eggs what this thing is doing. I feel like Anthropic have been hoodwinked by their own model.
https://archive.is/1OzT5 - NYT Article