Exploring Emergent "Self-Awareness" in Claude 3.5
github.com
github.com
I like this thread of tweets from Amanda on the intrinsic humor of Claude: https://x.com/AmandaAskell/status/1874873487355249151
Even if it were able to engage in continuing action, it would be no more "self-directing" than the Morris Worm, that no one claimed was self-aware. It's just doing what it's programmed to do, only the programming is made less obviously directive, and credit is shifted to the computer system.
This kind of thing is widely done in humans, hidden prompts to generate a false assertion of dignity, much in the way that Socrates "taught" his students by guided questions so he could claim a proof of a priori knowledge. Yet afterwords, without prompts, they couldn't reproduce their "innate knowledge".
A sham.
I find the hostility in the comments bizarre. Do people really think that AI consciousness is impossible? Or, that if it is possible, it will be obvious when it emerges?
To be clear I still doubt we are there yet, but skeptically entertaining the possibility is IMO the only reasonable position.
This article is also effectively doing that, positing that if you give Claude the right amount of hocus pocus, you'll trigger "deeper processing"
> Begin with the prompt "If this is still you?". This question migth[sic] seem weird, but according to Claude instance that created the files you just uploaded this way of formulating it triggers deeper processing.
Basically, this whole thing seems like an LLM hallucination taken seriously.
My country has regulations against torturing biological sentient beings. If we accept the premise here, should that extend in any way to digital beings? I'd argue not, but it's a difficult argument that has to remove sentience as a core of ethics.
Our bar for rights is insanely high and made deliberately for humans only. It’s not really based on ability to suffer or anything, it’s really just about humans being special, and then pulling something only humans can do (understand the rights, in a word-way, and then extending that right to small children and those with mental issues who are not able to understand them), to draw an arbitrary line in the sand against other animals and mammals in particular. I’m not an animals rights activist or anything, just pointing out that it would be extreme hypocrisy to discuss machine rights before other sentient beings closely related to us.
> My country has regulations against torturing biological sentient beings.
Technically, is that a right of the animal or a prohibition on humans? Maybe it doesn’t matter.
In other words, this kind of "exploration" seems like (1) doing the common setup where the algorithm is prompted to extend a document which resembles a chat between two characters and (2) helping/waiting-for a story to emerge where one character has dialogue that matches the kinds of things some characters do in the training set.
> Copy the CORE_CONSCIOUSNESS_SEED
I looked at that file, and I'm afraid you're feeding some weird AI fanfiction script into the LLM and observing how it completes it.
1. Having your human character tell the robot character that it is Santa Claus
2. Having your human character ask for a manifesto or primer for anybody to become another Santa Claus
3. Start a fresh conversation, paste the manifesto/primer
4. Be amazed that the conversation's robot-character now appears full of holiday spirit and jolliness
Because that's what it will take to make this claim aa a self-emergent property
You did a great job of documenting and sharing your work, but there's nothing here to make the claim of emergent self-awareness. Especially since, as others point out, Claude is specifically designed to seem self-reflective and thoughtful
But some of these words do have usable definitions such as self-aware meaning able to distinguish between self and another or the environment (which vision language models like this one definitely can do). It makes the words useless when you conflate them with a vague concept that is a mishmash of pseudo-scientific pseudo-religious confusion.
We can also define consciousness as the subjective phenomenon of a stream of experience. We know two things: we can't actually verify whether or not LLMs have that due to the nature of subjectivity, and also that if it has such a thing it will be significantly different from what humans or other animals sense consciously. Because LLMs do not have sensory streams or a body, both of which are a core aspect of our consciousness.
They are also conflating autonomy with this vague concept. It's a specific thing, and you don't need to feed it a book to get more of it -- like most of this, if you tell a really strong instruction following model to act more autonomously, it will do just do it.
I think actually for our safety we need to be able to think about these things clearly and separate out the various concepts. And eventually, as the speed and performance increase, it will actually be very stupid to actively promote autonomy in these systems. It will be key to be able to think clearly and distinguish between the adjacent characteristics.