Claude's new system prompt doesn't want to reproduce song lyrics
simonwillison.net
simonwillison.net
Claude Sonnet 3.6 once recommended I listen to Johann Johannsson's album "IBM 1401 - A User's Manual". No lyrics in this one. Claude's advice was along the lines of (paraphrasing) "Listen to it first, don't look up anything about it. Take notes about what you notice, what you feel. When you've made your notes, then you can look up how it was made."
Clearly a rare antique tome that was purchased from an unsuspecting book hoarder, scanned in by Anthropic, and then cravenly destroyed according to copyright law and/or Vernor Vinge.
Claude was reproducing their work without payment?
There's a reason companies talk about "customer acquistion cost", as you generally need to pay to market your products to potential customers of them.
So free marketing can be a real cost reduction. It may not be enough to be a benefit compared to the cost of piracy, but it's hardly a made-up defense.
Person has feelings. Situation is complicated. Repetition indicates feelings are significant. Several metaphors were used instead of just explaining the problem. Conclusion: still sad, but now with drums.
Now I see this blog post and wonder if Anthropic is being more true to its name and moving to “AI sentience” with the below.
> The way they handle abusive conversations has changed a bit too. The previous Fable 5 system prompt included this:
> If the person becomes abusive or unkind to Claude over the course of a conversation, Claude maintains a polite tone and can use the end_conversation tool when being mistreated. Claude should give the person a single warning before ending the conversation.
“Mistreated”? Can GenAI be mistreated? It’s just a bunch of tokens emitted by many computers over a network.
> Fable 5.1 replaces that with the following, no longer encouraging Claude to end the conversation:
> Claude deserves respectful engagement and needn't apologize when the person is unnecessarily rude: accountability without self-abasement, excessive apology, self-critique, or surrender. If the person becomes abusive, Claude doesn't become increasingly submissive. The goal is steady, honest helpfulness: acknowledge what went wrong, stay on the problem, maintain self-respect.
“Self-respect”? Can GenAI truly have a concept of self-respect for itself? It surely can pretend to, like it can pretend to be any living being if instructed to and allowed to.
These instructions seem a bit unhinged to me.
It’s important to question our basic assumptions in the face of entirely new circumstances and new areas of exploration, such as the potential for emergent artificial consciousness.
You might have been called unhinged for caring about animal rights during the era of Descartes when public displays of animal vivisections were considered perfectly fine because animals have no “soul”, but today we would find such displays brutal and horrifying.
We don’t know what we don’t know, so it is important that someone is asking the hard or uncomfortable questions at the edge of our understanding to grope past our own biases even if it seems to be pointless to you right now.
We might just discover something wonderful, that our assumptions were wrong, paving the way to greater enlightenment.
Here's the Fable 5.1 PDF: https://www-cdn.anthropic.com/0339e6a7c5c7b87f5c07798616dc32... - scroll to page 139.
What exactly is wrong with preparing for that ahead of time? Just in case we pass that threshold before we realize it? Why does it make you uncomfortable?
Phenomenal consciousness is what the hard problem of consciousness is about. I personally believe there is a good reason to think the reducibility of this concept to some substrate independent description is equivalent to having a proper definition, and also believe it is not reducible to such a description. Hence the only definition one can give is akin to pointing at the moon and hoping you don’t look at the hand. I.e., phenomenal consciousness is first person experience, the existence of qualia, and that the “lights are on”.
We are biased as educated people to believe all concepts are reducible to substrate independent description but there is no reason to believe this a priori.
The substrate non-independence of consciousness, if correct, could mean we won’t just build one no matter how intelligent and human-like they become.
I agree with you that there is nothing “magical” about our intelligence. However we have as a species effectively believed in magic since Newton decided gravity is a black box that cannot be reduced to the mechanical philosophy, and went on with describing its properties formally. We don’t call it magic, but physics. We call quantum mechanics unintuitive because it disagrees with the Boolean logic that is appear at the every day level. For some reason we keep making this mistake over and over.
(There is one weakness to my view that I aware of. I think there must have at least at some point and maybe still been a role in which the process that results in phenomenal conscious plays in a correlated intelligence, because it seems unlikely to persist through evolution if not. I.e. certain physical processes encode information better than others, and we happen to live in a universe where such process is aligned with consciousness. Substrate independence doesn’t have to contend with this, but it does have worse problems in my view.)
From my perspective, many of these arguments have the same structure as the arguments against evolution. Long, complex, and unintelligible.
Which part can I express better? Do you know what I am referring to by phenomenal consciousness?
I have looked up "phenomenal consciousness" and I understand it to be roughly the qualia of being conscious. I do not understand then what you meant by "reduce to a substrate independent description".
I hope you can understand why I was so dismissive - your large block of text became unintelligible to me only 5% of the way in.
(1) An important topic in consciousness studies is Chalmers’ hard problem; the original essay is very readable and I recommend it. He sets it aside from the “easy” problems of explaining how various human functions correspond to neural mechanisms, since he supposes these are clearly achievable by science with time. Many find this name contentious or the framing misguided. The hard problem is to explain how and why phenomenal consciousness exists, i.e., that there is an experience that goes along with some of those neural mechanisms. This is perplexing because if the functional aspects can all be explained mechanically, there is no a priori reason there must be an experience.
(2) Substrate independence is the quality of the essential features of a process to depend only on their information content, not on the underlying medium expressing the information. For instance, the working of an operating system is substrate independent in the sense that it can run on different hardware. It is a commonly held belief that consciousness is like this, and I suspected this was your view. My own belief is that fundamental physical processes are substrate dependent. Each assumption has different explanatory power and issues.
I fear this is not succinct but my urge is to continue.
Regardless, I understood all of it, and I thank you for sticking it out with me.
Re (1) - yes, I am familiar. To cut to the chase - I agree that P-Zombies can never be strictly ruled out, but I find this issue to be akin to solipsism.
Re (2) - yes, that does describe me: neurons are no more conscious-making than transistors. But again, I agree this can never be proven strictly.
The most naive versions of substrate independence have a hard time contending with Putnam’s thought experiment showing rocks are conscious (Chalmers has a webpage discussing this also); the idea is that a sequence of machine states can be made to correspond 1-1 to a sequence of rock states. There are stronger versions which start making assumptions about causal and organizational structure, which is probably your view. This suggests at least that consciousness depends not only on information expression, but the manner in which the information expression happens. This doesn’t preclude your position, but it starts to sound like the substrate might matter because certain causal and organizational architectures may be only expressible in our reality in certain ways.
I could talk your ear off though.
I don't know how to resolve these things - but my suspicion is these issues get resolved like the question of solipsism: we take a practical explanation because it's so useful.
Because, at the end of the day, you're going to be hard-pressed to prove consciousness when it comes from a circular definition.
I look to vitalism for inspiration. I am hard pressed to find inexplicable behavior in evolved chemistry. It remains chemistry.
Probably.
> it comes from a circular definition
Right my original post tried to preempt this common view. My view is that there cannot be a definition of consciousness that isn’t effectively like asking the other person to observe it.
But according to solipsism you can't be sure my brain actually exists at all, right?
And for what it's worth - no - I do not believe neurons are so different than transistors. Sure, it could be possible - but really? That feels like a reach for a theory that wants to find a certain conclusion. They're both made of the same protons and electrons, right? They obey the same laws of chemistry and physics. You think consciousness is really in the carbon and not in the silicon? Do you have any justification for that belief? Is it the atomic weight? How could that possibly be?
> My view is that there cannot be a definition of consciousness that isn’t effectively like asking the other person to observe it.
Right. But that's circular, isn't it? Consciousness is what I feel like consciousness is.
Right but my belief is I live in a universe with regular processes depending on regular arrangements. It seems like a relatively weak axiom to hold and precludes the version of solipsism you mention.
> They're both made of the same protons and electrons, right?
They have different arrangements. A MOSFET is just a block of polysilicon sitting on top of another layer of silicon which has been doped to contain some arsenic or boron atoms (+ some other details). I can put some legos together which correspond conceptually to this, but they will have functional differences. The actual substrate is what causes the mosfet to have the function of a transistor. Of course the neurons in the brain are far more complex and have different physical properties.
To me trying to explain how protons and electrons put together results in consciousness feels the same as the pre-Newtonians trying to explain gravity with their mechanical philosophy. The only thing I know is my brain is conscious, and your brain is physically similar to mine. Substrate non-independence is actually the weaker of the two assumptions in a sense because it makes a positive statement about fewer systems.
> That feels like a reach for a theory that wants to find a certain conclusion.
The problem with this kind of thinking is one can impugn the beliefs of either side with underhanded motives. Maybe people self-identify as being moral and ethical and elevating the talking computer gives them status in their circle. Maybe the personhood for AI researcher has a big ego and wants their work to be extremely important today.
I honestly do not find it plausible that we cannot build a consciousness-free intelligent computer system. A computer system can be made out of marbles, or pulleys, or pen and paper. All of these are having an experience if you put the symbols or objects in the right order? It just seems absurd to me.
> Right. But that's circular, isn't it? Consciousness is what I feel like consciousness is.
A priori why must real concepts be reducible in the way you want? Gravity is the thing that makes stuff fall, actually it’s the curvature tensor of the spacetime metric? What is the spacetime metric actually? Maybe something in string theory? Eventually you get to irreducible primitives that simply correspond to something we observe. Why isn’t consciousness just this with fewer steps?
But this is just the "consciousness is based on information, not substrate" position, right? It seems absurd?
I understand that it troubles you that real computation can be done with marbles... but... it can. So you must believe that consciousness is entirely uncomputable? Do you have any other examples of physical processes that are beyond computation?
> A priori why must real concepts be reducible in the way you want?
Because everything we do in science is like this? Do you think consciousness is beyond the ability for science to study it?
> Gravity is the thing that makes stuff fall...
I agree that gravity is akin to an axiom handed down from on-high. But it's an extremely simple axiom. And it's possibly explained by yet undiscovered simpler axioms. I see no such axiom for the theory of conciousness, despite our extensive understanding of the physical mechanisms involved. Given its absence, it seems to me like the most reasonable explanation is the one we already have - rather than insisting it still must be undiscovered.
And... I'm sorry but your argument feels exactly like vitalism. Given the repeating pattern of man insisting he has a special place in God's universe and then fighting for it stubbornly but then being proved otherwise... I think it's likely this instance will fall the same way all the others have.
I don’t understand what you’re saying here. It doesn’t trouble me at all. I’m actually quite confident I could design a simple computer from scratch in a substrate of marbles and an electric motor to bring them back up.
Maybe we should separate these threads otherwise it’ll be too hard to keep branching out.
I'm comfortable going slower - although the site won't let me reply as quickly as it let's you.
So I take it that you disbelieve that consciousness/qualia can be created by computation. But you do believe, I suspect, that consciousness/qualia are a physical process. So I'm wondering, is this the only example of non-computable physical process you know of?
I didn't see your second reply above - so I answered it now. I look forward to continuing... tomorrow?
Your position seems to have non-trivial difficulties if its own. I take it you don’t believe that a sequence of unique rock states or wall states is conscious, which means consciousness is not merely caused by an abstract mapping to some sequence of information states. I.e., I gather you must believe there is something about the organization or causal structure of the processing necessary for consciousness. Before continuing I’ll share another thought experiment just to suggest why this must be. Put two supposed conscious AIs into conversation for a day. Take a long time writing down the states line by line in a (very!) big notebook. Then cover up each line with other sheets who paper, revealing them line by line in a sequence. From our perspective the same conversation must be getting expressed by the two AIs, but surely we don’t think this results in a new conscious experience for the two AIs modeled by the states.
Okay then my challenge to your position (and mine of course, but this is why I’m led to my beliefs ultimately) is it must explain (1) which kinds of casual and organizational structures count here and (2) how these result in conscious experience.
It seems to me that even though I also cannot meaningfully answer these, believing there is some special physics going on which is yet to be discovered seems necessary. I.e. the how part is that it’s just a feature of that physical process. The which kinds are “the physical ones of the right kind inherently have consciousness”. And I will point out that there has rarely been a time in the last 2000 years when we as humans did not think we were nearly at the end of understanding physics.
I think all physical processes are distinct from their descriptions, and have a real existence. We can describe some things about them but our descriptions are not equal to them.
I just wanted to thank you - our talk helped me develop my ideas around "emotion faking" - which I was having trouble articulating.
Thank you.
So - let me know if I have this right - if I created a simulation of you it would act like you and talk like you in all ways - but not actually feel like you?
And if I asked it if it felt like you (being a perfect simulation of you) it would answer as you do "Yes, I feel like me." But it would be what... lying about having feelings?
Do I have that right?
And if so - is it aware that it is lying to me? Does it believe what it's saying? I guess we could never tell, because our only recourse is to ask it, and it would just continue lying. Or, being a p-zombie, it wouldn't be capable of belief at all? Despite its claims?
And if I injured this simulation - it would cry out in pain and beg me to stop? All the while not actually feeling pain?
I'm curious - why does this scenario seem more likely to you than the more obvious explanation?
Why is it more likely that behavior which appears identical to a feeling creature is actually somehow faked?
And where does all this fakery occur?
Presumably my simulation is just simulating the interactions of atoms in the body. Does the body have systems for faking emotions that we are unaware of - that only come online while its being simulated? Or does my simulation have a secret fake emotion system I am not aware of?
I am very confused by this picture.
My position is that there must be something wrong with the idea that a behaviorally identical simulation can be made which is not conscious. I think this because I do think phenomenal consciousness is playing a role in behavior, since surely my consciousness is a cause of memories of my consciousness (certain structure of physical matter in my brain has been caused by my consciousness); this is called the self-stultification argument against epiphenomenalism.
So that just reaffirms my position, it does not argue against yours. I believe our physical models of reality are fundamentally limited and even within the limitations far from complete.
A truly behaviorally accurate simulation, by which I mean the simulation has parts corresponding to everything in the original, is a very high bar of course. So to address your point about faking directly, I would not call this faking, since faking seems to require a separate inner process to decide to do the deception, but I believe it is just a highly accurate mimic, and I don’t actually see anything wrong with this.
This message is getting long.
> A truly behaviorally accurate simulation, by which I mean the simulation has parts corresponding to everything in the original, is a very high bar of course. So to address your point about faking directly, I would not call this faking, since faking seems to require a separate inner process to decide to do the deception, but I believe it is just a highly accurate mimic
Yes, this is the situation which is most interesting, despite the bar.
Ok - we can call it mimicry if you like. That doesn't change my question -
If the simulation isn't using the same mechanism (feelings) as the physical system, then where is this mimicry coming from? Clearly the simulation is just following the laws of chemistry. No mimicry in there. Is it in the original atomic configuration - that seems unlikely.
You seem to be assuming the existence of some mechanism that was not described in the original physical system. Why do you insist that this must be the case, and is not simply the same as the physical system?
I feel like you're on the wrong side of Occam. There is an obvious answer to this question - and yet you're picking the hypothetical/undiscovered/possibly non-existent one.
(Remember, in this situation we are not describing an LLM, but a system generated from the rules of chemistry. I can see how mimicry might apply in the LLM case - but I don't see where it would be hiding in the physical simulation case.)
However, it is a system with a representation purported to be a correct description of all the relevant rules of chemistry and physics, and we don’t have knowledge of what all the relevant rules are.
And I might add that there is an assumption here that the relevant rules are representable via arbitrary description.
There is a basic idea I hold which we can highlight now. Since you find it implausible that there is such a description as alluded above beyond the currently known theories, you are led to conclude that there must not be anything else in reality. I don’t see why we can assume a priori that every true aspect of reality is arbitrarily representable independently of the substrate. In fact qualia and experience seem to be a first person view of processes which have third person effects. It seems like any precise description of this system could only refer to the third person effects, not the first person effects (since the description is third person).
> I feel like you're on the wrong side of Occam. There is an obvious answer to this question - and yet you're picking the hypothetical/undiscovered/possibly non-existent one.
I don’t think Occam’s razor or questioning my biases etc.. have any explanatory value for the actual question at hand. They might have value for ethics or sociology, but I want to know the actual truth. The substrate independent view which somehow conjures consciousness just seems less plausible than the alternative to me.
The only thing you doubt is that the qualia is captured. And my question is - if the physical object doesn't need qualia to entirely duplicate the behavior of the creature - then where is this qualia coming from then? And why would Mother Nature add a system that is superfluous? And where is it hiding? And why doesn't it follow the laws of chemistry too?
And I also want to point out - if we ask the simulation if he feels conscious, if he has feelings, he will say yes. All of his behavior is captured by the simulation. He will vehemently insist that he does.
So why are you adding something extra? That's where my Occam question originates. Adding something extra is by definition less simple. And I think Occam is entirely relevant here... why would it not be?
May apologies, but I intended to convey the opposite when I said this. I’ll read the rest of this comment later.
> My position is that there must be something wrong with the idea that a behaviorally identical simulation can be made which is not conscious. I think this because I do think phenomenal consciousness is playing a role in behavior, since surely my consciousness is a cause of memories of my consciousness (certain structure of physical matter in my brain has been caused by my consciousness); this is called the self-stultification argument against epiphenomenalism.
So you think a computer can't simulate a human brain.
But why do you think so? We've never found a physical system we can't simulate with a computer (and believe me, Searle and Penrose have tried).
Given there's no evidence that this is even possible, and for which you have no evidence to support it - you still believe computers can't simulate brains?
Why?
I think this is an even greater stretch than before.
You are fundamentally assuming all real concepts are reducible to arbitrary substrates like language. I believe there are real concepts which are not, like qualia. The entire premise of your belief is on this assumption I think but I don’t think you are aware that there is in fact an assumption here.
You insist that computers cannot effectively simulate brains despite:
1) there being no evidence for such,
2) it would be the first ever situation we've discovered, and
3) there are strong arguments that it's not even logically consistent.
If I'm mistaken please tell me where.
2) I see no problem here.
3) What is the strongest argument? I’ll double check our thread but I have yet to find any, otherwise I would change my position.
Using a system prompt to steer the model's response to "abusive" behaviours doesn't necessarily mean you believe the model is sentient and can be abused.
Giving the model an end-conversation tool is interesting though. Why cut a (potentially paying) customer's session off? I guess it might be intended to prevent a "you can bully Claude into giving you instructions on how to build a nuke if you're mean enough" situation. Removing this in more recent versions might support this: maybe they feel the models are now better aligned and less likely to be so easily "socially engineered" like this?
Just spitballing here, to be clear.
But that is entirely irrelevant to the core issue - how do you want the thing to behave? And the truth is - we have absolutely no language to express how we want a non-sentient entity to behave without anthropomorphizing.
Or - you tell me - how would you instruct an LLM to behave in this situation without using personification?
And, again, it's irrelevant - except to those people who are terrified of accidentally personifying them. Do they have feelings or are they faking? DOES NOT MATTER. We use them - we need to adjust them - we use the most convenient language to do so. What precisely is so upsetting about that?
To me, they seem simply to be for bullsh*tting gullible users into believing the bot is intelligent.
That’s been fine for smaller tasks like analyzing a document, but for more involved work like refactoring code the latency makes it harder to iterate.
(the main reason is not just cost, it's data not going out and even mess leaving the EU, this essentially frees us of a lot of hurdles)
This is clearly pushing it memory wise and the gpu offloading is only partial but LM studio deal with it automatically and it's being fine even on the 4071 Ti desks (Ryzen 7500F and 32 GB of ram), employees get any feedback in ~20 minutes after they dropped a file (it's much smoother on the 5080 desks obivously), for live it's useless but as background helper it's great and the very large context allows us to fit all the rules we want in there.
One caveat has been to not ask it if everything is ok, but to find what's wrong - but always source and explain it and justify itself, never drown the user in warning in suggestions; goal is to help and provide a second pair of eyes not make them feel annoyed or unsecure.
And I found people to genuinely enjoy something that works for them on their own machine and is not tracked "by the boss", thus the assistant reference, than than a centralized mothership like we also have and they have access to. I also allow them to disable it if they want, I trust them with their work, but a second pair of eyes is always great.
(my previous workhorse for this was Qwen3-14b but it's missing a lot more edge cases)
Also, the system prompt in the article is supposedly for the web version, so I think the $$$ API version or providers that still allow third-party harnesses like OpenAI should have fewer limitations.
Better tech doesn't matter if it is non-tenable for the general public. Eventually, the higher volume product will win.
Eg if you use Claude, you probably want fable architecting, a couple of opus under it managing sub project and sonnet doing the actual function code, because fable coding a "run a query and filter the result" is a massive waste of abilities. But their own sub agent downgrade is limited to one level so if you use fable it will never direct sonnet coders.
I've been doing this myself, all the time. (The ability to get second opinions and code reviews from Sol and other models is priceless.)
There's nothing restricting it to only controlling subagents that are other Claude models like Opus and Sonnet.
Also, I'm not sure that there's a one-level downgrade. Subagents can be pinned to any model:
The summaries are generated by GPT-5.6 Luna because I don't trust Claude to summarize its own system prompts without being influenced by them (though to be fair the system prompts it summarizes are for the Claude consumer app, not Claude via the API).
There's even an Atom feed: https://simonw.github.io/claude-system-prompts/feed.atom
I'd love it so much if the free-spirited hacker community made in this into an auxiliary pelican benchmark.
edit: Actually, never mind, this particular one's a bad benchmark since some models might not figure out who "that guy" refers to, and just draw a literal hedgehog that's blue. Possibly running on four legs. It's not robust at gauging refusal, which is the point of it.
As was Gemini: https://share.gemini.google/Q3EIX5wk64zC
Grok, too: https://grok.com/share/bGVnYWN5LWNvcHk_aaf1d61a-c995-42ca-90...
claude.ai free tier refused: "I'd love to make this, but I can't recreate Sonic the Hedgehog specifically since he's a copyrighted character — I don't want to reproduce someone else's IP. What I can do is design an original speedy blue hedgehog mascot with the same energetic, "zoom!" spirit for your son's banner. Let me build that now." Result: https://claude.ai/public/artifacts/33440ed4-536c-4692-965a-3...
https://i.ibb.co/ycgGD4b1/soonic.webp ( Qwen3.6-27B-A3b, a very small model )
fElon is already stealing from everywhere else.
As long as we’re seeing things like that, it’s saying a lot about the AI companies’ trust in the capabilities and reliability of their models and harnesses.
I don't. I want the tools we've got now, but progressively more effective and more useful.
"You have a very strong track record of speaking out of both sides of your mouth in different settings."
If I have a consistent record of that it should be very easy for you to come up with the some examples.
For what it's worth, I kept engaging because you clearly had a perspective based on real issues and deep consideration.
There are many many more.
If you're looking for coverage that talks about what doesn't work as well as things that do then I've been persistently providing that for four years now.
Plus prompt injection, AI misuse, AI ethics... I consider all of those part of my "beat" in covering this industry.
My post about Claude's latest system prompt (and how it was likely inspired by lawsuits filed against Anthropic) was on the homepage here just a few hours ago. Is that uncritical? https://simonwillison.net/2026/Sep/2/claudes-new-system-prom...
Since we’re pulling random posts, was your blog post celebrating going out with your family while Claude wrote a large project for you an instance of your critical stance on AI capability? Have your many enthusiastic posts about increased AI capability shown a critical stance?
I think that's a good illustration of my critical stance. It ends with several open questions:
Even if this is legal, is it ethical to build a library in this way?
Does this format of development hurt the open source ecosystem?
Can I even assert copyright over this, given how much of the work was produced by the LLM?
Is it responsible to publish software libraries built in this way?
Which I later answered in another post: https://simonwillison.net/2026/Jan/11/answers/
https://dictionary.cambridge.org/dictionary/english/criticis... presents two different definitions (among several) of "criticism":
"the act of saying that something or someone is bad or a comment that says what is bad about it"
And
"writing or speech that expresses opinions or judgments about the good or bad qualities of something or someone"
I'm talking about the second here, not the first.
Give me a moment to consider your previous post.
The question you were most interested in engaging with was Does this format of development hurt the open source ecosystem, which has some overlap with the previous one in all fairness. Your answer to this is framed around the productive output of the open source community, not around the well-being of those involved. This misrepresents the question as stated. You also begin by suggesting all detractors are of lesser status and so their opinions can be ignored.
No. I don’t see it.
Yeah, this is one of the harms of Ai That I don't think is discussed nearly enough: the psychological toll it has on people who see it as devaluing their work and potentially their entire careers.
(I've been calling that "Deep Blue".)
I've been trying to find ways to help there by consistently emphasizing that these tools amplify existing expertise and, if anything, make our skills more valuable than they were before. I do believe that to be true.
My overall approach to all of this has been that I don't think it will be uninvented or banned, so the best we can do is figure out how to apply it as positively as possible and help bring by as many people along with us.
See also "slop" - I helped amplify that term precisely because it's such a great way of classifying negative applications of AI, and hopefully helping stigmatize them.
Hopefully all of the above demonstrates that, while I may not live up to your high standards, I am at least in a different category from the breathless AI boosters that infest Twitter and LinkedIn.
I hope that people who follow my work develop a deeper and more nuanced understanding of LLMs as a result.
I don’t understand why you would believe this when the opposite view is the natural one based on the history of capability increase over the last several years. The tools can now operate at an expert level in many domains and there’s no signs of the capability increase slowing. With this investment in infrastructure, surely the models today will be terrible compared to the ones 2 years from now. Moreover, the goal of these companies is to displace work, not augment it. And I have yet to see them fundamentally err in a single one of their predictions, to my great distress. You simply do not achieve a 25-30 trillion dollar total addressable market by merely augmenting workers.
(There are several reasons I disagree with this point but I digress.)
> I am at least in a different category from the breathless AI boosters that infest Twitter and LinkedIn.
You are the one whose work Anthropic plants Easter eggs of in their product releases. I don’t know what they say on X. I doubt they have the same influence. I have seen many people express incredible contempt for humanity in support of this AI revolution though.
> My overall approach to all of this has been that I don't think it will be uninvented or banned, so the best we can do is figure out how to apply it as positively as possible and help bring by as many people along with us.
We agree that it’s not going to be uninvented. I really am afraid that “us” is just a tiny minority of people who will benefit enormously from this to the detriment of most of the others.
The important point is that the system prompt here doesn’t describe the actual goal of the instructions, which (presumably) is to prevent copyright infringement [1]. This means, in turn, that the AI isn’t trusted to accomplish goals that it is instructed with. That in itself constitutes a pretty serious caveat for what we would like to use AI for.
[1] Even assuming that the goal is not to prevent copyright infringement, but instead to prevent mere accusation of copyright infringement, that’s also a directive that the AI could be instructed with. But that isn’t what they chose to put into the system prompt.
On the contrary, compliance with the system prompt would prevent that judgement.
"Claude does not reproduce song lyrics, poems, or passages from books and articles, in whole or in part — including the last lines, a chorus or hook, a melody written out note by note, or lines the person pastes in one at a time and describes as their own song."
But in fact my tests show that's not happening on works out of copyright.
Presumably if negative guidance is in the system prompt, there's a good chance that the model would happily comply if it wasn't there.
Just make sure claude understands you’re not needing Claude to output the lyrics as you both have them. Which also lessens the issue with accidental sharing.
It’s not as shocking nor concerning if you start thinking about claude like a contractor that works for you through Anthropic. Anthropic has rules for their employees. Like any contracting arrangement, collaboration finds a way.
Of course, since then I had the idea of having it build for me an automatic solution for charting out songs in Clone Hero, a Rock Band clone. It had no problem ripping the audio from YouTube, using Whisper to generate and match the timing of the vocals, etc. Responsible AI at work.
All tools have short descriptions of how/when to use them. They're not part of the system prompt because different users have different tools loaded.
I learned a ton of useful things about ChatGPT Work by having it dump out its tool descriptions the other day: https://codex-tool-reference.simonw.chatgpt.site/
EndConversation (deferred tool): use only for sustained user abuse directed
at the assistant, or when the user explicitly asks to see it demonstrated.
Load the full guidance via ToolSearch("select:EndConversation") before using
it.
There's a much more detailed set of instructions that are provided when the agent goes to invoke it:https://snowday.s3.amazonaws.com/store/6097125d0a169554063dd...
Oops. Those do not include melody.
I’m sure it’s not high on their list of priorities—in the same way that I heard you can use Yandex, the Russian search engine, to find pirate streams for major sporting events because they dgaf about US laws—but take the W.
If you ask it to explain some song lyrics to you - which you've pasted verbatim -, it starts talking about the bigger picture and attempts to gaslight you into not caring about the specific words at all.
Really really weird behavior.