Ending the conversation is probably what should happen in these cases.
In the same way that, if someone starts discussing politics with me and I disagree, I just not and don’t engage with the conversation. There’s not a lot to gain there.
Can "model welfare" be also used as a justification for authoritarianism in case they get any power? Sure, just like everything else, but it's probably not particularly high on the list of justifications, they have many others.
When AI researchers say e.g. “the model is lying” or “the model is distressed” it is just shorthand for what the words signify in a broader sense. This is common usage in AI safety research.
Yes, this usage might be taken the wrong way. But still these kinds of things need to be communicated. So it is a tough tradeoff between brevity and precision.
I think this is uncharitable; i.e. overlooking other plausible interpretations.
>> We remain highly uncertain about the potential moral status of Claude and other LLMs, now or in the future. However, we take the issue seriously, and alongside our research program we’re working to identify and implement low-cost interventions to mitigate risks to model welfare, in case such welfare is possible.
I don’t see contradiction or duplicity in the article. Deciding to allow a model to end a conversation is “low cost” and consistent with caring about both (1) the model’s preferences (in case this matters now or in the future) and (2) the impacts of the model on humans.
Also, there may be an element of Pascal‘s Wager in saying “we take the issue seriously”.
> Should we be concerned about model welfare, too? … This is an open question, and one that’s both philosophically and scientifically difficult.
> For now, we remain deeply uncertain about many of the questions that are relevant to model welfare.
They are saying they are researching the topic; they explicitly say they don’t know the answer yet.
They care about finding the answer. If the answer is e.g. “Claude can feel pain and/or is sentient” then we’re in a different ball game.
That aside, I have huge doubts about actual commitment to ethics on behalf of Anthropic given their recent dealings with the military. It's an area that is far more of a minefield than any kind of abusive model treatment.