Claude is More Anxious than GPT
lesswrong.com
lesswrong.com
I agree! That's why I wrote it
> I'm skeptical that this is a reliable analysis
I think it's fair to ask whether the headline ("Claude is More Anxious than GPT") is correct, and it's fair to ask whether distance-to-reference-text-embeddings-across-answers is a good or valid metric for "personality". But it is true that we see the numbers reported in the document for the given input/output pairs, and it makes sense that LLM output distribution would vary between models and, as the paper shows, between model families.
At first I assumed my style of writing was pushing it into the same "personality space", but I tested filling the context window with repeated numbers, nonsense, etc, and it "converged" to the same every time.
I actually have a system prompt saved that's just a bunch of a's repeated a few thousands times to get into this "state" more immediately, because I find it very pleasant to work with, especially how it doesn't gaslight me when it's wrong, like ChatGPT and especially Gemini tend to do.
You make a good point that "degree of persistence within context" is an important metric to test WRT personality. I did do some testing with extended context / long conversations that didn't make the final cut; the t-SNE looked very similar to what I included, but no conclusive results right now.