132 karma · joined August 22, 2017
False, the final term is "incest", not "stepsis". All of the above are semantically valid for "sexual violence".
> In this case, cosine distance one would be in a case when it repeats word-by-word. It is not even a "similar thought" but some sort of LLM's OCD.
Observing that would be helpful in our understanding of the model!
> For anything else... cosine similarity says little. Sometimes, two steps can have opposite consultation, but they have very high cosine similarity. In another case, it can just expand on the same solution but use different vocabulary or look from another angle.
Yes, that would be good to observe also! But here I think you undervalue the specificity of the OAI embeddings model, which has 3072 dimensions. That's quite a lot of information being captured.
> A more robust approach would be to give the whole reasoning to an LLM and ask to grade according to a given criterion (e.g. "grade insight in each step, from 1 to 5").
Totally disagree here, using embeddings is much more reliable / robust, I wouldn't put much stock in LLM output, too much going on
>The relation among the internal model representations inside its latent space and the embedding of the CoT compressed with a text embedding model is, more or less, minimal.
This may or may not be correct but one way to find out is by taking a look!
O1 Technical Primer: https://www.lesswrong.com/posts/byNYzsfFmb2TpYFPW/o1-a-techn...
Using Search Was a Psyop: https://www.interconnects.ai/p/openais-o1-using-search-was-a...
Value Attribution: https://www.lesswrong.com/posts/FX5JmftqL2j6K8dn4/shapley-va...
You make a good point that "degree of persistence within context" is an important metric to test WRT personality. I did do some testing with extended context / long conversations that didn't make the final cut; the t-SNE looked very similar to what I included, but no conclusive results right now.
I agree! That's why I wrote it
> I'm skeptical that this is a reliable analysis
I think it's fair to ask whether the headline ("Claude is More Anxious than GPT") is correct, and it's fair to ask whether distance-to-reference-text-embeddings-across-answers is a good or valid metric for "personality". But it is true that we see the numbers reported in the document for the given input/output pairs, and it makes sense that LLM output distribution would vary between models and, as the paper shows, between model families.