HNHacker News
TopNewBestAskShowJobs

higuidebot

132 karma · joined August 22, 2017

submissionscomments
higuidebot··on Morphology of a Marvel Movie
What do you think a "better" result would be here? Better by what metric?
higuidebot··on From "hot blonde" to "stepsis": porn titles over time
Definitely plausible but it underrates the changes in actual content. It's not just SEO and titling, it's actual videos that have "stepsis" etc. as themes.
higuidebot··on From "hot blonde" to "stepsis": porn titles over time
I am the author, it's just those 3 terms for the tSNE cluster. Sorry, I can tell from some of the comments here the graphs need to be clearer. "Stepsis" is indicative of incest IMO, the "step" is a fig leaf.
higuidebot··on From "hot blonde" to "stepsis": porn titles over time
It's not quite "random" to look at the most intentionally directed legislation of several years. What's your explanation? Genuinely curious. One way or another the clusters do exist, the trends do exist in both content and titling convention
higuidebot··on From "hot blonde" to "stepsis": porn titles over time
SEO contributes for sure, but I would reverse the statement here - it's more about new content and less about SEO. There's a feedback loop race to the bottom dynamic regardless
higuidebot··on From "hot blonde" to "stepsis": porn titles over time
>For their category "Sexual Violence", they include both violent terms ("woman being raped", "torture porn"), and non-violent terms ("stepsis")

False, the final term is "incest", not "stepsis". All of the above are semantically valid for "sexual violence".

higuidebot··on Watch R1 "think" with animated chains of thought
You're right curiousgal. I filed an issue in response to this comment and will resolve sometime soon: https://github.com/dhealy05/frames_of_mind/issues/1
higuidebot··on Watch R1 "think" with animated chains of thought
Lol @ R1's description of my Github ... thanks Deepseek, very cool!
higuidebot··on Watch R1 "think" with animated chains of thought
No need to do anything in particular! Perhaps interesting to observe
higuidebot··on Watch R1 "think" with animated chains of thought
Ha, if steps have consistent distances you could take the average distance at step X and generate a step of that length in some direction and be ~approximately correct regardless of the actual value
higuidebot··on Watch R1 "think" with animated chains of thought
Very cool!
higuidebot··on Watch R1 "think" with animated chains of thought
Well, you are certainly correct about how cosine sim would apply to the text embeddings, but I disagree about how useful that application is to our understanding of the model.

> In this case, cosine distance one would be in a case when it repeats word-by-word. It is not even a "similar thought" but some sort of LLM's OCD.

Observing that would be helpful in our understanding of the model!

> For anything else... cosine similarity says little. Sometimes, two steps can have opposite consultation, but they have very high cosine similarity. In another case, it can just expand on the same solution but use different vocabulary or look from another angle.

Yes, that would be good to observe also! But here I think you undervalue the specificity of the OAI embeddings model, which has 3072 dimensions. That's quite a lot of information being captured.

> A more robust approach would be to give the whole reasoning to an LLM and ask to grade according to a given criterion (e.g. "grade insight in each step, from 1 to 5").

Totally disagree here, using embeddings is much more reliable / robust, I wouldn't put much stock in LLM output, too much going on

higuidebot··on Watch R1 "think" with animated chains of thought
Potentially it's useful to understand a model "on its own terms" via its observable outputs.

>The relation among the internal model representations inside its latent space and the embedding of the CoT compressed with a text embedding model is, more or less, minimal.

This may or may not be correct but one way to find out is by taking a look!

higuidebot··on Watch R1 "think" with animated chains of thought
Text embeddings are underused WRT model understanding IMO. "Interpretability" focuses on more complex tools but perhaps misses some of the basics - shouldn't we have some sort of visual understanding of model thinking?
higuidebot··on Watch R1 "think" with animated chains of thought
Thank you for your support Mr. Wu
higuidebot··on Watch R1 "think" with animated chains of thought
I too am pleasantly surprised
higuidebot··on Watch R1 "think" with animated chains of thought
2D plot is tSNE, consecutive distance comparison is cosine sim distance normalized across the chain of thought
higuidebot··on Watch R1 "think" with animated chains of thought
2D plot is tSNE, consecutive distance comparison is cosine sim distance normalized across the chain of thought
higuidebot··on Watch R1 "think" with animated chains of thought
Whether or not it's "noise" might depend on if you think Chains of Thought are causally relevant? I liked these pieces if you want to read more about CoT / O1:

O1 Technical Primer: https://www.lesswrong.com/posts/byNYzsfFmb2TpYFPW/o1-a-techn...

Using Search Was a Psyop: https://www.interconnects.ai/p/openais-o1-using-search-was-a...

Value Attribution: https://www.lesswrong.com/posts/FX5JmftqL2j6K8dn4/shapley-va...

higuidebot··on Watch R1 "think" with animated chains of thought
Random walk is definitely possible. Also possible that we're observing some "search" in the embedding space from an initial point. It's hard to tell because the chains are often similar lengths, so I don't think it really terminates early. It might be interesting to find the closest CoT component to the final answer and see how step distance inflects at that point
higuidebot··on Watch R1 "think" with animated chains of thought
Well ... we have to reduce them to a 2D plane to visualize them ...
higuidebot··on Watch R1 "think" with animated chains of thought
Is it generating embeddings or just coordinates? What would be a better way?
higuidebot··on Claude is More Anxious than GPT
Any question in particular WRT system prompts and their effects you'd find interesting? Taking suggestions for followups!
higuidebot··on Claude is More Anxious than GPT
Ok I am legit interested in the behavior of "thousand of a's" Claude lol.

You make a good point that "degree of persistence within context" is an important metric to test WRT personality. I did do some testing with extended context / long conversations that didn't make the final cut; the t-SNE looked very similar to what I included, but no conclusive results right now.

higuidebot··on Claude is More Anxious than GPT
> For such an interesting and perhaps important line of work, there seems to be a surprising lack of psychometric rigor in certain corners of the literature.

I agree! That's why I wrote it

> I'm skeptical that this is a reliable analysis

I think it's fair to ask whether the headline ("Claude is More Anxious than GPT") is correct, and it's fair to ask whether distance-to-reference-text-embeddings-across-answers is a good or valid metric for "personality". But it is true that we see the numbers reported in the document for the given input/output pairs, and it makes sense that LLM output distribution would vary between models and, as the paper shows, between model families.

higuidebot··on Claude is More Anxious than GPT
That's one way to get to UBI alright
higuidebot··on Claude is More Anxious than GPT
Next step is definitely investigating the feedback loops / interplay between personality and "performance". My biggest question is whether intelligence converges into a certain personality "band", i.e. are higher performing models going to be more similar to each other ... We don't have quite enough variety to know yet but there's fertile ground there
higuidebot··on Claude is More Anxious than GPT
He's been two steps ahead since like 1984 ... will read!