HNHacker News
TopNewBestAskShowJobs

mp187

14 karma · joined February 17, 2024

submissionscomments
mp187··on Stereogram Tutorial (2020)
Divergent mode is much easier for me. I just unfocus my eyes (the same muscle that blurs them).
mp187··on Kolmogorov-Arnold Networks
Why was this your first thought? Is a limiting factor to transformers the MLP layer? I thought the bottleneck was in the renormalization part.
mp187··on Chain-of-Thought Reasoning Without Prompting
I wasn’t thinking something like beam search, I think this seems kind of unnatural. I can imagine that the human brain is doing something like GPT, but I can’t imagine it’s doing something like a beam search.

I was more thinking a model that writes to a piece of scratch paper to gain confidence. But it doesn’t have to actually output the scratch paper that it uses, it’s totally hidden from the user.

You could take this a step further, and have something like a “two-brained” model, where the original model falls back on a secondary model if it’s not confident in its response. This resembles a “fast” and “slow” brain.

I think the scratch paper idea has been explored to some extent, but I’m not sure if people think it’s a dead end.

mp187··on Chain-of-Thought Reasoning Without Prompting
A common theme in papers like these is that the model chooses word predictions greedily, instead of “thinking” and gaining confidence in its next word prediction.

This begs the question - why don’t people force the model to generate more tokens, until it has very high confidence in its next word prediction?

I can imagine several ways of doing this.