Divergent mode is much easier for me. I just unfocus my eyes (the same muscle that blurs them).
14 karma · joined February 17, 2024
I was more thinking a model that writes to a piece of scratch paper to gain confidence. But it doesn’t have to actually output the scratch paper that it uses, it’s totally hidden from the user.
You could take this a step further, and have something like a “two-brained” model, where the original model falls back on a secondary model if it’s not confident in its response. This resembles a “fast” and “slow” brain.
I think the scratch paper idea has been explored to some extent, but I’m not sure if people think it’s a dead end.
This begs the question - why don’t people force the model to generate more tokens, until it has very high confidence in its next word prediction?
I can imagine several ways of doing this.