Has anyone tried to see what the output looks like if the model is never allowed to be uncertain?
For example, whenever certainty drops below a threshold the sampler backtracks and chooses different tokens. Such that at the end every single token had an above threshold certainty.
I doubt it would entirely eliminate undesirable outputs, but it would be interesting.