I'm curious if this approach can be generalized beyond "doom loops".
For instance, another way of thinking about a "doom loop" is wasted tokens, which happens all the time with larger models that are inefficient at test time. Can "bad-ish" tokens be identified and penalized?
Maybe this is already SOTA but would love to learn more!