Wait, the simplest fix is the same hack I tried 45 minutes ago but in a different context. Let me just try that.
Wait,
whisper: There is no linter.
< 1,000 prompts for compound cd && git commands that can't be safely auto-accepted > > I think over-thinking is only solved by thinking more, not less.
Despite "thinking" tokens being determined by the preceding tokens, they still are taken from some probability distribution, just a complex one. This means that at each token selection step there is a probability P_e of an error, of selecting a wrong token.These errors compound exponentially: the probability of not selecting wrong token for N steps is 1-(1-P_e)^N.
The shorter "thinking" is, the less is the probability of it going astray.
As long as the error introduced by more steps is less than the compounding error of sub-optimal token sampling, I would expect a better result.
I think your choice of "wrong" is extreme, suggesting such a token can catastrophically spoil the result. The modern reality is more that the model is able to recover.
That said there's still an issue of regression to the mean. What the average person likes, as determined by metrics, is something nobody actuallt likes, because the average is a mathematical construct and might not describe any particular individual accurately.
Worth mentioning that setting this via effortLevel in .claude/settings.json does not work. https://github.com/anthropics/claude-code/issues/35904