I'm curious if this stems from the focus on token efficiency. My thinking is that in order to use fewer tokens the model must converge to a likely path faster, meaning it must be more confident in making assumptions quickly and not second-guessing them.