It's pretty dumb below xhigh reasoning as far as I know.
It seems weird to me that just using a ton of output tokens manages to produce a decent result in the end.
It seems to work well though. Sometimes I fear that it might be more likely to eg. run an incorrect, destructive command, but maybe that concern is not justified.