The transformer neural network itself only computes the next token probabilities; the sampling of the token, and the mechanism by which the token it generates is concatenated to the input to make it generate the token after, are technically not part of the LLM.
So we can either study the LLM strictly from single-token outputs, or from how it combines into larger constructs. The former excludes not just CoT, but a large part of the benchmarks on which they are tested.
Looking at it another way, how would humans solve PARITY? I don’t think anyone would read the string, close their eyes, then give an answer without a single thought. At the very least, I would place my finger on each digit successively, effectively using a tool that acts as a memory storage (storing a cursor through the string). The reason I would do that, is that I would have devised an algorithm that requires going through digits one at a time.