LLMs generate their output one word at a time (and don't themselves even know what that word will be, since it's randomly sampled from the output probabilities the model generates).
Chain-of-Thought simply let's the model see it's own output as an input and therefore build upon that. It lets them break a complex problem down into a series of simpler steps which they can see (output becomes input) and build upon.
It's amazing how well these models can do without CoT ("think step-by-step") when they are just ad-libbing word by word, but you can see the limitations of it if you ask for a bunch of sentences starting with a certain type of word, vs ending with that type of word. They struggle with the ending one because there is little internal planning ahead (none, other than to the extent to which the current output word limits, or was proscribed by, the next one).