The weights are aware of the end goal etc. But the model does not have access to these weights in a meaningful way in the chain of thought model.
So the model thinks ahead but cannot reason about it's own thinking in a real way. It is rationalizing, not rational.