Similar to exams where both the progress to the solution and the final outcome/value of the calculations are part of the grade.
To have the cake and eat it too for chain-of-thought reasoning, one way is to ask for a "final answer" so the final response token logprobs can be evaluated https://chatgpt.com/share/67239d92-b24c-800a-af8c-40da7be1f5...
Another trick is using JSON mode to keep intermediate results and final response separate, so each can be graded accordingly.