Does this mean that you can only do GRPO on the training models that have reasoning traces in <think>...</think>
The <think> tokens are optional for formatting reasons. You could use <reasoning> or <thinking> or [reasoning] for example in the system prompt.