The model supports both instruction and RL fine-tuning. It looks like people have started training it already [1]. The RL tuning may be only available to the Cloudflare customers, I couldn't be sure.
If it needs to be fine-tuned, then why not fine-tune smaller and cheaper models that are only as big as the task requires?