Train your own R1 reasoning model
unsloth.ai
unsloth.ai
Also, the different version of the same og Colab didn't make a 135M model to learn the XML tags, so do you think 8 billion should be the minimum for use this?
Yep - bigger models should help! Definitely give Llama a try!
Do you think that your training setup could help train these models to be better at agentic work?
So with GRPO and reinforcement learning, the OSS model creators now have one more tool to make OSS models much better, since we now don't need vast amounts of labeled CoT data, but rather just questions and answers, and we let RL / GRPO figure out the CoT itself after using some reward function.
So I guess it definitely can help in agentic workloads!