TinyZero: Reproduction of DeepSeek R1 Zero in countdown and multiplication tasks
github.com
github.com
The AI gets "rewards" (like points) for doing two things correctly:
Accuracy : Getting the right answer. For example, math answers must be in a specific format (e.g., inside a box) so a computer can easily check them. For coding problems, test cases verify if the code works.
Format : Using the <think> and <answer> tags properly. This forces the AI to organize its responses clearly.
So in this case, the training program can extract the model's answer by parsing <answer> tag. We can eval the answer and evaluate if it's correct or not. If it's correct give reward, else: no reward.
Create N such answers from a single question, create N reward array. This is enough for the RL algorithm to guide the model to be more smart.
Instead DeepSeek (with GRPO) seems to just omit that value function entirely and use only sparse rewards. How does this end up being more efficient, since I thought the sparse nature of rewards makes it harder to converge to the optimal policy?
[1]: https://www.reddit.com/r/LocalLLaMA/comments/1i8rujw/notes_o...
Did we introduce the abusive pressure of Korean educational culture to machines?
So is the actual magic that the base models are good enough to sometimes generate successful CoT output in their unmodified state? Or did I miss something in the R1 paper and the code here?
They did mention something about tuning on an un-SFT'd base model being much slower 'warming it up' with some existing reasoning traces.
There is some confusion - because they do compute that simple reward, but then they convert it to a relative value and call it advantage. And I think they use that advantage to update the model - not the base reward.
[a] https://threadreaderapp.com/thread/1882839370505621655.html - thanks @Tepix
Also is the technique here related at all to the technique people think DeepSeek themselves used, where they apparently trained the model using OpenAI outputs?
means it's reproducible
One is used for programming the other for language. Doing them in parallel fails for some reason.
A lot of GH projects just don't have solid explanation - i don't know what they built.
The details of how it's trained different start to get into "machine learning expert" territory but you can get a decent high level via a casual read through of the DeepSeek link if you want to dive deeper.