New research shows RL may not help a model learn new basic skills
arxiv.org
arxiv.org
This means if the core data (ex. additions, subtractions, etc) were not there in the pre training stage, RL on complex math problems would not lead to the model developing improvements in the core areas.