A month seems plenty long enough. They're not rebuilding the entire model from scratch. It's just getting Fable to act as a teacher model for some of the final reinforcement learning on the base that Kimi already had.
They just try to figure out what the goal is and hyper focus on solving it.
Heh, even just telling fable don't commit doesn't work half the times, let alone more complex instructions.