Where do learning signals come from when there is no ground truth in post-training?
New paper shows how to convert inference-time compute into high quality supervision for RL training.
Up to 30% rel. improvement on a realistic non-verifiable tasks (HealthBench), with the models own self-synthesised rubrics!