ParentFull threaddjhn·Research seems to on balance point towards RLHF&RLVR merely increasing subjective sampling efficiency within the pretraining data.View on HN