Lots of RLHF - reinforcement learning from human feedback.
Also: "distillation" - seeding or running training sessions on the output of frontier models.
Also: "distillation" - seeding or running training sessions on the output of frontier models.
No comments yet.