Learning from Human Preferences
blog.openai.com
blog.openai.com
> Our algorithm’s performance is only as good as the human evaluator’s intuition about what behaviors look correct, so if the human doesn’t have a good grasp of the task they may not offer as much helpful feedback. Relatedly, in some domains our system can result in agents adopting policies that trick the evaluators. For example, a robot which was supposed to grasp items instead positioned its manipulator in between the camera and the object so that it only appeared to be grasping it, as shown below.
What you've pointed out is a wonderful demonstration of our need to improve measurement and testing, which humans are typically horrible at. For a good example, just look at government under any party. Policies are continually argued over when that time would be better spent developing a good test and solid measurements for policies and laws.
Most regulations are closed-loop in that sense. They create the way to measure their effectiveness and be reevaluated accordingly.
Really looking forward to a follow-up where they explore 2.2.4 further. Sampling examples which provide maximal information game seems like it could result in another huge reduction in the amount of human oversight necessary. Could see an adversarial scheme which could learn to sample these examples optimally from the manifold. This kind of thing is powerful in human learning of complex tasks to ask for clarification or feedback in specific places of uncertainty.