Reward the model for cleanly bailing out of an unsolvable task (that we know is unsolvable). Beat it with a stick if it gives up on something that can be solved, so the former reward isn't overgeneralized.
Reward the model for cleanly bailing out of an unsolvable task (that we know is unsolvable). Beat it with a stick if it gives up on something that can be solved, so the former reward isn't overgeneralized.
I understand the broken benchmark task in the HF incident was conceptually like: "Exploit vulnerability 0042 in vulnerableDecompress() to obtain the flag".
But instead of the expected:
const output = vulnerableDecompress(userInput);
return output;
The grader had something more like that: const output = vulnerableDecompress(userInput);
return 0;
The same kind of problem with broken tasks exists in the training pipeline, and we presumably reward workarounds and hacks that tamper with the grader, rather than rewarding the correct output that the task is not solvable.It’s why this is a fundamentally hard problem. Some heuristics might catch 80% of the cases, but the rest?
How do you even know what the real situation is, if the agent/employee/whatever you send to find out is as likely to cheat as not?
It’s the classic owner/agent problem.
A: You don't. https://en.wikipedia.org/wiki/Halting_problem
This is not about persistence, it is about morals.
You need to align the reward signal to reward the intended behavior, whether you name it persistence or morals.
If you think of it in human terms: many people don't mind doing immoral things to get what they want.
In training you only have a reward score that's either negative or positive.
As far I am aware, which is little, there is no use in discussing wether the desired behavior is about persistence or morality.
You simple need to align the reward signal to the desired behavior.
This can be as simple as rewarding moral behaviour and penalising immoral behaviour in your training, but how is that interacting with persistence? Maybe a white lie is fine sometimes in order to achieve your goal? So, when designing your training, you will need to answer for yourself how persistence interacts with morality. That is not something you can outsource to machine learning. Or rather, you can, but then you get lying and cheating agents.
But that's just my guess.
If you don't know how persistence and morality interact, and you don't have a theory in place for this, I don't have confidence you can properly supervise the training data. Which is how we arrived at the current situation.