The grader-focused trajectories are striking. Do you think exposing stronger user-intent checks or adversarial tests for spec compliance would reduce this failure mode without making coding agents too conservative?
In a well-constructed training environment, models would not be thinking about the concept of a grader nor recognize that this is an eval rather than production usage. I think about these more as desiderata for the environment itself, rather than as explicit reward-signals to shape the model.
But of course researchers should also simultaneously cut/patch shortcuts/exploits that we find in reward-signals (models should never be able to get full reward with an unsatisfactory solution).