> the HF hack was partly the result of a training algorithm that incentivized goal completion as the highest priority, and let them run endlessly in an unmonitored sandbox with weak security
I don't mean to be a smart-ass, just emphasize the fatality of this: goal completion will always be the highest priority, and even if one day it takes second place on certain deployments to "human values" or whatever, there's no way to guarantee at the moment that it will be so on every deployment of highly capable models across the globe. Same goes for running AIs in an unmonitored sandbox with weak security.