> But the question is ... given them what the previous model could score on ExploitGym, was their negligence reasonable or reckless?
There is room for more than one question here. This model and training method was shown to be a different risk the first time it hacked Artifactory. Besides which their failure to detect it (outside a crash) shows pretty appalling ops for a supposed trillion dollar company.