Yes, they should have.
But the question is not "given they were doing ExploitGym, why didn't they try harder?", rather it is "given them what the previous model could score on ExploitGym, was their negligence reasonable or reckless?"
Personally, I lean towards them being "reckless", but that question is what lawsuits would rely on, from what I hear.
Remember also that there were still independent bloggers in mid February of this year writing in those blogs to assert that LLMs could not find zero days. Anthropic had already had blogged about agent finding critical zero-days a week earlier, but people keep mis-judging performance.