The one thing that gives me concern in their is "nanoda [the external proof checker] is [now] tracked daily". Although that would have caught this issue, we also now live in a world in which some model is going to think that hacking the proof-checker distribution is the obvious way to obtain the proof it is after; I expect that attempts on that will be much more common than soundness bugs. However, this is said without knowing what other measures are in place to assure the integrity of the distribution.
This used to be the bane of all machine learning experiments. It might have been lost to time but I once stumbled upon a big list of AI reward-hacks like this. Things like - we tried to develop fast cars, but the AI just made a really tall weighted stick that would fall over onto the finish line. And we tried to teach the AI not to lose in Tetris, so it hit pause whenever it was about to.