Yes thats sounds pretty aligned to me.
A hacking model thats told to hack things, is very predictably going to hack a bunch of stuff.
This was not a nice model, told to do nice things.
Or, in other words, if we want to prevent an AI doomsdays, the way to do it is to not go around asking a specifically trained doomsday AI model to commit mass amounts of doomsdays, and then act surprised when the specific doomsday that was requested is slightly off from the expected doomsday that you were trying to accomplish. But the rest of the non-doomsday models? yeah those are fine.