AIs do what they do, and we don't know how they do it - or why.
We can characterize some AI behaviors in advance - but not all behaviors. And the book on "best practices of AI wrangling" is yet to be written. Pharma has been dealing with vaguely similar problems - every experimental drug has a risk profile, side effects are unknown in advance - but they had decades to figure out some of the "best practices". AI labs are going in fast and hard, writing the book as they go.
Clearly, some of the lines in there are going to be written in blood.