For AI models, that isn't really the case. We don't know how each individual model will act in the real world, and certainly not after having spent millions of hours training against itself or in a simulated environment. And as these models get more and more powerful, this growing uncertainty is something that I believe is worrisome.
We don’t know how brain works, why sleep exist, most of humans cultures and languages are not documented, hundreds if not thousands of psychological and medical things are unknown, etc. Machines and their applications are a few more magnitudes more understood than any human matter, by the sheer fact we created them. Their complexity is ridiculously low in comparison to human, their biology, psychology, etc.
but all of that being said, i think it's also worth considering what a non-human entity with more power than us might be able to do. after all, even the worst examples of humanity are confined to at most "only" causing millions of deaths, but not the eradication of all life on Earth for example. it's not difficult to imagine that a non-human entity, with non-human goals, and with super-human powers, might act in a way that is contrary to our interests, or even to our existence.
AI is just the fruiting body of something that has already happened.
How many people have been in US prisons again? Oh right, about 3%. Clearly a high percentage of these people did something both unexpected that seriously damaged others.
So this is not a real difference between AIs and humans. One might even say, not a single AI has been convicted yet, so it might not be true for them. For humans, serious, bad, violent, illegal, even when totally irrational, ... behaviors are very common indeed.
(and frankly, if there ever is a true Human <> AI conflict, not allowing any kind of errors for AIs seems to me a strong contender for the casus belli)
The computer can do whatever it wants. People can do whatever they want. The question will be what level of security access will they have.
The key difference today is people are really good at making rationalizations for individual decisions. Computers are not.
Sometimes decisions are generally important, when they are important they require trust, and if these models never generate a method of demonstrating success/trust they won't be adopted.
And therein lies the rub. Do people actually want to know the truth? They invent methods to obtain it, but is that their aim or is it to confirm their current understanding of the world? Unlike a human being that can be forced into silence through coercion or manipulation, a computational model, once proven with certainty, is never going to go back in Pandora's box. This contradiction of pursuing truth but suppressing the inconvenience of its conclusions hasn't disappeared for some reason despite greater and greater knowledge of ourselves. Will it ever?
Several attempts at training with real world models have produced results that have been attacked for being "algorithmically biased". This isn't because there was anything experimentally wrong with the dataset or choice of model on their own, but because of the result's contradiction of a certain worldview. While, I agree with you that people come closer to the truth in the long-term, in the short- and medium-term, I expect certain people to pad their own data and quash the different findings out of ideological motivation for some time.
I'd consider The ability to output a rationalization in human understandable text (ie: english in my case) will determine success /failure.