Actually I think this is one of those things that sounds so obvious, and yet isn't right in practice.
I'l roughly segment projects into those that can fail a bit and those that can't.
Now obviously, for things that can't fail even a little you can't just take someone's word. But why can you take the word of two people, or ten? The failures (in programming and code review) are absolutely correlated. If one person fails at something another likely will fail in the same place or when reviewing it. When people die if there's a bug, or you're even losing a lot of money, you should not be trusting best-effort human anything. This is where imperative programming against a test suite breaks down and just can't cope. You need to switch out your internals for something provable in its domain. (eg, the math to land a space shuttle in a fixed-time infrastructure (ie before it lands itself)).
And for everything else, it's a $/$ calculation. How much do you want to spend to have some unknowably smaller amount of risk? And usually the best value for the dollar is having the original engineering team work on the tests, with outside oversight and good metrics. If an integrated team is having problems getting full branch coverage (the only worthwhile coverage metric...) in a given method, they rewrite the method. External teams have to test what's given (or waste a ton of time in communication delays) and that usually ends up with suboptimal tests and suboptimal coverage. What you can't do though is simply ask a developer if their work is properly tested and trust their answer. You need real metrics and to know what they mean for you. (Like benchmarks.)