Your argument is just as applicable on human code reviewers. Obviously having others review the code will catch issues you would never have thought of. This includes agents as well.
The tests many of us use for how capable a model or harness is is usually based around whether they can spot logical errors readily visible to humans.