It's common for different models to find holes in another's work. There are various good reasons for that.
FWIW, we use ChatGPT for our primary model and use Claude to do the reviews. This works better than ChatGPT doing it's own review even with a clean session/context.