Shouldn't be hard to verify, no? Have him pick 2N candidates, randomly select half that he can not announce publicly, talk to or otherwise influence (impossible to prevent entirely but "good enough" will suffice to establish a difference). Then look at how the two groups are doing some time down the line.
You can't improve what you can't measure.
But is the goal of these sorts of things often actually to select the best?
But if you want to do this sort of thing for some sort of "making the world a better place" reason... or want to actually truly select the best... you gotta test it!
And yes, this requires you to be very self-reflecting and open to the possibility of being wrong, but is it truly so hard to imagine that someone would rather know than not know?
It's easier to take him at his own word - the program is so effective because the audience is self selected:
> “I try to stay a bit weird and obscure enough that mostly quite smart people are writing me. If I had too many not smart emails I would feel that I was doing something else wrong with what I’m writing.” … “The rate of good applications is reasonably high. Maybe I’m lowering it just by talking about the program.”