I'm not sure why one would even think an academic test could determine success. One of the biggest lies many of us have been told is that if you're just smart enough you'll be successful. Sorry, but that's not how it works.
[0] https://edpolicyinca.org/publications/predicting-college-suc...
[1] https://www.latimes.com/california/story/2019-12-22/grades-v...
One very important piece of insight I gained throughout that time is that after you use some exam to get the top 25% (ish) percent of students, the exam performance really does not matter: it does not correlate with scientific output or originality of work. If anything, I had to unlearn some of the skills that made me good at olympiads because they were severely limiting my creativity. At an olympiad you know there is a solution, while true research problems might be unsolvable and need to be approached differently.
TL;DR: Exams (selective or not, hard or not) are great at giving you the top 20% of students, but they are inherently terrible at telling you who among the top 20% will be a productive scientist or engineer.
In particular the famous Benbow study [1] that established that even among the top percentile of the population there are massive differences: the top quarter percentile (99.75-99.99) was 2-3 times more likely to have authored academic research later in life, and about 1.5x more likely to have attended postgrad education, to have gone for a STEM degree, to have gone for a PhD than the bottom quarter of the top percent (99.00-99.24).
I am asking, because if I have to choose between my empirical observations (anecdotal as they are, given it is n=1 observers), and a single piece of research without followups, I do feel justified to stick to what I have seen myself. But I am open to be convinced otherwise.
To start testing the validity of the claims in this paper a study needs to be performed where the strength of effect has to be considered, when comparing a similar one-percenter group and a larger ten-percenter group. My prediction is that the strength of effect would show only negligible differences, confirming my hunch that the special treatment[1] is what created the new subdivision in this new group.
More reading material on this topic would certainly be interesting, if you have anything in mind.
[1]: The special treatment in this paper being telling a kid "you are only a 7th grader, but you are as smart as a high schooler".