The set of 5 hits discovered in the discovery dataset were at the usual 10^-8 threshold. They then meta-analytically pooled those results with the older PGC results and got a set of hits at 10^-5. Finally, they took those two datasets and pooled it with the held out replication dataset, yielding the final set of 17 hits at 10^-8. The final results are at the significance level you want, and almost all of the signs for the top SNPs are the same between datasets/cohorts (presumably why they reported broken-out sub-analyses rather than skipping straight to the final results, to demonstrate consistency). Those 17 are the ones used in the rest of the paper.
> Many GWASs report much less stringent p-value and many don't run sub-replications.
This isn't true. Most GWASes use the standard genome-wide significance level of 10-8. If they do not, it's because of well-motivated other considerations such as being replications of previous hits. (If you are testing replication of 5 earlier hits, rather than 500,000 SNPs, your p-value threshold ought to be looser.)
> I wonder if this study is subject to that requirement.
The supplementary gives the top 10k hits, which is the most critical part of the data, which you can use for polygenic scores, gene sets, heritability & genetic correlations via LD score regression etc. It's only top 10k SNPs because I believe 23andMe imposes that as a requirement on people using its data - something about possible reidentification if too many SNPs' values are released. (There are, of course, other cynical business-related reasons for why they might impose such a requirement.) I've seen that done in a few other GWAS studies like educational attainment, and they said it was because of 23andMe.