Having worked at a company that went through the FDA process for a genetic test product, where the link between gene in question and phenotype were fairly well characterized, what we want to do for our analytical studies were to basically push the tests out to their failure modes. That meant seeing what happened with common contaminants, what happens with incorrect sample collection technique, what happens with incorrect storage, etc. For our clinical tests, we had to prove not the lab rate, but the actual customer use rate.
For example, 23andme ships you a test package, you sample it yourself, send it back and then they process it. While quantifying that last rate is simple (as you say), it's crucial to understand how well the product will work in practice, with an actual customer on the other end.
For something like 23andme, that would likely mean getting a whole bunch of untrained test subjects, having them preform the sample collection as instructed by the instructions shipped with the sample collector, having them shipped over by standard means to the lab, and then processed. The test subjects would likely then have to have their DNA fully sequenced with some gold standard test, and then the results compared. 23andMe would be given leeway, in the sense that inconclusives don't "really" count as a wrong result.