More likely is that you're helping them write the results section of their research paper, with a sentence like this, "In a blind trial of 5,000 internet users, over 93% of people were unable to tell the difference between generated audio and real audio at a statistically significant level (P<=0.035)."
Note: if we assume a binomial distribution, you need to get 7 or 8 correct to reach the magical P<=.05 barrier. If you assume that there are a fixed 4 generated clips and 4 real clips, then you need to get them all right (https://en.wikipedia.org/wiki/Fisher%27s_exact_test). I think it's fair to use the binomial distribution because the website does not tell you the number of real and generated clips before you take the test.