The main issue here is mixing up a statistically significant effect and that effect being significant. With a very high powered experiment you can detect very small effects but that doesn’t mean they have a significant effect on the outcome.
I really think the simple hearing test is best but should incorporate common tools of measurement science such as repeated tests and multiple, varied listeners. Audio “quality” is very subjective and can come down to preferences driven by culture, age, experiences, etc. Some sort of consensus rating over all reviewers should be done.
Edit: also the rule of 3s! Never measure twice, always 1 or 3 times, is particularly relevant to this article :)