If you're checking whether a coin is fair and you toss out significantly more tails than heads because you didn't feel they were proper tosses, of course you're going to reject the null hypothesis that the coin is fair. Even if your criteria for what was a proper toss is objective and reasonable, you're no longer testing whether the coin is fair, you're only testing whether your criteria for counting a toss is fair assuming the coin isn't.
When you construct your experiments to reject a hypothesis - this particular model can not be right if we see this - then you can make real progress towards truth.