Understood that that is how you test in the early stages. But at some point you have to run tests against the actual deployed perceptual hash database.
No reason why you can't just generate random images and add them to the real database, indistinguishable from actual hashes. As long as the test images aren't public it's low risk, and even if they leak they could be removed from the database and new test images could be generated.
Wouldn't that mean that people created a system that can be used with normal pictures? So your code was literally to flag an innocent photo?