The thing I really don't understand is this seems like such a simple thing to test! Search for X in our data set vs. competitor data set: flag for review if we are not within 1km radius. Don't release until the number and scope of flags hasn't been reduced to an acceptable margin of error (don't even have to test the full data set, just need a good enough statistician to help you figure out how bad your overall data is based on your sample tested data).