It depends on the signature.
On the "probably too simple, but..." side, record 8 seconds of audio. Normalize it. Bucket the net volume audio levels across quarter-second intervals. Fuzzy match a bit. If the levels have some change in them, and the changes match, you're in. Virtually no useful information in that signature.
By contrast, if you get "clever", and want to go with that basic approach, but you want more bits to fuzzy match, so you break apart the frequency bands into 64 bands and sample on 50 millisecond intervals, and you're getting perilously close to revealing speech in the signature with enough processing. (You lose a lot, but you end up with more to work with than we privacy sensitive folks would like.)
The point isn't that either of these are what they are doing. I don't expect either of those would "work" on their own; I kept them simple to describe and simple to understand for the purposes of this post. The point is the info leak on the signature can vary quite widely depending on how they do it, and without knowing the algorithm you can't speculate very effectively. It could be anything from "fairly safe" to "you might as well just send up the audio".