> because nobody is doing this kind of processing client-side.
This is a half truth at best. Audio recognition can be based on one of two techniques. The first is audio watermarking which is completely client side (albeit not applicable for this FB feature) or fingerprinting. They'll likely use the latter which records a segment of audio, does the rough audio equivalent of a hash on it and then send it the server for an optimised pattern match. It's somewhat unlikely they'll send raw audio to their servers although that's an assumption.