You could go deeper and compare samples of audio that could be uploaded separately (eg, siren sounds), check out MFCC processing https://en.wikipedia.org/wiki/Mel-frequency_cepstrum#Applica... to do Shazam-style audio comparison.
I wonder too if you process it on a per-frame basis or you can take series of frames too (eg analyze the last 5 seconds of frames) to detect things like a "hand wave".