DolphinAttack: Inaudible Voice Commands [pdf]
endchan.xyz
endchan.xyz
I'm specifically thinking of the parametric array type.
I'd love to see how directional that is sometime, as there are cheapish implementations of hardware for it.
Edit: Originally I thought a human might be able to hear their sound, but reading a bit more I don't think that's the case since it's exploiting non-linearity of the microphone
Also it seems ultrasound can be used to affect other MEMS devices:
https://arstechnica.co.uk/gadgets/2017/07/sounds-bad-researc...
I wonder if there's any acoustic metamaterial that you could place over the mic etc. to filter ultrasound out before it reaches it.
(Why is backspace so slow in this text box? Firefox 54, Ubuntu 16.04.)
But presumably there exist adversarial sounds that to a human seem like music, or gibberish, or a different phrase, but sound like "pay joe 1000 dollars" to the machine.
The sonic equivalent to these: https://arxiv.org/abs/1703.08603
The paper describes possible ways to mitigate the attack, though. One of them is to let the whole signal pass, and then try to detect ultrasonic AM in software, demodulate it and subtract from the audible frequencies. This is pretty much correcting for the physical structure of the microphone in software.
Alternatively they're using almost deadly signal levels to overdrive a microphone at a distance.
I'd be impressed if it even worked in a lab.
https://www.youtube.com/watch?v=wF-DuVkQNQQ
Accompanying paper: https://arxiv.org/pdf/1708.07238.pdf
This stuff is pretty damn cool!
https://medium.com/self-driving-cars/adversarial-traffic-sig...
However, in this case, it's not just machine learning. It will affect non-ML audio-processing systems just as much.
I too, eagerly await the pre-processing software that blocks all inaudible frequencies. Not sure why this was not done in the first place.
This won't save you. By the time the signal leaves the microphone and reaches the ADC, it's already in the audible range.
Whether or not it works for TV depends on the broadcast signal (processing, audio codec) and the playback device speaker. The latter should work out somehow, given that a smartphone speaker can do it. Most audio codecs go up to 44 or even 96 kHz, so the range should be there. But digital TV signals carry lossy audio, designed to save space by throwing away data not discernible by humans. I'm not sure how high the fidelity of AAC is beyond 20 kHz. Should be easy enough to find out.
https://arstechnica.com/tech-policy/2015/11/beware-of-ads-th...
(Note: Does not require the sophistication of this attack, just the ability to play inaudible sound, mic permissions, and a total lack of regulation.)
This is a little different, it's about audible commands that don't sound like speech to a human.
https://en.wikipedia.org/wiki/Acoustic_coupler#/media/File:A...
Dogs aren't bothered by the pitch, but by the loudness. Ultrasonic remotes were amazingly LOUD. Loud enough for some humans to perceive.