Using ultrasound attack to disarm a smart-home system
theregister.com
theregister.com
I know Siri is much too dumb to facilitate a money transfer…
Why would voice recognition software be interpreting ultrasonic (or near-ultrasonic) signals at all?
First, it doesn't make sense they'd be trained on them. So why would models be interpreting these as speech at all?
And second, it doesn't make sense they'd make it from the microphone to the recognition engine -- surely there's a low pass filter in there to remove all extraneous noise above the vocal range?
I don't get it.
(Edit: could it be some kind of downsampling aliasing artifact that is interpreted as normal vocal frequency, precisely because they skip a low pass filter that would prevent it?)
I also remember seeing a presentation on the first gen Echo which went into its noise cancelling tech, making sure that stuff coming out of the speaker wasn't received by the mic, so the success of the speaker-to-mic attack vector also seems totally bizarre.
1: https://www.wired.com/2016/11/block-ultrasonic-signals-didnt...
Seems to be the inverse - if the wake word _lacks_ these frequencies, then Echos ignore it.
Listening in on a frequency to disable word detection is a whole other thing...
When I say across a room to set a reminder or add something to my shopping list my Homepods will just silently do so, with no indication but a flash of the screen. I have no idea if it's registered what I was saying or not.
When I ask to turn on the lights in a room, it'll do a bing-bong noise at me to indicate that it's registered despite the fact I can see the lights turning on. It's utter nonsense.
Once Siri is active, using the hardware volume buttons control the feedback volume.
https://cse.engin.umich.edu/stories/researchers-take-control...
I don't get it. Why? You can speak to smart assistants much longer than that isn't it?