Alexa and Siri Can Hear Hidden Commands
nytimes.com
nytimes.com
Edit: Yep, omit / reduce tones in the 3000 - 6000 Hz range https://www.reddit.com/r/amazonecho/comments/5oer2u/i_may_ha...
Expecting TV's to play tones outside the range people can hear is ridiculous.
For example you could embed a 6kHz tone in a way that's inaudible to humans due to the other frequencies in the waveform.
https://www.howtogeek.com/338409/hundreds-of-smartphone-apps...
I don't think they're using infrasound for that. I think they use a technology similar to Shazam where it just analyzes the sound to determine what's on.
Similarly, if the video is heard through a phone line, the call might be using a symbolic voice codec, which would also restore the range (in the sense that it’s not even storing sound, just phonemes.)
Rather than journalistic oversight I think this verifies what people have commented many times: that the fact that GA does not have a personalized name makes it to refer to it. SO much so that a very distant third product is included rather than GA.
I couldn’t have added the subtitle because as you say, it would have been too long.
"Alexa, Siri, and Google Home, can hear commands that humans can't"
"We evaluate these attacks under two different threat models. In the black-box model, an attacker uses the speech recognition system as an opaque oracle. We show that the adversary can produce difficult to understand commands that are effective against existing systems in the black-box model.
Under the white-box model, the attacker has full knowledge of the internals of the speech recognition system and uses it to create attack commands that we demonstrate through user testing are not understandable by humans."
And then "digital assistants like Amazon’s Alexa or Apple’s Siri are set to outnumber people by 2021".
Why is Google mentioned in one context but not other?
I would also recommending changing your default word at a minimum. Then again, I might also recommend ditching the device entirely but I happen to have one in my kitchen that I like OK sometimes.
With everyone scrutinizing the web traffic around them trying to prove NSA/Google spying with them we'd definitely have found something by now.
So it's probably just yours that is sending off to the NSA. I'd send it back to them for repairs, see if they don't give you a free home mini in compensation!
They also said they found a lot of recordings in their Google History that really shouldn't have been there.
Here's another article about the history: https://qz.com/526545/googles-been-quietly-recording-your-vo...
https://www.theverge.com/circuitbreaker/2017/10/11/16462572/...
We literally dont use it for anything else.
Might get one of those philips light sets since our living room is weird... Tbh, id rather not use alexa..
Someone could clearly make an offline device that did voice recognition clock and timer, but does anyone?
It is also voice activated, which means it can be done when your hands are full (another common occurrence in cookery).
The last is the most important for me; I live hundreds of mile from both, and the key is that the echo is a super easy device for them to use.
I can send a chatty audio message telling them I'm looking to call, or asking a question, which is asynchronous in nature, and when they are ready they can call or message me.
It is conference calling by default. Which is great for family stuff, they wouldn't meet my girlfriend otherwise.
I got them the echo show (I have the normal echo) and it is fabulous. Video calling with no effort. I use my phone to do video on my side.
For this, the echo is worth every damn penny and then some.
Say something like "My firetruck is red" -then- manually activate siri, she was listening the whole time.
Though it makes sense that it would have a rolling buffer
How do I do that for Google home?
(To be fair, I recall this specific point coming up in an HR thread about Alexa being able to open your door for Amazon deliveries, but I thought it was worth reiterating here.)
so the question is, shouldn't they be able to detect the wavelength of what they are processing to weed out some of the more obvious tricks? with voice recognition could it also not be limited to a voice it is trained to know?
I run 3rd party roms, Linux on all my dev/tv machines, disable Cortana on my gaming laptop, and hope there isn't something listening in all that trusted, untrusted and oss code I'm running.
I told my roommate I'd move out if he ever got an Alexa or Google home device. I do want to run Jarvis, or one of the OSS alternatives. Many of them send your data to Google/Amazon as well if you enable using their Speech-to-Text services, but they also have options for using local OSS decoders as well (and typically enable those by default).
Our phones are so powerful today there is no reason to send your speech to the cloud (someone else's computer). It should just be done locally; and tech should be improved so accuracy is improved locally without needing the larger datasets that Google/Amazon/Apple use.
More devs need to use the OSS assistance instead, and maybe that will push other engineers to no go the easy route and opt to protect their privacy instead.
As for IoT cameras, though, it's the opposite. I assume all of them have already been pwned, so I never buy them in the first place.
iPhones already do this. My wife's iPhone won't respond to me, and vice versa.
Though I don't know if this is enough to mitigate the attack mentioned in TFA.
Not that digital assistants are worth the risk they bring. Until now I can’t get Siri to do anything useful that isn’t very artificially and carefully phrased.
Given all the hype from my friends and the commercials, I expected something outstanding.
Nope, significantly worse than google's assistant.
That was the start of my complete disappointment in apple as I continued to use an iphone and wonder- Why is anyone buying this?
Personal anecdote, I dislike Apple computers, but I like their phones. So I am at least one person who doesn't buy them for brand.
This notion is outdated.
Things that drive me crazy-
>No widgets
>constant reminders to sign-in
>constant reminders to update
>no double tap/settings seem harder to find
>Finger print scanner sucks so bad.
>Little annoyances like the animations to change screen take 0.5 seconds too long.
I'm not sure what I'm supposed to be enjoying on my iphone.
Flick to the left from the homescreen or from the notification screen. They have them.
> constant reminders to sign-in
After an update, sure, but TouchID got rid of most of these.
> constant reminders to update
iPhones actually get updated. Reminding you to install them is good practice.
> no double tap/settings seem harder to find
Please explain this gesture.
> Finger print scanner sucks so bad.
You probably (like me) only scanned part of your fingerprint.
> Little annoyances like the animations to change screen take 0.5 seconds too long.
Dig into the accessibility settings. They let you change a lot more than you think you should be able to.
This is demonstrably false. iPhone's second-gen Touch ID sensor is one of, if not the best in the business. It's fast and ridiculously accurate, and is the reason people were disappointed in Face ID when it was first released.
Having not used Photos on macOS in years beyond ensuring it actually imported my photos, I opened it up and was surprised at the level of analysis it had. It made an album for each city I visited in Mexico. "Puerto Vallarta 2017" and such. Even had a "Furry Friends in Mexico" album that was all the furred beasts I met along the way.
Really well done, and all done locally on my computer.
This is the sort of thing I have no problem voting for with my dollars.
All of those seem like expected features in any OS/phone.
I switched back because of that + Google Assistant's unwillingness to work without Google tracking my location history constantly (assuming that the "off" switch there even actually stops them).
Also, tons of people choose certain companies to buy from because of their status and reputation. That's how most industries work. Some people only buy American made cars, or only Chevy or only Honda. There's nothing wrong with brand loyalty especially when the brand is consistently delivering quality products to it's customers. iPhone isn't an arbitrary status symbol. Apple put years and years of effort into building up the reputation they have.
Also, Siri is a joke. If your priority in a smartphone is to make use of it's virtual assistant, then an iPhone is not for you.
And here is another article I found that mentions the normal 20-20kHz frequency response range: http://blog.shure.com/mic-basics-frequency-response/
Isn't that mostly overlap with the human ear capability? I understand each person is different, etc. But just curious the specifics.
However, there's a lot of play within the space. One difference is that microphones do a very direct recording of the sound waves, but what we hear is actually very distorted compared to the "real" sound by the nature of our ear. One of the big differences is that if there is a very loud 4000Hz sound, we can't hear a soft 4005Hz sound near it very well, but the microphone "hears" it just fine. So for instance, you could put out a loud sound for a user, but embed a very quiet command in frequencies the human couldn't hear, but if the listening model doesn't account for that (and there are reasons it wouldn't necessarily want to, because it wants to hear commands even in the presence of significant background noise), you could get commands in to a system. See https://en.wikipedia.org/wiki/Psychoacoustics for discussion about how our ears fail to pick up the "real audio" signal, and how much we've exploited that in music compression.
Now, that was a very brute force example. It sounds to me like what this article is talking about are called "adversarial examples" (https://blog.acolyer.org/2017/02/28/when-dnns-go-wrong-adver... ). Voice recognition doesn't listen the same way we do, it doesn't necessarily take a holistic view of the signal, but is looking for specific frequency patterns and changes and turning that into phonemes, into words, etc. (There's a lot of ways of doing this and I don't specifically know what Alexa and Siri are doing, so that's a really vague overview.) If you know what they are looking for, you can use filters to very, very selectively remove the patterns from a bit of music or something that Alexa might trigger on, and then insert just the bare minimum skeleton of the sounds that it is really recognizing. A human won't be able to hear the difference (most likely; depends on how badly the original is mangled but even if it is audible it is almost certainly not audible without an A/B test and very good ears), but the probably-neural-nets monitoring for sounds will end up superstimulated and interpret the adversarial example as words.
While the adversarial examples work best with tuning to the target network, widely-shared networks like Alexa or Siri mean that such tuning is practical where attacking some custom-trained model used by one person isn't, and experiments have shown that adversarial examples travel between separately-trained nets and even non-neural-net models to a much, much greater degree than what at least my own intuition would have suggested before hand. (See previous link and look for the discussion of "Practical black-box attacks against deep learning systems using adversarial examples". It is extremely counter-intuitive to me how easy this is.)
So this attack is kinda a "reverse-MP3" that adds those lossy bits back in, but shaped with an attack payload. Or at least it adds enough pieces of the attack payload that the neural net pattern recognition triggers, while the humans say "Doesn't sound like anything to me".
Is that a close-enough explain-like-im-a-freshman?
(As another sort of philosophical sidebar, this either proves, or provides very strong evidence, that whatever it is our brains are doing, it is not what deep learning nets are doing, nor anything else vulnerable to such trivial adversarial examples. I've seen adversarial examples against another technique that do seem to work against humans as well, but it requires such a distortion to the image that "I can't tell if that's a dog or a toaster" actually makes sense; it's not just some sort of attack against human vision or something, it's a fancy morphed thing halfway between the two that would probably confuse anything and anybody.)
What can Siri do that's dangerous?
> Send an email to my mom that says, "I have an emergency and I need $2000. Here's the account number to send it to: 12345. Mom, please don't ask questions. This is urgent. Send the money now."
Would Siri lookup your mom's email and send that?
There are way too many interaction steps required by the device owner to make this specific one a feasible attack.
I think it can do Apple Pay actions too.
Degrees of dangerous. I don't have a homepod but presumably it couldn't do anything with Apple pay or your messages. Having Alexa or Siri control home automation stuff seems like something you might want to think about a little, leaving the lights on all day and burning some energy is a very different thing than re-configuring your HVAC or a security camera.
A secret command to "paste clipboard into new email, send to [address]" is a shiny new attack vector without any apparent straight forward way to plug the security hole.
I used to have my iPhone lock 5min after I pressed the sleep button. Now that TouchID makes it very easy to unlock, I have it locking immediately.
When I let my friend's 4yo use my iPad, I triple tap the home button and press "Guided Access", which can prevent the user from accessing other apps until I disable it. (I do this because I'm worried about what he may accidentally search on the web, not because I'm worried he'll steal my data!)
The audio system has an A/D converter which samples audio at a specific rate -- say 48 KHz. Aliasing occurs when the input to the D/A convert is above 1/2 the sample rate. A 24001 Hz signal is indistinguishable from a 23999 Hz signal. A 25000 Hz signal is indistinguishible from a 23000 Hz signal, etc.
To eliminate these types of problems, there will be an analog lowpass filter before the sampling circuit. There is a gradual rolloff of signal sensitivity. Aliasing still occurs, but the energy of the aliased signals is significantly reduced.
My guess is you take a voice command, even if it constrained to be say 200 Hz to 2KHz, then invert the spectrum and shift it to the 46-48 KHz range. When this high frequency is played back, due to aliasing, the software after the A/D converter sees it as a 0-2KHz signal, though greatly attenuated. To overcome that, the source audio can be tremendously loud. Humans can't hear it, so it remains stealthy.
Based on flipping through the pages of the paper (https://arxiv.org/pdf/1708.09537.pdf), it looks like they're taking advantage of the non-linearity in the response at high frequencies to effectively demodulate a lower-frequency signal that was mixed up to ~22 KHz.
Which, if that's what they're doing, is totally awesome!
Pet peeve, I really wish that this was a link.
Was I just blind? Is the actual paper linked anywhere in the article?
It's probably this paper: https://nicholas.carlini.com/papers/2018_dls_audioadvex.pdf
discussed in January when it went up on Arxiv: https://news.ycombinator.com/item?id=16220376
Why hide it in radio content? Couldn't they just play it out loud when I am not home?
So, security by obscurity?
Obscurity is a perfectly valid layer in a security system. It's just not sufficient as the primary security mechanism.
Are there docs for Siri so that i can learn what it can/can’t do?
I have tried skipping songs, playing a genre, set random play, and similar in iTunes — generally a failure, often initiates an unwanted phone call.
On the phone I can successfully call the intended contact about 50% of the time, possibly because I have ~250 contacts.
I suspect that if i knew the right words to interact with the API I could have a more enjoyable Siri experience.
Alternatively, is there a way to disable it completely — as in long hold on headphones button does not initiate.
https://techpinions.com/wp-content/uploads/2018/04/Screen-Sh...
It's funny hoe Google is somehow perceived as evil while Apple or Amazon not.
If I have to trust someone with mt data (and we all do), I will choose Google over anyone else.
But really I don’t think anyone deeply believes that these companies are good or evil in the personal human sense, rather it’s a question of incentives and interests.
Google makes money by selling me to advertisers. I understand the business value but I’m personally not comfortable with it.
Amazon makes money by selling me other people’s stuff. I’m comfortable with the business, but sometimes I’m concerned that what’s good for Amazon isn’t what’s good for the people who make the stuff I like.
Apple makes money by selling me stuff that they make. This is the business model that I like best, because when they make stuff I don’t like I don’t buy it, and when they make stuff I love I’m happy to give them my money in exchange.
Buying from the maker is the best win-win virtuous cycle, in my opinion.
Google makes money by knowing everything they can figure out about me, and they're not especially forthcoming about what they know (or worse, what they think they know). Weirdly, I actually sort of trust them at some level, so if they offered me an option to pay up for a guarantee they won't track me or sell my information to the highest bidder, I would be more interested in their services.
Apple is unapologetically interested in getting the largest capital investment from me while being sufficiently committed to keeping my stuff private that the FBI periodically tries to use law to force them to provide a backdoor. At this time I feel that my data is safer with them than any other viable provider. Also, please note the gov't does not seem too concerned about Android devices. That tells me what I need to know, even if the constant security holes and utter lack of updates for devices more than a year old weren't obvious enough (I've had a bunch of Android phones, I'm not an Apple fanboy)
You are absolutely welcome to trust whichever corporation makes you most comfortable, no quibbles from me :). It's still a mostly free country.
As far as I can tell, Google doesn't post a comprehensive reference, based on:
google now command reference site:google.com
Most any list that's been published is from 3rd party sites, and usually from 2016.
Google's documentation that I've found tends to be of a form of a random list of various different scenarios you can do, but nothing comprehensive.
And besides, my sense is that new development is on Google Assistant, which (I think) requires web search history to be turned on, which in my opinion is stepping over the line. I'm getting tired enough of Google's invasiveness that I'd like to switch to iOS, but I can't stand the UI, and the hardware is all too expensive for my tastes.
[1] https://www.google.com/search?q=site:assistant.google.com/se...
There is probably more than just an AI that does speech to text and then a second phase interpreter. I suspect there is some AI in the first layer of Siri/OKGoogle/Alexa that uses context clues to narrow down what you're asking, but who knows for sure. It's a big black box.
Eventually it's like the 90s again where you type "Get ye flask" and you get a box saying, "You cannot get ye flask" and you're left playing Peasants Quest asking, "Why in the world can I not 'get ye flask?!'"
"Alexa play <whatever> in the kitchen from spotify"
No other combination of words works. I'm not sure why it needs me to say "kitchen", all my Sonos systems are connected together, but if I say anything else it'll either play on just the one speaker or not work at all. I'm not sure why I must say "from spotify", but apparently I do or it ends up playing some random radio station from some other service.
I find things like this quite the mouthful:
"Alexa play Black Sabbath by Black Sabbath in the kitchen from spotify"
..and with a statement so complicated it often misunderstands and starts doing something random.
I would MUCH prefer a simple voice based API. Attempting to understand conversational speech properly rarely seems to work effectively and often just ends up with users memorising a command just to get it to understand.