Alexa and Google Home expose users to vishing and eavesdropping
srlabs.de
srlabs.de
1. https://www.vice.com/en_us/article/43kga3/amazon-is-coaching...
2. https://www.eff.org/deeplinks/2019/08/five-concerns-about-am...
Edit to append link and quote:
Quote: However, he noted, there is a workaround if a resident happens to reject a police request. If the community member doesn’t want to supply a Ring video that seems vital to a local law enforcement investigation, police can contact Amazon, which will then essentially “subpoena” the video.
Link: https://www.govtech.com/security/Amazons-Ring-Video-Camera-A...
I mean, something like this should be viewed in the context of comparable IRL vendors. If I rent a 3rd-party storage unit from U-Haul or similar, a warrant is generally required.
(one exception I found was a case where police, on-site, witnessed a drug deal. They then used the defendant's key to open their unit without a warrant. It was judged lawful, that finding drugs and keycard on the defendant was sufficient probably cause. That makes sense, given that if police witness you in front of your house, or car, etc selling drugs, that would be sufficient as well to search.) [0]
https://www.govinfo.gov/content/pkg/USCOURTS-ilnd-1_14-cr-00...
I assume it's legal for Amazon to give them these videos, since the images they're asking for are things that the cops or anyone else could have seen happening outside your house if they would have been driving by (or that your neighbors could have told the cops). There's no legal expectation of privacy, as there would be inside your house.
They can do this because it hasn’t yet been determined unlawful.
We are in a dire need of cyber ethics framework that enshrines user privacy.
Newsflash: computing device with the capability for user interaction can request information that you might not want to give it.
In other words, how is this situation different from any software running on any other type of computing device?
Back in 1966, the makers of the Eliza AI chatbot program were shocked to learn people inherently trusted the program and told it things they didn't want other people to hear. So I propose vishing capitalizes on this phenomena.
i.e. - it's not a new technique, but a new instance of the problem, and that makes it worthwhile (especially for something widely used in private environments) to explore and expose.
It'd be nice if we could reach some kind of device/phone capability plateau and reduce consumption of new equipment. And ideally settle on a small set of software to use on those, which could be hardened and made reliable over time.
Until then, ...
As far as many users are concerned, they're talking to Alexa. The third party app is Alexa, too.
And because of the opaque single-dimensional nature of voice interfaces, even a savvy user doesn't know who's really receiving their intent -- there are enough glitches where you think you're sending to the active skill, but you're back in Alexa's lobby again, so the inverse case the researchers are playing with is a good vector.
I think they could solve some of this because Amazon/Google are gatekeepers -- they get user input no matter where it goes -- they could easily automate detecting anomalous user input and flag for review (that would of course miss the first victims, but it's better than nothing).
I think the "Who's listening?" part is a little harder to solve. Maybe by forcing the third party app to always announce itself as itself? But that does add some friction to the "experience" they want to provide...however, a little friction is better if it means protecting your users.
Developers (on Alexa, at least) can optionally do this now with SSML, but making it a requirement would be an audio cue to users that the “actor” has changed — without adding any delay to the interaction.
Visual UIs generally offer a host of cues to indicate what program is running, and take special efforts to make security-sensitive interactions and dialogs hard to fake. Using these techniques in a voice UI is tricky. There's no good way to tell where the last output came from, or where the next input is going. How can a user be certain that a request for privileged information is coming from a trusted source? In this example, Google clearly tried to create a signature sound (the "Bye earcon") that lets the user know when an app has exited, but an app was able to fake it. The attack leverages the user's trust that was built up by Google.
I think this article provides a useful example that highlights the particular difficulties securing a voice UI system from phishing attacks.
How is this not a massive red flag?
One thing Amazon or Google could do here for voice apps, though, that Apple and Google (Android) can't for standard phone apps, is audit voice responses for anomalies or user input that matches a suspicous pattern and flag apps that trigger it.
They can do this because every utterance a user sends a voice assistant passes through Amazon or Google systems. If an app has access to user PII, they could add some automation to flag suspicious user responses or anomalous activity that differs from x days previous and pass it up the chain for review.
One thing I do like about Alexa development is that if you, as a developer, are privacy-minded (and don't need nor want user data for anything), you can protect your users by configuring your apps not to collect any of your users' info . As a developer, you don't even get IP addresses as everything goes User > Amazon > Developer > Amazon > User.
You always get session and Amazon-assigned user id, but they're typically pretty anonymous unless the user says "I am Jane Doe" -- which, I guess we should be honest, probably does happen more than it should, and this is what the OP researchers are exploiting.
If like me you were wondering what it meant:
"Vishing is the telephone equivalent of phishing. It is described as the act of using the telephone in an attempt to scam the user into surrendering private information that will be used for identity theft."
But thanks for the explanation.
In practice we know that trigger words for all these devices occasionally misfire (There was also an issue with one type of google home device a while ago which was shipped with a faulty physical button that caused it to be turned on at intervals as if the user had pressed the physical button to start speaking).
And would we know if they had been recording unnecessarily?
And you can look at network traffic (e.g. from wifi router stats) to be pretty confident they're not constantly live-streaming audio up to the cloud.
Of course most people will not actually do this monitoring themselves, but there are enough of these devices out there that if a significant number started recording constantly somebody would notice pretty quickly. And that would be terrible PR for the company involved, so I think google and amazon and apple have a pretty strong incentive not to do this.
The PR angle isn't that reassuring to me either, they've already absorbed some pretty bad PR hits on these devices and they're still going strong.
a) viewing what they store via their log tools (though this isn't guaranteed to show everything, ie if they are recording everything they couldhide)
b) monitoring outbound network connections
It's got a small couple-second buffer (enough to store "Amazon" or "Computer" or "Alexa" or "Echo") where it takes what it hears and compares it with its internal model for a match.
If there's no match, the buffer is overwritten with the next bit of noise. Once the device gets a wake word match, it transmits the statement that follows to home base to transcribe and handle.
My simple solution is to not use these kind of devices.
Momentary switches only activate when pushed and as son as their not pushed deäctivate.
Was it a momentary switch or a normal one (push to activate, push to deäctivate)?
Edit: at least the physical slide switch on the Home Mini is a hardware cut-off; I assume the same is true of the Home Max.
Edit: I hate people making claims with zero evidence, so here's some evidence for you: I just took apart my Google Home Mini. The mics are digital PDM mics, connected to a shared line (in stereo config), that goes to what is almost certainly an AND gate (tiny IC, can't quite find the part number, pinout matches a SN74LV1T00), with the other input connected directly to the mute switch (via some resistors), and the output to the SoC (via a resistor divider, probably because the SoC input is likely 1.8V logic). When the mute switch is engaged, the output of the AND gate, which is normally a TDM train (average half of 3.3V), goes to 0V. This is the output that goes to the SoC. So when the mute switch is engaged, the audio input from the mics is electrically cut off from the SoC.
https://security.stackexchange.com/questions/154343/can-a-sp...