Alexa Guard has had this functionality for a while, and I’d expect the folks here at HN to be able to infer a few things from the support link and basic reasoning.
Support link https://support.ring.com/hc/en-us/articles/360028205592-Usin...
So: 1) if an event is detected, you can listen to a 10 second clip or drop in (2-way call) to listen in or look. 2) Echo devices have relatively small amounts of RAM 3) Echo devices aren’t constantly hammering WiFi connections
From this, one should be able to deduce that the wakeword engine detects events and streams clips to servers only in situations that match events and settings to support these features. Why? Because processing, transit, and storage aren’t free, and one can’t store data in RAM that isn’t there or transmit data over WiFi without the physical layer showing signs of it. Furthermore, Amazon hasn’t cracked the code on hyper-efficient GB into KB lossless compression only to squirrel it away only for use in voice assistants.
Take the number of Alexa devices sold and run the numbers for all of those devices sending audio data to AWS all the time. The costs would be astronomical. The same goes for Google (though not with AWS). They’re no doubt incorporating the detectors into their on-device models.
I thought kbg archiver solved it eons ago.
The HomePod for example uses the Apple A8 chip which is a very capable chip used to power the iPhone 6 than did way more than encode/decode audio.
Interesting point, my guess would be that not many would notice since a HomePod/Alexa/Google Home would usually sit somewhere in a corner of a room/under the TV and not be regularly touched since you don't need to touch it to control it most of the time.
I am not even sure it would be that much heat, my x86 laptop can play video for a very long amount of time before getting noticeably hot, granted with a fan, however these ARM CPUs get noticeably less hot than your average Intel chip, even without a fan.
Even for the cheaper devices, the CPU is probably capable enough, (maybe excluding the cheapest Echo Dot/Nest).
The Echo Show devices even have a screen and are actually designed to play videos from all kind of sources, (decode), as well as for videocalling, (encode) and they're £60 right now on Amazon UK.
They could do speech recognition on the device and then ship off the plain text. I don't think they do this, but it is most certainly within their technical ability.
As a practical example, I have a copy of a 458,045 word audiobook on my computer and I just downloaded a copy of the e-book. The audiobook is just over 1 GiB, while the plain text of the e-book compressed with bz2 comes in at 800 KiB.
Even when they actually do ship audio off the device for processing, I'd be surprised if it's done losslessly.
Not with the tiny cheap CPUs that are one these devices they couldn't. Or at least not very effective speech recognition.
If they don’t today, they will one day. They don’t have a choice on the matter because it’s a money maker.
EDIT:
Thanks for the reply exittheone, I see GMail actually stopped this ad scanning practice in 2017[1], likely with google sign in on chrome and so many places it’s very possibly they just don’t that extra info. They still can let third party extensions read your email so I wouldn’t say we’re exactly in a better world...[2]
1. https://variety.com/2017/digital/news/google-gmail-ads-email...
2. https://mashable.com/article/google-reading-your-emails-resp...
This is most definitely false. Google still states that they do not scan your emails for as targeting.
"That Veronica Vaughn is one hot piece of ass. And I know from experience."
"No you don't"
"Well not me personally, but this guy I know. Him and her GOT IT ON."
"No they didn't"
"But you can imagine what it would be like if they did."
Google didn't originally scan emails. Then they did. Then they didn't.
Countering "Google scanned emails" with "well the don't right now" is the only ridiculous behavior in this thread.
(I work on Gmail)
[0]: https://privacyrights.org/resources/privacy-and-civil-libert... (from 2004, the year Gmail launched)
Please don't retort "they don't sell PII", as P2 most definitely is what they use to sell X to 3rd parties in whatever form the ubiquitous NDAs prevent the public from knowing.
https://en.wikipedia.org/wiki/Gmail#Automated_scanning_of_em...
Edit: P2 > PII, am watching motogp & removed unnecessary personal jab.
Why do you think Google would sell pii? It doesn't make business sense.
(I work on Gmail)
In these discussions, "selling PII" is sometimes a short for "providing a service that allows targeting ads based on PII without actually releasing PII itself to the service users", which is only a little less bad.
Also, from a privacy perspective, there is a vast difference between the two. One person who you trust to keep your data safe (but not to use it in ways you find ethical necessarily) is vastly different from them giving it to other parties who you don't know.
What's currently stopping evil actors from exfiltrating data from Google's PII through buying narrowly targeted ads, each time with slightly different targeting, and intersecting these results to build a more detailed picture of people who viewed the ads?
> It's great that they do that, but it's still the bare minimum over here on the Old Continent. We can and should demand more.
Sure, but its been possible to do this since well before the GDPR mandated it. Now maybe you can argue that the threat of regulation is what keeps Google in check here, and ok fine that's an unfalsifiable claim but maybe it's true. But even still, that doesn't actually justify "we should demand more". Maybe more privacy regulation is justified, but "we already have some" isn't actually justification.
> What's currently stopping evil actors from exfiltrating data from Google's PII through buying narrowly targeted ads, each time with slightly different targeting
The snarky answer first. PII has a specific meaning. It means personally identifying information. Your ZIP code isn't PII. Your name is. No matter what ad targeting tricks you do you can't pull my name or address out of what Google sends you. So you don't get PII.
Now the less snarky answer. The actual attack you're describing does this repeated targeting thing, which ties private data to some pseudonymous ID, like a browser fingerprint. At this point they don't have any PII. Then, you get the victim to enter their personal information on your site. Now you can tie the PII to the other information from the shadow profile you've built.
So why isn't this useful? Mostly, cost. To get this to work, you need to have some one or some group click on multiple different ads you control ($ + time cost) and then enter their identifying information on a site you control. Click through isn't assured, and conversion to entering information is very unlikely. When you're, you know, actually selling a product, this is a worthwhile investment.
But this attack is essentially paying to advertise to people with the goal of learning who you are advertising to. As a result, this only really makes sense in the context of targeted attacks or generic blackmail. Targeted attacks don't work because now you need a specific person to enter their PII in your site (and then what?, you've learned that someone is interested in LGBT topics. I'm interested in LGBT topics and straight). And similarly broad blackmail doesn't work.
But I'm interested in how you think an attacker could do something in a cost effective manner.
As for your red herring, I can only guess how mal-scans can be monetized. If anyone can do it, I have no doubt Goog will.
Edit: O-T content removed.
Well yes, they're very clearly monetized: Google supplies a email service which it sells to businesses as part of GSuite. One of the selling points of this email service is spam protection. Is that a bad thing?
> You use PII to create whatever product(s) you sell to 3rds.
Ok, so if your concern is that a company has your PII, that's a concern I guess. But it makes it difficult to use the internet (or, like, shop at stores), since there are all kinds of companies that gobble up your PII but don't tell you or give you control over it (whereas Google does, for example, let you delete the data it has on you and control collection of much of the data it collects).
If you're going to argue that Google is bad for the same reason that your credit card company and Amazon and CVS are bad, then sure they all have your PII, but I'd still argue that Google behaves more ethically than the others when dealing with it.
> whereas Google does, for example, let you delete the data it has on you and control collection of much of the data it collects
That's table stakes under GDPR. It's great that they do that, but it's still the bare minimum over here on the Old Continent. We can and should demand more.
"Google also places advertising on Gmail based on key words that appear in messages transmitted through our system (it’s a good example of ads helping to pay for the free services we all enjoy online) - so if you’re emailing a friend about a trip to Paris, for example, ads might appear on the right hand side of the page for trains to France. Google does this using software similar to the kind that scans emails for viruses, to filter out spam and turn the bits of data received into the characters on the screen. No human being other than the user ever reads the messages sent or received on Gmail – it’s simply a computer matching up key words in peoples’ emails with targeted ads." -- https://static.googleusercontent.com/media/services.google.c... (~2008)
(Disclosure: I work for Google, speaking only for myself)
To address both links: 1) As a user I'm actually happy about the tight integration into calendar and other apps. These necessarily require reading my mail though. I'm already trusting Google with my mail so using it for better integration into other Google products is fine with me.
2) They are not randomly handing out access to your emails to third parties, they just have an API that lets third parties read your emails _after you gave explicit consent via the regular oAuth flow_. Which is also completely reasonable for me.
On the other smoke alarm sound is very easy to detect with classical methods so might be extremely cheap to run it without any ML.
How power hungry is it actually? I'd only have to run it on the big model if a pretty dumb model thinks that the voice is even in the audio stream.
The advantage of an ML model is that you can do multi-class prediction for a sound clip with somewhat arbitrary complexity and the cost of execution is more or less the same even if you add an extra class or two. It's just an extra row in your output prediction vector. By complexity I mean signals that don't have an obviously characteristic spectrum, like glass smashing. A typical CNN backbone is capable of classifying hundreds if not thousands of classes with high accuracy - even the edge architectures. Always-on detection tends to use very compact networks (kB in size) that will run on low-power ARM cores, or even specialist ASICs, but even so 20-30 common audio event types seems very feasible.
For people worrying about sending data back, there's no reason why you'd do this off-device. The exception might be that you feedback to Google that there was a false alarm, so they can use your sound clip as a negative training example. Just a guess there, but Tesla does this extensively for Autopilot - they deploy models in your car and specifically ask it to capture images of rare events (Andrej Karpathy gave an example of tree-occluded stop signs).
You can toggle glass break and fire alarm detection on any of your Google Home devices in the Home app, and this includes the old Google Home mini pucks and hubs, as well as new Nest branded minis and hubs.