And I fully believe such devices are going to be used (if they aren't already) to identify persons engaging in "wrongthink" by listening in on the private conversations in peoples' homes.
And I fully believe such devices are going to be used (if they aren't already) to identify persons engaging in "wrongthink" by listening in on the private conversations in peoples' homes.
In other words, I don't think Amazon ever receives the "wrongthink" Alexa hears, unless you say it directly to Alexa.
Without it being open source, there's no guarantee though?
Would it really be the first time we were lied to/surveilled?
When will we stop giving the hyper-growth oriented Silicon Valley startup world the benefit of the doubt?
I know lack of evidence doesn't mean it doesn't exist, but if that came up empty then it'd be hiding pretty good.
Although honestly I'd delay transmission until user interaction and then hide in that noise - it'd be the first thing I do.
Eh, look at the traffic anyway
You can do multiple runs of feeding pre-recorded messages into say multiple speakers and do the trial over many days. Then on a series of other speakers you can do a robust sequence of pre-recorded conversations followed by the same pre-recorded messages at the same time and then do statistical analysis on traffic volume.
I just presume these things are listening to everything and recording everything. I think that should be the general assumption if you bring essentially an "internet microphone machine" into your home.
If not by the company who sold it to you then by 3rd party hackers, clever app developers, or some other group. Every marketer wants to know what their customers are saying in the privacy of their own home.
As a tangent I've long wanted to have fun with this ... start a campaign to start collectively talking about a ridiculous product (say a vacuum cleaner with elephant ears that flap in proportion to the amount of dirt it picks up) in private conversations and see if a company releases it by listening in. "There's significant consumer demand for the dumbo-vac!"
Isn't this equivalent to the halting problem? Even with source code, there is a chance the compiler was compromised. In practice, these devices are closed source, so you would need to verify all the possible code paths.
Moreover, we know that NSA coerced phone companies into exposing metadata. What is the probability NSA has not requested backdoors of Amazon, Google, and the like?
This is just the nature of indirect observation. People in the natural sciences deal with this problem all the time.
https://moniotrlab.ccis.neu.edu/smart-speakers-study-pets20/
I think the data stream from uploading compressed audio for an extended period would be difficult to hide.
Yeah, but there are mistriggers as well - I think you should see them in myactivity.google.com with Assistant filter enabled.
https://www.theguardian.com/technology/2019/jul/26/apple-con...
We have a Nest Hub, and saying Google twenty times a day was a deal breaker so we all use some other variation that works 99% of the time, and looking at history it accidentally triggers itself a few dozen more times during the day.
It has a real value for us for now, but privacy issues are real in my opinion.
The Google Mini that Spotify sent me, however, went straight into a pile.
You needn't use your real name, of course, but for HN to be a community, users need some identity for other users to relate to. Otherwise we may as well have no usernames and no community, and that would be a different kind of forum. https://hn.algolia.com/?sort=byDate&dateRange=all&type=comme...
Voice recognition is 'on' all the time as it needs to recognise 'keyword'. All you need then is simple transcription into text.
Its certainly possible for Amazon/Google/Whoever to send your device a firmware update that turns it into an always-on microphone, but it doesn't do that by default
Not to say it's not a concern, but that's not really how these devices work – at least in the Alexa case. They're just matching for a specific hotword, rather than constantly performing speech-to-text (which is computationally expensive and done remotely). Think of it more like Shazam or the other audio fingerprinting services – you don't have to actually transcribe the text to understand if a particular word has been heard.
And if one of them is doing it, they all are, they all think the same and have the same incentives. This entire play is about the data.
1) There has been various cases of such devices being triggered incorrectly and uploading chunks of recordings
2) It's all implemented in software. It's extremely easy for the vendor to enable more keywords or record for longer times.
3) It's impossible to prove that 2 is not happening already in limited cases
4) There is a proven long history of very effective global surveillance programs targeting every electronic device (phones, cell towers, carrier-grade routers, PCs, servers).
This is the intended behavior.
But - there is a non-zero rate of false positives (when the device detects something that it thinks is the wake word, but is not), in which case, the audio is streamed to the cloud. This audio could (should?) be used for model training to improve the future precision of wake word detection. But, it could also potentially end up being subpoenaed by a law enforcement agency.
Of course you can't completely trust it. But make it hard for them.
Also I pretty much only use it to listen to the radio.
Bottom line is that Amazon and Facebook somehow decided to show exactly the random stuff we fed them. None of the stuff fed to Google, Microsoft, or Apple ever made it out in an obvious enough way for us to notice it.
It's not something to draw a solid conclusion on, I'm sure the experiment had plenty of flaws but the degree of suspicion it raised was way above the noise floor and it was enough for me personally.
If I search for something, it immediately shows up in the facebook feed of the person I live with. We aren’t even friends on facebook.
However we share an IP address via NAT, and I’m sure location data has leaked enough to correlate us.
I have never yet come across an example of this where there aren’t correlating variables other than the always listening mic theory.
We even tried to make sure the terms are "plausible" given all other data the companies may have had on us. Age, social status, etc. We picked things where we're comfortably but not too obviously in the target audience (no "energy drink for student gamer" type thing).
I can’t tell from your description whether this was adequately controlled or not.
I had an Echo that I installed in a spare room in the house and used for a short time exclusively to have these made up conversations next to it and keep talking about a "Whirlpool washing machine" without ever using the Alexa hot-word. The keyword really couldn't leave that room except via the Echo. After a short time to my surprise I started seeing this in my Amazon. I have no doubt that the Echo is (at least occasionally) listening and sending information without any indication that it does.
My friends tested their own stuff in their household with their own keywords in much the same way that I did. Google Home, Apple Homepod/Siri, Facebook, Microsoft Cortana. The only 2 people who saw their keyword pop up again were myself with Amazon and one other with Facebook.
I can't draw the conclusion that Google does not do this, maybe they just do it smarter. But I can certainly say I cannot under any circumstances give Amazon the benefit of the doubt.
So - are you absolutely certain that nobody in your household used the term ‘whirlpool’ in any text based online interaction?
For example - is it possible that you emailed someone while you were coordinating these tests and you or they are Gmail users?
That, and lists (shopping list, costco list, etc), are our two main uses. I hardly ask it the weather, and only use it as a kitchen timer 1/3 the time too.
In 1984, they got caught because they thought no one could hear them. That was their mistake. What I learned from it (well ok many things), but what that taught me is to be careful if you are going to do something you shouldn't. Never assume that someone isn't watching when you're doing something illegal. I'm just posting signs for everyone else in the house, this area isn't secure.