It’s not shocking at all. I find my Echo to be incredibly useful and I’ve decided that it’s worth the trade-off. Just like I did with the microphone I have on my wrist and the one in my pocket.
It’s not shocking at all. I find my Echo to be incredibly useful and I’ve decided that it’s worth the trade-off. Just like I did with the microphone I have on my wrist and the one in my pocket.
The "smart speaker", on the other hand, is 100% useless if the microphone is disabled. It also has microphones designed to hear as far away as possible given its design.
Finally, the phone in your pocket and watch on your wrist likely handle the audio data very different than that speaker manufacturer. You might find it useful now, but don't discount now you'll ever find it intrusive.
>My expectation
This is the key phrase here. Just as you have your expectations, Alexa/Home owners have their expectations. And if Alexa owners find out that Amazon/Google is using their inputs in ways they don't want it to, they'll drop it or find an alternative.
Different people have different expectations. I was late into the whole smartphone thing, and refused to buy one for years because of the privacy leak it represents (heck, I refused to buy a dumb phone for years for the same reason). Back in those days, I would have been as unsympathetic for your rationale of owning a smartphone as you are of those owning smart speakers.
Do you not see that there are different levels of comfort, and not everyone has the same level as you do? Furthermore, do you not see that on that continuum, even you have made compromises?
Without being aware of what is going on, people can’t make these choices for themselves.
The Snowden leaks revealed that this had been an attack vector available and used for years.
The method they were using has since been patched, but one would be foolish to think that it has been impossible for interested parties to have found another way to do the same.
I'm also not disputing there are baseband controls a nefarious entity could abuse but I'm saying that people who choose to trust Amazon with always listening devices vs a smart phone are not comparing apples to apples.
There are of course more subtle ways to do this, but not for real time monitoring. You can transcode the audio with an encoder that does advanced mathematical quantization and batch upload it. If all you care about is human voice, then you could upload an hours worth of audio in a couple hundred kbytes, if even that much.
Days worth of text could be compressed and sent off in literally dozens of kilobytes. And that's not even accounting for side-channels (a hypothetical targeted attack could conceivably look for "target phrases" and send a single bit of information using a side channel to some server to alert that the phrase has been said).
Back when I used to help the FBI monitor cell phones, they were very slow processors and the user could tell when it was sending anything due to the heat. Best we could do was play with MWI on/off to get cell locations and use the mainframe (cell switch) to do audio monitoring, but that was very limited at the time.
The coprocessors are really just ASICs designed to run neural networks efficiently. The real cool shit is in the compression and simplification of the networks that allow constant recognition from a small power-sipping chip.
They're also doing a bunch of work with federated machine learning processes. So all of the training data can be kept on the device itself, but they can (i'm grossly oversimplifying this) basically "merge" a full neural net from somewhere else with data learned on-device.
It's that last part that I'm extremely interested in, as it means we can have all the benefits of machine learning systems, but also keep all of our privacy in that all the processing happens locally, our training data doesn't need to leave the device at all, and the algorithms can still adapt and improve for each person individually. It apparently also has side effects that the neural network on your device can adapt to your usage individually much faster and more accurately than a "master" neural network ever could.
Are there now co-processors running all the time that could do full speech recognition (not just keyword recognition, which is way easier, but full speech recognition) all the time without such significant battery drain that it would be noticed? I ask not because I consider this some sort of technical impossibility, but because if that is the case, I'd like to know so I can update my understanding of the situation.
You are correct that the actual processors that are "always" listening are fairly limited (They describe it as a DSP, so that should give you an idea of its capabilities), but when layered the way they are in Pixel phones, it makes it possible for the DSP to detect "something" interesting is happening, then wake up and pass the information to an "AI processor" (they call it the "Pixel visual core" but i find that a misnomer because it's a pretty capable TPU on it's own) which would then do the transcribing, and go back to sleep.
I'm making some assumptions that I haven't researched myself, but I'm assuming there are patterns, frequencies, and other simply detected "signals" present in most speech that make it easy to detect "someone is speaking", and from there can wake up a more powerful (but still surprisingly efficient) processor to do more work.
And Google also has some simplified and compressed neural networks dedicated to transcribing speech which are designed to run on lower powered devices[2]. I don't know if they are actually using the PVC in the Pixel phones to do audio transcription or if they're just using the general CPU for that, but I have a feeling that a TPU would be able to scream through buffered audio very quickly and would go back to sleep very quickly leading to minimal power draw.
In fact I'd be willing to bet good money that an "always transcribing" system could be done today with current generation phones in a way that wouldn't impact battery life enough to be obviously noticeable or cause the user to start looking for what was draining it in most cases. And considering how often I'm personally actively speaking in a day, I bet it would have virtually no impact on my personal device if implemented this way.
(and man did it feel wrong to type that last paragraph out!)
[1] https://ai.google/research/pubs/pub46522
[2] https://ai.googleblog.com/2019/03/an-all-neural-on-device-sp...
Both do local processing for the keywords until it hears one then sends the following data (and I believe a few seconds prior) to the servers for processing.
My phone works the exact same way with the Google assistant. I don't know if other "assistants" (like Bixby) work the same way.
I enjoyed the conversation with Alex Stamos (former CSO at Yahoo/FB), where pointed out that had George Orwell been told that people would willingly pay for a tracking device and keep it on them at all times and powered on at all times, he would not have believed it.
I should add, any time I say anything even remotely negative about cell phones, I get beat up here at HN pretty good. This is a very good psychological indication that people have cognitive dissonance and denial around this topic.