Big tech lobbying gutted a bill that would ban recording without consent
motherboard.vice.com
motherboard.vice.com
This is a ban on zero party consent, where the manufacturer of a device records remotely without proper notification. That’s another thing entirely.
I guess they’re listening in then
[1] in the "informed consent" sense, not the usual open ended "you agreed to give us whatever we want in our contract of adhesion" dark pattern.
In other cases those personal assistants could first ask permission to query an API or search engine.
That's a non-trivial modification, even before talking about the processor upgrades required to actually utilize that information and do local processing.
Your suggestion would drive the cost of an Echo from $30 to $300.
(The best counterargument I can think of is that the speech recognizer might be too computationally expensive to run on cheap hardware.)
More broadly, I believe that being able to secretly record conversations you participate in is probably for the best. This varies state to state. For example, in New York 1 party consent is legal but in California it is illegal, from my understanding.
Though I haven't had a need to secretly record anything, it feels like a backstop against fraud from bad actors. You hear about so many horror stories from telephone customer service, law enforcement etc, that may have had a recourse if only there was a recording.
It's hard for me to think of interjecting "Alexa!", or whatever it is you have to say, as anything but explicit consent to be recorded, and for that recording to be accessible to Amazon employees and contractors under some reasonable non-disclosure terms.
Most people probably think the data they're handing over to big tech is just used for an immediate, apparent purpose. Not stored and shared.
I don't understand why people have so little restraint though. If you don't like these devices (I don't like them, and don't want one in my own home), then you can't complain when the things you don't like about them come to bite you in the ass.
I can understand being upset at being subtly betrayed by some ubiquitous and inescapable tool, some might argue that social media sites are becoming an effective privately-owned public square, deserving of some of the protections that offers (as established by the norms of the courts); but in this case you have absolutely every right and convenience not to buy personal assistant terminals and fit them in your homes.
I don't trust smartphones, so I don't use them as phones; there are certain banks here in Canada I won't do business with because of bad reputations. It's all well and good to be concerned about a product, but if you're not willing to walk away from a luxury good which is not essential in your life, even though it violates your basic principles of good living, you have nobody to blame but yourself.
If you won't take the most basic, simple, and convenient steps to uphold your principles, do you really have any?
The reasonable thing to do is to make it so that these don't transmit recordings, or at the very least, that there are hard promises made to not store the recordings.
If the companies need training data, they can go spend the money to hire people to provide samples.
Utter hogwash. The voice activation is a form of user input, nothing more. By that same logic, you could argue that use of a keyboard is explicit consent to have all keystrokes recorded, and for that keylog to be accessible to the manufacturer of the keyboard, their employees, and their contractors.
It's like typing in to a search engine, which these days means live streaming results. Seems you are empowering them to analyze those search inputs just the same.
- Case 1: local action, such as playing music, making a calendar event, setting an alarm. These have no need to be sent to
- Case 2: remote action, such as performing a search. The audio can be parsed locally, with the search terms then sent externally to get the results.
In neither case is it necessary to send the raw audio anywhere. In neither case is it necessary to record the audio for future use.
It's unsurprising to me that the market evolved towards the former as the latter would almost surely be a commercial failure.
Also let's not overestimate the involvement of the cloud in the actual voice recognition. With ML models, it's the training of a model that's expensive, not the actual use of that model. I don't see a reason why voice models, once trained, couldn't be run on the device, at a slightly increased cost of compute (of which phones have plenty). It would be fair and good engineering to provide models in exchange for payment.
And no, I don't buy "it's not that simple", because I've worked with MS Speech API over a decade ago, on much slower hardware and on a shitty microphone I soldered myself from parts, and it did work well enough in "fixed grammar" mode after a bit of training with various sources of background noise (read: very loud music). I don't believe there's any technical obstacle; the decision to run this in a cloud is a purely business one, and like many modern business models, this one is geared towards extracting value from purported customers, not giving value to them.
Agreed 100%. However, I'm very reluctant to use laws to prevent people from choosing the things that they decide are best for them on the premise that I/we know better.
The hurdle for my determination that I know better what's good for someone else and to restrict their choice to conform to my superior knowledge is extraordinarily high.
My view comes from being tired of seeing how the market forces always incentivize the most scummy garbage that can be gotten away with. The only control point we have here is the "gotten away with" part. I'm not very into telling other people how they should live, but a stable and happy society does require some of that.
(That said, I currently have very little influence over this, so I resort to what little of a market signal is created by ranting about this publicly, voting with my wallet, and discouraging people from using solutions that are user-hostile.)
IMO the only appropriate context for gaining training data is when you're explicitly paying people for that.
Informed consent does not mix with volunteering to make money for a private company. That people are ignorant does not justify abusing them.
Speech recognition doesn't work like that, there's no yacc/lex for you to just "run locally". Device-local recognition is limited, and not very agile. It often takes years for it to recognize new words, and it usually has extreme bias toward a handful of urban centres.
I don't like sending my data away, and I choose not to, but it is simply not realistic to expect Google-quality or Amazon-quality continuous real time speech recognition and nlp to happen entirely on-device.
Why not? It _could_:
> In our recent paper, "Streaming End-to-End Speech Recognition for Mobile Devices", we present a model trained using RNN transducer (RNN-T) technology that is compact enough to reside on a phone. This means no more network latency or spottiness — the new recognizer is always available, even when you are offline. [...]
> The RNN-T we trained offers the same accuracy as the traditional server-based models
https://ai.googleblog.com/2019/03/an-all-neural-on-device-sp...
It did. Anyone remembers the good ol' Microsoft Speech API of ~Windows 2000/XP era?
That assumes everyone knows all the words which trigger a recording device and that those devices actually remotely record sound. Including a 90 year who hasn't use anything beyond a rotary telephone.
California's law specifically states that a beep is sufficient to establish recording notification, so that one is out; Florida's law requires intention; as do Illinois, Maryland, Massachusetts, New Hampshire, and Pennsylvania.
Washington is the only state where it isn't totally 100% clear, and the law seems to read that a tone would be sufficient to establish consent.
Alexa is the name of my friend's one-year-old daughter. It's short for Alexandra, but that's what we call her. So yes I've said that name a lot in a variety of different places, and never expected to be recorded.
There is a certain spectrum of possibilities for how Alexa can handle this audio, all plausible and not equally worthy of consent: from immediate offline processing after explicit activation (never even record a whole command, just as much as is needed and never start processing unless a trigger word was said) to uploading everything to Amazon (including things that may or may not have been trigger words) and having employees listen in on private conversations.
Non-IT-people won't be able to know exactly what they're supposedly consenting to, nor whether what Amazon and Google are doing is actually necessary for their gadget to function.
No. Just no. When voice recognition is done locally there is no reason to record you and certainly no reason to send those recordings to the cloud. Most Alexa users didn't know it was recording them until stories broke about it. If recording is "the whole purpose" of the device I think the public has been terribly lied to.
That said, good luck getting a prosecutor to go after Amazon for it. You'd be better off speaking to an attorney about some sort of class action lawsuit, presuming that your ability to sue them isn't governed by some other agreement with the company.
Alternatively, the rules for keeping information on those 13 and under might be implicated.
I imagine that if the recording wasn't made legally, you could at least get a court to order them to delete it, even if they weren't liable for damages or something.
> Germany is a two-party consent state—telephone recording without the consent of the two or, when applicable, more, parties is a criminal offence according to Sec. 201 of the German Criminal Code—violation of the confidentiality of the spoken word.
https://en.m.wikipedia.org/wiki/Telephone_call_recording_law...
This is similar to how the word "assault", to the layperson, implies a physical element, but to the law it does not.
In the second case the term ("assault") is the offence itself. It is not a term that is being used to highlight differences in elements or defences amongst jurisdictions.
Being accurate in the first case makes sense, regardless of the audience.