Siri, Privacy, and Trust
daringfireball.net
daringfireball.net
I wouldn't be surprised if those recordings also include classified conversations. Some extremely high levels of classification require you to go into sound isolated rooms and leave your phones outside, but there are still plenty of levels of classified information where that is not a requirement.
Watch history seems like something that would be difficult to analyze in tiny chunks, you'd really like your researchers to see large periods of watch history. So maybe anonymization is, "here's a months worth of videos, and it's linked to an ID, but not a username."
Without looking it up online, for anyone who wants to take a guess, how do people think that Google's Youtube algorithm training works for recommendations? How much of your full watch history do you think a contractor can see at a time? What anonymization efforts do you think Google uses? Or do you think it's mostly automated based on automatic classification?
If the Verge or Motherboard put out an article tomorrow that said that contractors could see your entire watch history, would you be surprised? If an article came out tomorrow and said that to help train email classification, human contractors were reading GMail messages, would that surprise anyone here, or is that something we collectively expect already?
On the one hand, it's hard to really disagree with Gruber (and his quote from Steve Jobs). Being fully respectful of privacy means being as fully, transparently, and granularly opt-in as possible--in plain English.
On the other hand, doing so will almost certainly slow progress on a voice assistant and make it less generally useful relative to other assistants that don't have such protections in place (all other things being equal). Even if some opt-in fully, it probably won't be sufficient to make up for the many who don't.
Is making everything opt-in the right tradeoff? Probably. We do it for program crash information after all and (I guess?) that's ended up being good enough. And, honestly, if voice assistants don't progress quite as quickly as they might in a privacy oblivious world, it's not the end of the world.
One can debate whether one or the other is the better business decision but that's a separate discussion.
That isn't a decision for Apple, but their users. Personal data is private property. In the EU at least, that's not a slogan or an ideal - it's the law.
https://www.nytimes.com/2019/08/16/technology/ai-humans.html
I’d certainly opt-in to help Apple, for example. Hopefully, within a decade it becomes a solved problem then more work will be done on device.
This recent news has made me feel a little gross. A voice assistant has negative value to me if I can’t completely trust it. Last year there were literally tech news pieces about how Siri data stays on your iPhone, which was obviously an unconfirmed rumour that journalists bought into
I’m not throwing my household full of Apple shit in the dumpster and rage quitting, but I’m quite disappointed in you, Apple. Most disappointing is that I now have to keep an eye on you.
Siri records your queries too, but she doesn’t catalog them or provide access to the running list of requests. You can’t listen to your history of Siri interactions in Apple’s app universe.
While Apple logs and stores Siri queries, they’re tied to a random string of numbers for each user instead of an Apple ID or email address. Apple deletes the association between those queries and those numerical codes after six months. Your Amazon and Google histories, on the other hand, stay there until you decide to delete them.
http://themillenniumreport.com/2017/03/not-only-are-alexa-si...
From Wired, “Apple finally reveals how long Siri keeps your data”, in 2013:
Once the voice recording is six months old, Apple "disassociates" your user number from the clip, deleting the number from the voice file. But it keeps these disassociated files for up to 18 more months for testing and product improvement purposes.
"Apple may keep anonymized Siri data for up to two years," Muller says "If a user turns Siri off, both identifiers are deleted immediately along with any associated data."
The updated version of the milleniumreport lists:
> Update: But what about the bigger pool of data, the aggregated voice queries from each system’s user base? According to this Bloomberg story, Amazon, Apple, Google, and Microsoft are using all that variety to hone these systems’ understanding of spoken language even further. Several times a day, Amazon uses the entire stack of Alexa queries to educate its A.I. about dialects and casual speech. Microsoft has mysterious fake apartments(!) set up to record and understand natural speech patterns. Google slices and dices the audio it’s already captured, then remixes it to help train its system. All these methods are meant to make your voice assistant smarter in the coming years.
Nothing in that paragraph would suggest to a normal user that actual human beings listen to their recordings. In fact, it would suggest the opposite: that in order to get around human review, Microsoft would go to the trouble of setting up fake apartment buildings to do tests in. Amazon educates its AI about queries using the entire stack of queries? Well certainly, that's not happening with a human being, it would be too labor intensive.
This is in an article designed for non-technical users, that has to explain much more legitimately obvious things like wake words:
> Listening to what you say before a wake word is essential to the entire concept of wake words. The process borrows a page from the pre-buffer on many cameras’ burst modes, which capture a few frames before you press the shutter button. This just does it with your voice.
Yet when it gets to the "low-paid contractors will occasionally listen to you having sex" part, suddenly everything is very vague.
Ordinary people did not understand this, and based on Gruber's article, a good number of technical people did not understand either. And in light of the quote at the end of this article, why should Gruber need to read Wired to figure out what's happening on his devices? That's not a good user experience.
Yes, maybe you can argue that he should have understood what "for testing and product improvement purposes" entailed. But he didn't. A lot of people don't know how AI works, and they don't know how ML testing works, and they shouldn't have to.
Recall that Siri was an acquisition. If you go back to media then, you’ll see many companies were doing this at the time. There was lots of coverage about humans training. Some companies even got in trouble that it was only* humans, even the supposedly AI parts. (An email transcription service, for example, had a human transcription center in Egypt.) Point being, this was known, and seems to have been forgotten.
Even now, all of them mean people reviewing quality when they say for product improvement purposes.
Anyone who has called any big company knows exactly what the phrase “for quality improvement purposes” means. They tell you they’re going to be recording your call for quality improvement purposes and while they don’t say this part, you absolutely know it means there’s a chance someone is going to listen to it.
Since that’s about the only time people ever hear the phrase not buried in legalese, it could be assumed that’s what it means in other contexts as well.
* EDIT: Linking to a history of Siri since it’s been over a decade now: https://www.bestaiassistant.com/siri/old-siri-history-siri-a...
But they haven't linked them where AI is concerned. There could be a couple of reasons for that -- it could be that phone recordings are primarily human analyzed, and because AI has more automatic learning, people just assumed it was different. Or maybe it used to be obvious and then over time people just assumed AI matured enough that it wasn't necessary any more. But regardless, it seems pretty obvious to me that a lot of people (even smart people that I respect) didn't understand.
So to me, talking about whether they should have understood doesn't change much -- I still think it would behoove a privacy-centric company to be a lot more clear now that we know that people don't understand.
[emphasis mine]
This doesn't seem like "acting surprised."
On another note I never use voice assistants. I just cannot be bothered to relearn how to speak for them.
But they definitely should let people choose whether their data is stored to help improve Siri, and whether it will be managed by humans.
Are you sure about that? A quote in the article reads:
> [Apple] says it keeps recordings for six months before removing identifying information from a copy that it could keep for two years or more.
Except here it is FAANG who just keep breaking the user's trust to get the data. But that is a necessary step on the road to some "El Dorado" system that will have been trained on so much unethically collected data that they are done and it can finally become an ethical system.
Edit: clarification