The reason it doesn't happen is because corporations don't like giving up control of data that they can keep to themselves and store forever.
Many of these apps could run clientside and access the cloud solely for information, but that doesn't fit in with the Google vision of organizing the world's information and making it useful for selling products and services to advertisers.
Data retention and collection practices that prevent widespread spying must be codified into law for this to ever work. Companies, Google being the shining example, (but Apple and Facebook not far behind) will simply not enforce these boundaries upon themselves, because the data is just too valuable in the context of future algorithmic analysis.
Also, it's not a "slight chance". All these major innovations HAVE been used for evil, from mobile phones to wide-area networking to the web. It's not just fearmongering.
This is precisely my point, these technologies have been used for evil, but would we be better off NOT having them in the first place? Everything has been used for evil, including pencils. If we loose sight of perspective and only play on the fears, then it's very much fearmongering.
I agree with you on your solution of decentralized collection being better. But I would argue that data analysis involving many data sources, including yours, is what makes a lot of the services being built USEFUL. Google Now being a good example of that. I would also argue, that targeted advertisement is much more useful to me, than non targeted advertisement.
The solution is to stop centralized data collection.
Centralized data collection gives me good auto-complete in search. It means when I start sending an email to John Smith and Jane Doe it can ask me if I might have meant Jane Roe instead. It means when I lose my phone I don't lose any of my data.Similarly for "Mark as Spam", Priority Inbox, Recommended Videos on Youtube, Voice Recognition on Android, etc.
Note 1: Yes, you could also do a pretty good job by having a model of your problem. i.e. computing a weighted levenstein distance where the weights are the probabilities of making that error. However, I'd argue that this would still be better with centralized data; you can compute much better probability vectors. And regardless, the best solutions in the field will be with the combination of both.
Note 2: All of the above is speculation. While I help write some of the tools that these guys use, I have no knowledge of how they write their software. This is just how I'd do it.
A nitpick:
Auto-correction for a user's contacts could probably be done on-device, although I'd guess that machine learning across all users will probably massively reduce your success rate. Consider an ambiguous correction; you accidentally type "Gob", but have contacts of "Rob" and "Bob". I imagine that ranking the suggestions can be improved using a globally trained model.
Data is what lets us keep hospitals safe by establishing best practices, lets us cure diseases, lets us avoid unnecessary treatment, and just in general, separate fact from fiction. If we can't aggregate data, we can't do those things nearly as effectively as we could, and simplistic views of privacy have prevented this from happening.
There are plenty of concerns about data, and plenty of ways of mitigating those concerns, but a blanket condemnation of any centralized data would hold humanity back immeasurably (and in fact, already has.)
This line of reasoning seems similar to arguing that, by not creating a society in which 95% of people live in sanitized hospital rooms and the remaining 5% of people are doctors and nurses, we have "already caused many deaths".
Yes, the fact that people value things like "privacy", "being in large groups where others could be ill", and "moving from place to place" will "cause deaths". But that doesn't mean individuals should be forced to abandon these things for better health outcomes. While some benefits could certainly come from centralizing medical records, I tend to think that benefit cannot justify compelling people to keep their records in a centralized database, in the same way I don't think the health benefits of not traveling justify banning travel. You may not value privacy, but I think that in a truly inclusive society that values the individual, privacy should remain a real option for those who value it.
> There are plenty of concerns about data, and plenty of ways of mitigating those concerns, but a blanket condemnation of any centralized data would hold humanity back immeasurably (and in fact, already has.)
I agree that a blanket condemnation of centralized data is also not a good approach. Each individual should be able to make this choice in a meaningful way, with respect by default that any given person might care about it.
The thing I find most aggravating about it is that the standard for harm for data seems to be "What could a totalitarian government do with it?" and there are very few useful things that couldn't be used for very bad things in the hands of a totalitarian government (newspapers, for instance.) Meanwhile, companies can't reveal all the useful things that are consequences of their data because that makes them vulnerable to both competitors and spam.
So we're pretty much stuck with only uninformed opinions and worst-case scenario analysis, which isn't a rational way to approach anything. The only way I can think to improve the debate is for privacy advocates to focus on actual harm that has actually happened to someone to at least keep things grounded in reality.
Some people would gladly share medical information for the greater good if asked. You'd have to pry it from the cold, dead hands of others. The problem is that we currently don't separate those who would like to provide their information from those who would like to refuse because we're greedy for as much information as we can get, or we don't trust people to make the "right" decision.
A system that worked well and that accomodated both views would be one that really respected a user's choice--one that could handle data in a centralized or individual fashion, accordingly. Data that was willingly given could be used under the terms of the agreement without risk of angering people who don't want their data to be analyzed. It would ultimately provide the same benefits while respecting individuals and not creating controversy.
In the medical records situation, you'd have a group of people at one end that would be quite pleased to contribute their information, a group at the other end that would immediately decline, and quite a few people in the middle who would likely fall somewhat evenly to either side. Sure, some people would probably be fear-mongered out of sharing, but I imagine there'd be a lot less of that going on if people knew they could actually make a meaningful individual choice for privacy if they chose to do so. When opt-outs are buried, hidden, and it's not clear that they actually work, skepticism and fear grow. In a system that doesn't try to railroad people to sacrifice their privacy--a system that actually respects the individual's choice--fear would be reduced.
The only way to cast questions around this type of privacy as "a debate" involves changing the issue to compelled sharing of personal information. In a debate over forced participation, I think it's pretty easy to see why thoughts start to drift towards totalitarian concerns.
I think a system that offered real choice for each individual may drastically reduce the current problems from both perspectives, at the price of some engineering overhead.
There isn't any coercion going on with any of the things we're talking about. Every technology product has at least the choice to not use it.
It's quite difficult to ask meaningful questions about what users are comfortable with, get meaningful answers, and then figure out how the answer they've already given applies to a grey area situation where the cost of getting it wrong is a lawsuit. It becomes no longer sufficient to treat the data with respect and only use it for beneficial and privacy respecting purposes. You now have to constrain it by another set of rules whose relationship to what's actually happening can be unclear and arbitrary.
Some products work without storing any data. Most don't. For those that require user data to fulfill their basic function, the engineering cost of exempting certain data from certain systems can be much higher. One programmer screwing up becomes a lawsuit.
This is essentially the situation with HIPAA. Everyone is too worried about liability to do anything innovative, so that sector doesn't improve.
If someone is actually harmed by something a company does with user data, then it's entirely appropriate to stop patronizing that company, or to claim damages through all the normal routes. The presumption of not trusting anyone with data in advance of any actual harm is what I object to. Data can do real and permanent good in the world, and some companies are worth trusting (particularly since all of their incentives are to remain trustworthy if they want to continue to exist.)
But that's exactly the problem: many of the modern technologies that pose potential risks to privacy don't in practice provide an opt-out for the people whose privacy they might infringe.
Sometimes, you do have a choice. However, if you don't know about it or understand the implications, you can't make an informed decision.
Sometimes you don't get any meaningful choice, for example if governments decide it's OK to share sensitive healthcare information now. Strictly speaking you do have a choice, but that choice is never to visit a doctor or hospital. Try contrasting the dangers from a significant reduction in public trust in the integrity and ethics of the entire medical profession with the hypothetical future benefits of analysing aggregate healthcare data, and let me know which one really seems like the bigger risk.
Sometimes you don't get any choice in practice because your data is collected incidentally. When you were in the background of someone's personal holiday snap, that didn't really matter. When you're in the background of a CCTV image, which is centrally recorded and subject to future data mining operations, it matters more. When you're in the background of numerous CCTV images just because you left your home, which are subject to geotagging, facial recognition, gait analysis, covert audio recording, correlation with other databases such as mobile phone history, ANPR scans and purchase history, permanent archival and any additional data mining techniques that anyone who gets hold of the data might find later... Well, now you're in the plot of a sci-fi short story that doesn't end well.
Except that of course, it's not a story any more. Insurers already bump premiums based on profiling, but that profiling is notoriously inaccurate. Lenders already check credit records, which again are notoriously inaccurate. Employers already not only Google job applicants but in some cases also ask for personal log-in credentials to read through their social networking history. Governments already sell personal data held for legitimate public interest reasons to private parties, and even in seemingly simple cases like the government's vehicle licensing authority in the UK providing details of the registered owner of a car with given plates, this has been widely abused. Where these things have been curtailed -- which doesn't happen nearly as often as it should -- it has mostly been because primary legislation was passed or the rules for government's own departments were updated to cover specific cases, and only after so many people suffered from the intrusion that it became a politically significant issue.
To be clear, I don't object to the idea that there are potentially great benefits to be had from data mining, including in sensitive cases like public health data. But I think you are almost completely ignoring the accompanying risks, despite a seemingly endless stream of failures resulting in serious adverse consequences for individuals whose privacy wasn't adequately protected. We need the rest of how society works to catch up with the capabilities of modern technology before we can reap the benefits without paying too high a price.
Such as? Typically it seems that things become a scandal based on hypothetical harm rather than actual harm.
Of course, in reality it's often difficult to prove that a specific outcome was the result of a privacy invasion. It's not like insurers or employers are going to document that they discriminated unfairly, whether illegally or otherwise, in making their decisions. But we know all the things I mentioned can happen, partly because too many times there have been cases where real evidence was seen, and partly because in some cases incentives are aligned with poor behaviour and it's just plain naive to think it won't then happen if there's nothing to balance those incentives.
Privacy is important because it removes the ability to make those unfair decisions in the first place.
First off, let me clarify what I'm advocating for so we aren't talking past each other. I would like for the public to be less skeptical of organizations collecting large amounts of data, and storing it to analyze in aggregate for a variety of purposes. Particularly, if access to the data is controlled, and if it is only used in a sufficiently aggregated form. Society will reap tremendous benefits from enabling things like this.
I don't think of most of the cases you are describing as being related to this.
That's a very one-sided view. Just this week, the latest attempt to do this in the UK, the care.data programme, essentially became so politically toxic that it's dead.
This happened for a number of reasons. Some of them were just incompetence, like claiming everyone would receive a leaflet explaining the proposals and the right to opt out, and then finding that not only was your leaflet heavily criticised by medical and IT professionals for being woefully misleading, but when surveyed about 2/3 of adults reported not having seen it anyway. That was the credibility of the programme operators you saw falling down the sinkhole.
However, other reasons for objecting would have stood up even if everyone were fully informed. The data in question wasn't actually going to be available to clinicians like doctors and nurses who might find it useful when providing care. And contrary to the laudible-sounding goals that some medical professionals have suggested, much like those you have been advocating yourself in this discussion, it also wasn't going to be restricted to people like medical researchers.
In fact -- and it is now well-established, beyond-any-doubt, clear-as-day fact -- the data could never have been protected to the extent that was claimed (numerous qualified people have debunked the effectiveness of the claimed pseudonymisation), and the proposed rules and "safeguards" for who would have access to the data and for what purposes weren't even close to restricting it to legitimate medical research of the kind you describe. Those advocating the scheme at government level have once again demonstrated a fundamental lack of understanding of the implications of this kind of technology. There are a few other questions that seem to have been brushed under the carpet, too, like how opting out would supposedly mean your information never physically left your GP's systems, yet paradoxically there were circumstances discussed a couple of weeks ago where some organisations, like police and security services, would be able to access the data centrally via the new system anyway.
The trouble with many of these privacy issues we've been discussing recently -- whether it's Google-esque creepy mass surveillance, or the NHS plans to consolidate and share particularly sensitive data about individuals, or governments monitoring surveillance networks -- is that they are all cases of Pandora's box. Once you've compelled people to give up privacy and they've been entered into someone's database, that data is out there, and it's subject to redistribution and repurposing at any future time, with or without the blessing of the data subjects. Our privacy laws are dangerously underweight and already fail to balance the heavyweight capabilities of modern technologies. Until that gets fixed -- and I mean fixed in the sense that privacy laws are actually enforced and respected at government level -- the only reasonable conclusion is that we should err on the side of caution with giving up personal data, and challenge every attempt to push the boundaries to make sure it's justified.
Of course we should encourage the good and discourage the evil, but that's a moral pursuit, not a technological one.
The world would be a better place if there were no defined country names, zip codes. We should decentralize all our data to prevent evil.