Google AI has access to huge haul of NHS patient data
newscientist.com
newscientist.com
For more than a decade researchers have been able to get access to large deidentified datasets. As a PhD student I have access to data on 150 Million visits by 45 Million patients. In some ways the data and access I have is superior to that of Google & NHS (since UK is a tiny country in comparison to USA), and I am just a PhD student. Though I have been working on it for last 5 years.
Recently there is a new qualified entity program run by CMS which provides access to Medicare data.
You can read more about my research and see the demo of the system in my past submissions and at
http://www.computationalhealthcare.com
Also for the actual govenment program http://www.ahrq.gov/research/data/index.html
Actually when it comes to Medical information its a much much more complex problem legally. This paper gives a good overview on issue of "Patient ownership of data" http://papers.ssrn.com/sol3/papers.cfm?abstract_id=1857986
I am by no means minimizing the concenrs that people have. But I think the article paints Google in negative light while ignoring current standard practices. I wish the discussion would be more rooted in the facts about how such data sharing systems work currently.
Because politicians keep making promises in regards to privacy that they don't keep in regards to "confidential" information.
Similarly, most of the medical value from such information [e.g. Frequency of X within a given population fitting certain characteristics] would likely deanonymize people.
> "All necessary safeguards would be in place to ensure protection of patients' details - the data will be anonymised and the process will be carefully and robustly regulated.
> "Proper regulation and essential safeguards need to be in place when it comes to patients data," he said. "It cannot be done in a way where essential rules are threatened."
The legality of these data sharing laws start off with public promises of anonymization, robust regulation, safeguards, and privacy.
https://www.england.nhs.uk/2014/01/geraint-lewis/
> Amber data are where we remove each patient’s identifiers (their date of birth, postcode, and so on) and replace them with a meaningless pseudonym that bears no relationship to their “real world” identity. Amber data are essential for tracking how individuals interact with the different parts of the NHS and social care over time. For example, using amber data we can see how the NHS cares for cohorts of patients who are admitted repeatedly to hospital but who seldom visit their GP. In theory, a determined analyst could attempt to re-identify individuals within amber data by linking them to other data sets. For this reason, we never publish amber data. Instead, amber data are only made available under a legal contract to approved analysts for approved purposes. The contract stipulates how the data must be stored and protected, and how the data must be destroyed afterwards. Any attempt to re-identify an individual is strictly prohibited and there is a range of criminal and civil penalties for any infringements.
The problem with psuedonymous data is the NHS basically admits it can be used to identify people given sufficient effort.
---
That is why people are "astonished" by these decisions. The politician provides the initial promises that imply anonymity, the implementation doesn't provide true anonymity but provides criminal penalties for pulling off the mask, and then the data is handed to enough 3rd parties if such data is leaked its likely impossible to know by whom unless the data was tampered with to provide a per-contract identifier.
I understand this specific decision did not involve a politician but the conversation was why people are surprised. How many people do you think really know the anonymity originally promised became a permeable pseudonym?
It seems you have disagreement with the system adopted by NHS and several others worldwide. This has nothing to do with Google or Politicians. The system emerged from decades of research and understanding of compromise between the need to protect privacy and advancement of medical research. You are free to suggest alternatives over current system. I have studied this problem for last five years and there isn't a simple solution.
2) You asked why they were astonished. Well, #1 is why. Most people don't have the time/energy/desire to study the implementation details on every facet of their lives.
This Google thing has fuck all to do with care.data - they're totally separate.
I understand "the" NHS is complex, but it's pretty frustrating talking to someone who has very strong opinions and who clearly doesn't know what they're talking about.
It's really weird to link to a document that talks about the severe legal penalties for anyone who attempts to de-anonymise the data, and then use that to say "look how flimsy these agreements are!", especially when the document you link to has a BOLD lead saying that things are even stricter in the newer document.
> The politician provides the initial promises
Again, not a politician. Chief data officer at NHS England, and a real doctor. http://www.nuffieldtrust.org.uk/about/our-people/dr-geraint-...
> The problem with psuedonymous data is the NHS basically admits it can be used to identify people given sufficient effort.
It's trivially easy for Google to do this already without the NHS data, and they don't face prison time for doing it. See all the pregnant teens outed by supermarket loyalty cards for other examples.
And how many members of the general population do you think are aware of this?
> I understand "the" NHS is complex, but it's pretty frustrating talking to someone who has very strong opinions and who clearly doesn't know what they're talking about.
It probably has something to do with the fact you are completely missing the point I'm discussing rather than the strength of my opinions.
As an individual, you may wish to hoard all of your personal medical information. Doing so may provide marginal benefit or protection against some theoretical harms. However, much of medical science relies on large population studies. Is knee surgery worth it? Are breast exams? Do COX-2 inhibitors increase heart attacks? Enormous amounts of good could be done by having large datasets of entire patient histories available for analysis by all, assuming they could not be de-anonymized (if not, then the datasets have to be analyzed under contracts)
Whenever I visit the doctor and I am given a form asking if my data can be sent to some group to participate in a study, I always answer yes. Never once have I had any negative repercussions from doing so, and it is my hope that the data was used to publish scientific papers that added to net knowledge of humanity. However, I have to wonder if asking me for permission actually interferes with random sampling. Are people who give permission vs people who refuse statistically more likely to have other behaviors that may influence the results?
That's where the problem is anonymizing data is really difficult and it takes only a few bits of info to associate that data with a name. It takes only 3 pieces of information to uniquely identify 87% of people, Zip code, birthday, and sex.
http://arstechnica.com/tech-policy/2009/09/your-secrets-live...
http://dataprivacylab.org/projects/identifiability/paper1.pd...
If this was about positive externalities, Google tech could have found its way into the NHS.
Instead, as always and every time, data found a way into Google.
Funny how that works.
And that's your choice. Unfortunately, someone else committed suicide because they didn't ask for the help they needed for their mental health issues because of justified concern that their condition would leak and it would lead to discrimination.
There are many reasons privacy is important. Among the most important is the ability to have completely open discussions with clinical professionals about any medical issues you have, without having to worry that anyone has any ulterior motives or that anything you say could come back to haunt you later.
It is difficult enough for some people to talk about some conditions even with their own doctor in the privacy of the doctor's office. Probably some valuable benefits would result if we dropped privacy altogether and allowed widespread analysis of clinical data. However, it is an absolute certainty that serious damage is done both when patient confidentiality is perceived to be threatened (through patients not raising issues in the first place) and when patient confidentiality is in fact compromised (through discrimination of various kinds, much of which may be based on misconceptions).
I've grown cynical. My first thought on seeing this kind of form is that a third party will be making money from the information I provide somehow - whether it's through patented drugs or treatments or through pay journals like those owned by Elsevier. Neither I, nor anyone I know, will get to benefit from from the information I provide without spending money (often exorbitant amounts of it). This is the flip side of privatizing everything. Why should I lift a finger to help for-profit entities when they will not do the same to help me?
But if releasing your data leads to research that can save lives, isn't that worth doing even if someone gets rich in the process? I'd much rather there exist a million-dollar treatment that can save my life at the cost of living in debt for the next twenty years, rather than die in six months because the research was never done.
(Yes, this is a false dichotomy - there should be a better way of administering treatments and doing medical research.)
The idea that someone, somewhere is making a profit on something you provided should make you happy, not sad. To make a profit, they had to be paid, and to be paid they normally have to have provided something that someone wanted enough to part with cash. That cash changed hands is not the main point of the transaction (it's a net zero for society, someone gained cash and someone lost it), the main point is rather that something of value was created.
You threw out the example of pay journals like Elsevier. It is true that for cases like that, there may be a market inefficiency that means not much value is created. But I think those are edge cases and relatively rare. Even if our worst fears about Elsevier and other rent seeking journals are true, it would still be the case that you're helping rather than hurting by providing anonymized personal data.
You might have sold your own data for a higher price if you had known.
This seems like a double standard. Why should I be happy when corporations will regularly go to court to make sure that nothing they create returns to the public domain within my lifetime? The norm seems to be capturing this kind value wherever (and for a long as) possible, not distributing it freely for the good of mankind.
I'm not making a normative argument that this is a healthy way to live or for society to operate. In a functional community resources, ideas, and capabilities may be shared freely for the benefit of all, but it can't always go one way. That's exploitation, not community.
I'd rather hideously expensive treatments exist, even if I can't afford them, than live in a world where they're never even an option. Price can and does change over time, and those treatments become more accessible to more people. But if the treatments are never developed in the first place, then that's an even greater tragedy because then there's literally nothing that can be done to make them accessible.
Especially when that private entity gains more power through my data and tries everything from questionable business practices to lobbying to make sure they, and only they, reap the benefits.
People value chemical weapons and forced labor as well. They're willing to pay a lot for them, too.
You aren't giving up something of value to you. There's zero opportunity cost involved. The value is added after your data is collected, collated, analyzed, and used by researchers. Individually, it's meaningless. Together, it's priceless.
Will someone eventually profit? Sure, but they'll do so because society benefits from the work they've done. And eventually, the cost will decrease. Would you rather have drug treatments that exist in the first place even if someone profits from them or nothing? Those are pretty much your options. For all of its flaws, the current research model at least works. The only other option, besides doing nothing and letting people die from diseases we could eventually treat, is some sort of public financing and that comes with a host of problems. Would anyone really be so foolish as to want to subject medical research to the congressional budget process, even if there's an abstraction layer between congress and the researchers? That means tradeoffs, and lots of them: for instance, Drug option A would be pursued, and option B ignored because we're already funding A even if B might be more promising later on.
We already see it with federal applied research when critical work is jeopardized for the sake of political grandstanding. Every so often, someone will trot out a cherry-picked grant they don't like and wave it around like a red flag in front of a bull. Drug companies might be bad enough, but the politicians would be worse.
I think in this case and others like it the family, at minimum, has a right to know. "...it was not until 1973, when a scientist called to ask for blood samples to study the genes her children had inherited from her, that Ms. Lacks’s family learned that their mother’s cells were, in effect, scattered across the planet. Some members of the family tried to find more information. Some wanted a portion of the profits that companies were earning from research on HeLa cells. They were largely ignored for years."
http://www.nytimes.com/2013/08/08/science/after-decades-of-r...
What are the criteria for who gets access? What are the constraints of that access?
This story covers the latter being blown apart, the constraints were poorly defined and implemented and thus even if the criteria is well defined access to far more data was made possible.
I'm sure that few patients desire an end to research, or would argue that such access isn't a good thing... but what of the insurance industry? Should they have access? Would the NHS be able to define and enforce those constraints?
Perhaps that's an obvious no.
What then of an insurer partnering with a medical research company, from the viewpoint of "This costs insurance a lot of money, we'd like to fund a way to reduce that financial exposure".
The grey areas emerge immediately.
If we cannot control access to patient data, data that would be trivial to either strip anonymity or just to have in aggregate enough to still produce net-negatives (i.e. correlated by post code would reveal enough with little extra work)... and if we cannot define and enforce the constraints of access... then we really shouldn't be sharing what is highly sensitive and personal information that was originally only disclosed between a patient and a Doctor under the premise that what is shared is covered by the explicit and implicit confidentiality of that conversation.
It's always worth remembering:
Data was acquired under doctor patient confidentiality.
If we considered that data to have a licence, it is the most restrictive licence possible. One could consider what has happened here as a re-licensing without permission. Such an act could have a chilling effect on the relationship between the doctor and patient.
I have seen a few of these sorts of deals killed because of data access concerns, and/or computation requirements ("you can have access to anonymized data, but you have to run your code in a sandbox on our health servers").
And, this is why we have legislation.
> The scale of the sharing program was apparently misrepresented to the public, originally announced as an app to help hospitals monitor patients with kidney disease with real-time alerts and analytics. But since those patients don't have their own separate dataset, Google has argued it needs access to all patient data from the participating hospitals.
No assumption there, they didn't have a separate dataset and so granted access to all patient data.
Yes, but under what conditions? Many privacy laws apply here, and treating Google as some monolithic entity where everyone working there can now read anyone's personal health history is inaccurate.
Here's HHS on what HIPAA has to say about this: [0]
[0] http://www.hhs.gov/hipaa/for-professionals/privacy/special-t...
happy to address criticism
Will definitely be looking into healthcare data more, as this story has resulted in some interesting leads
I mean, other than that time just a few years ago[0] where Google broke the law and then breached the contract they signed with the UK Government.
Google is a corporation. It can't have good intentions of its own. It's the thousands of employees who will potentially be working with and handling the data that you need to worry about.
Getting access to private information in Google is hard - my experience as a researcher here is that there's a strong incentive to find an open-source or non-PII dataset before touching user data. I'll go through my year here without ever touching even the most innocuous PII data.
It's very unlikely to me that thousands of people will have access to this data. It's much more likely that a small handful will, and that they'll be supported by others with no access whatsoever. From the article, in fact:
"The agreement clearly states that Google cannot use the data in any other part of its business. The data itself will be stored in the UK by a third party contracted by Google, not in DeepMind’s offices. DeepMind is also obliged to delete its copy of the data when the agreement expires at the end of September 2017."
From an incentive perspective, the potential value-add of abusing the data is tiny compared to the potential costs and loss of user trust. Google's very aware of how important it is to maintain user trust -- http://www.techradar.com/us/news/internet/google-we-have-a-c...
Corporations don't have brains, but they have cultures, and Google's culture -- composed of those thousands of engineers -- is quite fanatical about protecting user privacy. It's one of the non-technical things that's impressed me most during my time here.
The risk with a company like Google is if the economic winds and culture changes, but that's a long-term process, and is also the reason for legally-binding contracts to do things like delete the data (see above).
tl;dr: Google has the technical means to protect the confidential data better than almost any other agency, including from its own employees. The most important question to ask is whether the NHS structured the data sharing in a way that provides for long-term protection, and (IANAL!) it sounds like it from the article.
Source: I'm a professor who deals with our IRB occasionally, have colleagues doing joint CS-medical research, pushed patients around a hospital in a younger life, and am on sabbatical for the year at Google.
Do you mean that it's not sufficient to have a medical record that doesn't indicate knee issues? That you would need a medical record that confirms there are no knee issues?
Consider: I may have a medical record that doesn't indicate knee issues because I have correctly been diagnosed with no knee issues. Or I may have such a record because I have been incorrectly so diagnosed, despite the presence of knee issues. Or I may have such a record because I haven't been to a doctor about my knee issues in the years since they've developed, and so no opportunity has arisen for my knee issues to be documented. Simply from a medical record reflecting no evidence of knee issues, you cannot know which of these, if any, is true.
As I mentioned in a reply to your neighbor comment, clinical studies are designed to exclude such confounders as these - and that's a considerable part of why such studies are so expensive to design and carry out. It's very difficult to see how such exclusion could be achieved with data of the quality which seems to be involved here.
1000 patients come in with symptoms that look like cancer at year 0.
100 actually get diagnosed with cancer at some point between year 0 and year 5.
Presumably, the remaining 900 didn't have cancer at year 0.
Especially for the NHS dataset, since you will either see the patient in there or in a death index (unlike, say, US where they may have just gone to another hospital).
Also, the scope here is more like a longitudinal vaccine study than a clinical trial. 50M people will provide a lot of robustness that you wouldn't see in a 1000 patient trial.
Or you won't see any new information for the patient, because the patient is lost to followup. Or - worse - you'll see new information, but it'll be invisibly erroneous, because random GPs don't work to the standard that physicians administering examinations in clinical studies do.
> 50M people will provide a lot of robustness that you wouldn't see in a 1000 patient trial
Not if the 1000-patient trial is well designed, and the data of those 50M people is totally uncontrolled and unverified. This isn't warfare - Stalin's dictum has no place here. You can't overcome the flaws of a dirty dataset by adding more dirty data to it, especially when you literally cannot know either the magnitude or the nature of the inaccuracy, or even tell what's accurate from what's not.
So you will take this into account and emphasize the more reliable types of tests, like blood tests. Or you'll find ways of learning which doctors who do the tests more accurately(or train doctors to do so) and which people are more consistent/reliable with their relationship with their docs. Or maybe you'll get a few hypotheses which are relatively likely and that would incentivize the researchers/google to do small clinical trials on them.
It's worth a try at least.
They could also look at any advertising material and report that to ASA (if it meets the criteria for being regulated).
That trust has a Caldicott Guardian who will be responsible - legally - for keeping patient data safe. I would have liked a quote from them, although I guess that quote would be something like "No patient identifiable data has been shared with DeepMind". That would be scary, I have no doubt that DeepMind would do a very good job of de-anonymising data, but I know that Google would have to be monumentally stupid to try that.
There are other data projects happening in the NHS - care.data (that dot isn't a typo!) is one that got a lot of attention. That allowed (after some fuss) people to opt-out. (It didn't allow people to specifically opt in to show their support, which is something I would have done.)
I'm a bit wary of Vice's reporting here. They don't seem to know what they're talking about (there's nothing about controls over patient data in the NHS, for example); they don't seem to have approached the Trust involved; they haven't done a good job of explaining what's going on.
There are some really bad failures or data protection in the NHS (especially around mass email! People using CC instead of BCC to a group of people using an HIV clinic, for example) and there are some historic abuses (selling data to insurance companies) that led to changes in the law.
So I don't know if this is terrible and deserving of anger, or okay and poorly reported, or a good thing with misleading reporting.
Putting that here because I was confused about what NHS was in the first place (I'm French).
[0]http://www.hopkinsmedicine.org/institutional_review_board/hi...
Personally, I see potential in ResearchKit to solve the consent problem re medical data
I have a small team that is working with leaders in the space[1] to help many of the major EMR vendors support an open-standards based approach to medical record sharing.
We'll be seeking public feedback when some of the preliminary work is ready, but are very excited to get input from the community as we make progress.
Those two aren't inclusive of each other.
In this case the government would be the Department of Health. They've had no involvement with this.
"The NHS" would be NHS England, and they didn't set this up, although they might sta.rt being involved to check the controls.
The local NHS, the Clinical Commissioning Group, didn't set this up.
So we're talking about one hospital trust