> To conduct the study, Google obtained de-identified data of 216,221 adults, with more than 46 billion data points between them. The data span 11 combined years at two hospitals, University of California San Francisco Medical Center (from 2012-2016) and University of Chicago Medicine (2009-2016).
I know I'm not supposed to say 'Did you read the article?', so I won't say that.
Yes there are problems with combining that data with other datasets to re-identify. Google would be in a rather good place to do that.
http://www.zdnet.com/article/re-identification-possible-with...
Unless Google starts using fully homomorphic encryption or something similar, then you should consider that data non-anonymous. And until that happens, Google and national healthcare agencies should also require consent from the patients before giving the data to Google or any other company for such studies. Otherwise, I hope class action lawsuits will be started against both.
I've went through 10 research papers on this very subject last week, all from 2016-2017, and this is a huge problem.
Just because google does it, people assume it's infallible, where in reality there are bunch of methods to deanonymize data, from machine learning to pure mathematical models on binary vs categorical inputs and such.
In this day and age, somebody will publish the deanonymized, cross-referenced (from voting registry or whatever) set online for free searching.
Then again, most people are fine with this, as long as their names aren't on the list, so considering it's 200k sample out of 323M population, at most 0.7% will be really outraged. Which may explain the down votes he's receiving. 99.3% names aren't on the list, and from that set they want the scoop on the rest of their neighbors.
Come back with your criticisms when you have actual evidence of Google mishandling data or misleading the IRB boards.
And the core foundation of medical ethics is that the patient gets to decide if/how much of that risk to take.