Scalable and accurate deep learning for electronic health records
arxiv.org
arxiv.org
How much benefit does the average person get from being profiled like crazy across the web? Besides the (dubious) benefit of free services, there isnt much positive you can point to from such activity, and a shitload of really scummy exploitative crap.
I dont see how such medical analysis is really going to pan out much differently at scale.
Health insurers can't deny you coverage or raise prices due to any pre-existing health condition. By law, an insurer has to offer the same rate to everybody with the same age, zip code, and smoking status.
Insurers are allowed to rebate up to 30% of your premiums if you take healthy actions (e.g., going for a run, joining a weight loss program, medication adherence), as long as those actions don't discriminate based on health status.
Combine that experience with some good reference, research and consultation with lots of other doctors and you would have a much better outcome overall.
I do not complain when IntelliJ autocompletes my lambda function in Java 8 and get defensive and yell at the IDE. (Excluding who prefer no assist code editors)
We lament unrealistic conditions in white board interviews without Google, but have surprising double standards when it comes to medical professionals.
Lol, as someone who criticized Google on multiple occasions, just wanted to point out that it’s OK to call them out when they (may be) wrong and still like their technology.
I'm thrilled that my GP is willing to slog through research, or look up things she's not 100% sure of. Sure, I've got access to much of the same information. She's got the training to interpret it though. We were discussing vitamin supplements and fortification at our last visit. I wasn't entirely sure about a couple of the claims she made, so I dug up some articles and emailed them to her. A couple days later I got a response with an interpretation of the articles (essentially, nothing to worry about).
Or, how about this? I've got a rare variant of a fairly common condition (such that a GP probably wouldn't see this firsthand). I'd seen a specialist at a world renowned hospital before. Those visits were expensive and not very informative. Visit with the GP? Huh that's unusual, let's see what the options are. Her conclusion? You can go through a variety of expensive tests, but unless I'm symptomatic the best course of action is no action. Now, I was curious, so I looked all that up after the visit. Fuck if I know how to discern how expensive those tests are or whether or not I'd see any benefit. There's a ton of information out there, but the big bucks are spent on people who a.) know how to find it and b.) how to interpret it.
To bring it to a tech setting -- let's think about Jenkins. I've cobbled together a build pipeline for chef cookbooks for one of our apps across our prod and non-prod environments. Sure, we pull credentials from the Jenkins credential store. No, I can't tell you off the top of my head the exact syntax to pull a secret out of there. But I can tell you where I'd look for instructions. Being able to be productive with Groovy (yuck) is part of my job, but it's not the only part of my job. Being able to find the relevant information is far more important than having to memorize it.
Honestly though the biggest problem I have is that there's nearly zero pipeline documentation for Jenkins. The language itself is neither here nor there (but by virtue of not being a Java guy, I don't have a particularly robust development environment setup).
I say this because this problem is vastly more complicated than people believe for the following reasons:
* A substantial amount of data about patients is encapsulated in clinical notes which are free text. There is no standard format for notes; notes wildly vary in terms of quality, detail, and accuracy; and notes are full of non-standard abbreviations and clinical shorthand.
* A large amount of the data in an EHR is incomplete or incorrect. Information may be out of date or recorded incorrectly. Patient medication lists are a great example of this, and they frequently cannot be trusted by physicians even during the course of a single hospital stay (medications are started or discontinued without being properly recorded). Then there is the issue of non-adherence: patients frequently claim to be adherent to treatment regimens that they are not actually adherent to which is going to distort outcome data. In fact, the majority of people are non-adherent for many different treatment regimens. More generally, diagnoses may have been incorrectly determined and data may be incorrectly labeled or missing.
* Data exists all over the place even within an EHR, let alone across facilities. I myself have been to at least 5-10 hospitals and more outpatient clinics over my life, none of whom share data with each other. No one entity has all the relevant data about a patient, including insurers, specific care teams, or the hospital systems they got this data from.
* Data placed into a EHR may be intentionally falsified for simplicity or to get insurance coverage. It is not uncommon for physicians to intentionally misdiagnose a patient so their insurance will cover a diagnostic / treatment / procedure, or to input a less specific diagnostic code to save time when a more specific one is available.
Trying to train with data sets this fucked up is an enormous challenge.
In terms of incorrect records, I'm not sure this would be a major problem over a large dataset. Making this up, but I imagine the range of errors is fairly random across all fields and with enough data you can always identify an approach that appears to save a life more than some other approach, even if the diagnosis is incorrect.
Although I totally agree with you that this is going to be an enormous challenge for Google, I feel like it would be underestimating them to assume this has very little hope of succeeding.
The only important part is making sure that the training labels are accurate. If the training data has inaccuracies, the model should learn to work around them (with some accuracy cost). The main question simply is how much signal is in the data. And that's hard to predict in advance (see the recent surprising results about predicting sexual orientation from faces).
Those results were not so much surprising, as coerced. For example, training data included an equal amount of men of both sexual orientations and the same for women- an equal distribution, in stark contrast to the one in the real world. This was clearly done to overcome the inevitably poor accuracy when training with unbalanced classes.
In short: don't base anything on that paper, for it was a steaming pile.
We changed the URL from https://qz.com/1189730/google-is-using-46-billion-data-point....
If this, however, is something that takes multiple patients records into consideration, this would most likely never see the light of day. Most hospitals don't want anyone knowing that, for instance, their city has the highest number of leukemia related deaths per capita than any other city in the US.
(e.g. I had to give my address three times to the same outfit last month just to get a flu jab.)
But we can't go back to paper records so we'll try building a giant brain (AI) to try to decipher and manage millions of electronic records. Trouble is that that brain still won't have written the records or talked with the people they're about. So I doubt it will work.
How much did this data cost?
Considering they have the ability to pick from thousands of hospitals and universities (the latter being where they actually got the data) they probably have a fairly large amount of leverage to keep the costs low.
Also, 200k patients is actually kind of small. Granted this dataset is far more granular/robust than what you’d typically find in commercially available healthcare datasets, but to give you some frame of reference, the healthcare datasets I work with contain > 20 million individuals (again, with orders of magnitude fewer features).
> To conduct the study, Google obtained de-identified data of 216,221 adults, with more than 46 billion data points between them. The data span 11 combined years at two hospitals, University of California San Francisco Medical Center (from 2012-2016) and University of Chicago Medicine (2009-2016).
I know I'm not supposed to say 'Did you read the article?', so I won't say that.
Yes there are problems with combining that data with other datasets to re-identify. Google would be in a rather good place to do that.
http://www.zdnet.com/article/re-identification-possible-with...
Unless Google starts using fully homomorphic encryption or something similar, then you should consider that data non-anonymous. And until that happens, Google and national healthcare agencies should also require consent from the patients before giving the data to Google or any other company for such studies. Otherwise, I hope class action lawsuits will be started against both.
I've went through 10 research papers on this very subject last week, all from 2016-2017, and this is a huge problem.
Just because google does it, people assume it's infallible, where in reality there are bunch of methods to deanonymize data, from machine learning to pure mathematical models on binary vs categorical inputs and such.
In this day and age, somebody will publish the deanonymized, cross-referenced (from voting registry or whatever) set online for free searching.
Then again, most people are fine with this, as long as their names aren't on the list, so considering it's 200k sample out of 323M population, at most 0.7% will be really outraged. Which may explain the down votes he's receiving. 99.3% names aren't on the list, and from that set they want the scoop on the rest of their neighbors.
Come back with your criticisms when you have actual evidence of Google mishandling data or misleading the IRB boards.
And the core foundation of medical ethics is that the patient gets to decide if/how much of that risk to take.