Hospitals are selling troves of medical data
theverge.com
theverge.com
I’m not sure how to defend against this as it seems based on HIPAA and what it allows. Since de-identified data can be legally sold, I think it will.
The theoretical defense I’ve thought up is a class action lawsuit for synthetic breaches. Since these data are deemed de-identified by expert determination [0] and that’s hazy, if I could reidentify myself after de-id, and I didn’t authorize it, then I could be eligible for breach damages for HIPAA violations up to $50k per person [1].
Since these sets have millions of people, likely everyone in the country. And since expert determination can possibly classify an acceptable re-id risk as less than 1%, this could be a million or two people. So a big enough pool to attract big legal investments.
That would increase the cost and risk of doing this to outweigh the benefits. But currently it’s “free money” for any healthcare system that’s kind of impossible for me as a patient to opt out.
[0] https://www.hhs.gov/hipaa/for-professionals/privacy/special-... [1] https://www.injuryclaimcoach.com/hipaa-violations.html
I hope for all our sakes you are in the business of buying de-identified medical data :)
You may or may not have a private civil tort against the medical provider, separately from HIPAA.
I specifically linked to an injury lawyer to show the types of current HIPAA violation civil suits that have succeeded. My reasoning is that it’s lucrative enough for lawyers to solicit business.
I consider intentional breaches of medical privacy to be in the same league as physical assault and as such would expect a society to respond to these kinds of crimes with swift and severe punishments including jail time and fines.
Why not just arrest these people and take the profits from their crime instead of adding civil lawyers and bureaucracy to the mix?
Can you point to any country with a law on the books like this?
But of course that's not what TFA was about.
[0] https://inspiredelearning.com/blog/hipaa-101-guide-violation...
Right, there are criminal penalties possible under HIPAA (I don't know if anyone has actually gone to jail under them) but they explicitly don't apply to the situation I was replying to (or I misunderstood it).
Disclaimer: I work for a company that uses this data to match terminally ill patients with niche treatments and clinical trials, and the work literally saves and prolongs lives.
Talked with a company that made similar claims about what they are doing, but it turns out they were really ensuring that their users where choosing the medication from the highest bidder rather than the one that was really in the the patient's best interests.
The mental gymnastics they did to justify they were doing what was in the patient's best interests was fun to observe, but in the end they were just a tool of a large pharmaceutical company, exploiting sick people for profit.
Quick cynicism sanity check: who pays your company, your users or the "niche" treatment provider?
If it's the former, that sounds like it's good work.
If it's the latter I would recommend being a bit more skeptical of the business motives of your employer and their customers.
Insurers. They don’t want to pay for the expensive product.
Of course, not only is it distinctly possible to contribute to society and make a profit—it’s basically the only way. In the real world loads of researchers and engineers and other employees aren’t going to work for free to appease random Internet cynics (who rarely if ever make positive contributions to society of their own, mind you), so capital is required which entails investors and returns on investment. Yeah, that means insurers or someone else paying money.
It's surprisingly hard, which is why there is a cottage industry built up around it for clinical trials.
Regardless, getting informed consent for data acquisition is great if you are doing a clinical trial or really looking forward, but for a lot of things you want a ton of retrospective data that can be difficult or impossible to consent.
It's a hard problem. I hope in the future it will be an easy problem, but we aren't there.
I have literally yet to see a medical-adjacent company not default to assuming they’re on the moral high ground.
When I see companies say “we save lives” my initial reaction is no different from companies that say “we hire the best” or laws that “are for the children”.
Maps Street View of 1428 Bush St, San Francisco is more of what I have in mind. But maybe I’ve just passed that building one too many times.
Marketing drugs better isn’t exactly saving lives. Researching drugs should be done through IRB and include patient consent. Of course, there’s great benefit to bypassing ethical controls, but they are still worthwhile.
Conflict of interest regulations should be applied for data, If my data can be used to cause harm to me then you shouldn't have it that's all.
The Expert Determination data would likely be released to you with an agreement that you would not attempt to re-identify the data. So if you then attempted to re-identify the data, even yourself, you would be breaking the terms of the license and I would think liable to civil action from the entity that gave you the data (and their upstreams). I don't think you would win here.
You can either re-identify the people behind the data or you alter it that strong that it becomes useless for meaningful applications.
I recently had a crazy notion for losslessly scrambling the sequence as well. Mostly for protecting voter privacy (order in which ballots are cast). One of the major blockers to fully digital voting.
I haven't found any hits using terms like "cryptographic timestamps." Surely I can't be the first.
Sadly, others will continue to push their harebrained ideas.
One thing I learned as an activist is that offense beats defense. Meaning it's easier to promote a correct solution than oppose all the bad solutions.
So if a digital equivalent of the Australian Ballot system (private voting, public counting) exists, I better find it.
There are other concept in research to still use this data in research like differential privacy or using KI to synthesize data according to training data provided. But so far all concepts that tried to alter and then publish a datasets directly for research failed miserably.
GP said sex and year of birth that are usually perfectly fine for deidentification assuming some basic k-anonymity and l-diversity protections.
It’s frustrating when people bring up examples of reidentification out of data that were not properly reidentified in the first place.
Hopefully no one competent would think that having unique records based on sex, dob, and zip code are makes data deidentified. This is usually the case of someone not actually deidentified.
The bigger risk, I think, is when you have some threshold of at least 10 records sharing sex, year of birth and still iding individuals.
Sex and birth year only creates 100 total buckets for people born in the last 50 years. There is no way you could deanonymize that.
As has been said, I don't think there is a way to anonymize data, that is useful and not reversible when combined with other data. Even if you only have 2 buckets.
For example, perhaps you take data on favorite movies and strip away every bit of PII except sex and birth year (as you stated).
Now let’s say I take the stripped data and target some facebook ads to people of a specific sex and birth year and have the ad be for people who love a certain obscure movie, doubling down on a second obscure movie and so on. Eventually I could have a reasonable chance of determining which other movies a unique individual may like based on the stripped data.
Obviously irrelevant for movies, but more relevant for prescription drug uses, sexual preferences etc.
The point is that with enough outside data available, even data stripped of PII can be de-anonimized.
People don't yet have intuition about this. I still don't even know how to articulate it. Here's a stab:
"With enough data collected, you can uniquely identify people by ruling out everyone else."
Mid 2000s, Seisent was helping law enforcement solve cold cases using big data. In layperson's terms, they'd narrow the list of suspects by ruling out everyone who has a solid alibi.
At the time, that was achieved by building profiles of everyone, living and dead, simply by compiling 1600 publicly available datasets. Like court records and mortgages and whatnot.
Today you'd include location tracking, social media, all financial tracking, etc.
It's remarkable that any crime goes unsolved. Like the back log of rape test kits, it's only because no one cares enough to bother to look.
It's a pretty neat idea overall, reminds me of those link-following competitions using Wikipedia where you have to start from a topic and get to another completely unrelated topic solely by recursively chasing links. They both exploit the interconnected nature of our culture, a question (or an article) about a recent TV show providing a clue (or a link) about a pretty unrelated historical character.
The game desperately needs some notion of statistical proximity though, it kept asking whether the character was an actor even after knowing he is a president, it doesn't do _any_ kind of deduction on what it already knows, just mad slash-and-dash till it gets the identifying info.
Practically, this will likely yield the same results.
Remember I'm still struggling with phrasing:
I have no idea how akinator is implemented. 20 questions is traditionally a drill down decision tree. From their blurb, I infer that akinator divines its own decision tree, vs user supplied questions. Clever.
Instead of finding a needle in a haystack, the big data strategy is to progressively remove all the hay until the needle is revealed.
Like the big data equiv of Sherlock Holmes' deduction (logic). "When you have eliminated the impossible, whatever left, however improbable, must be the truth."
Again, am not an academic, philosopher. I don't know what to call this strategy.
It's not logic, reasoning. It's brute force. Test every single possibility to find the matches.
In other words, it's a database query.
Ok, maybe I'm just confusing myself here.
I worked with a hospital awhile back with patient radiological data (for free and for science). Patients had to explicly sign-off that they were sharing their data with us and what we planned to use it for. A lot of the metadata from DICOM wasn't even wiped, I had their names, street addresses, all sorts of stuff which was supposed to be wiped (de-identified).
Even then, I worked with data from a patient with a brain tumor and their neurosurgeon looking to remove it. That data was correctly anonymoized but it was their head, so I could basically reconstruct their face--so is a coarse geomerric sampling of a face de-identified? I guess it depends on how coarse it was.
Just look at what's done with social media, advertising, and browser data to get an idea of where things can go.
This was for a project paternship with a university and hospital to improve patient outcomes with some new exploratory tech approaches. A short document that followed some standard study participation format was generated using easily understandable language in about 1-2 pages IIRC (large fonts so it was easily readable by patients with poor eye-sight). Everything done, including the document, went through an external IRB process for human subject data and was approved. Everyone involved had to go through human subject training and what not.
Physician would mention the study to patients that would likely be good subjects for the work about the work, its goals, if they'd be interested in participating. Forms were then provided to patients involved to sign (explicitly) about their agreement to participate in the effort and how their data would be used, protected, etc. The process also required physician sign-off to confirm they read the document to the patient verbally, determined they were competent, cognizant, not under any sort of duress/intoxication, etc. The patient also needed to verbally acknowledge they agreed. Oh, and there was a clause they could retroactively pull out of the work, including their data at any point of they felt uncomfortable or changed their minds.
The patients and their data weren't the product, tech developed that would assist patients was the product of the data. For patients who agreed, some would also be permitted to see some of the products of the work related to their data. Im forgetting a lot of the data collection process because it was very rigorous and several years ago now, but everything above bar, no dark patterny ah-ha-gotcha! line buried in a 300 page liability sign off they had to agree to for some necessary life saving treatment or anything of that nature.
I even got to meet some of the people we helped which was a bit rewarding to see people's lives improve a bit with technology. The specific patient mentioned and their neurosurgeon even let me sit in on their brain surgery tumor removal (patient's suggestion), which was a very unique experience. So yea, they knew what was going on.
With that said, not all data usage was as transparent and ethical as what I worked with, and I saw a lot of mistakes there that make me cringe thinking what a less ethical business with no transparency might do, given the opportunity.
It'll be much more useful for legal and not-so-legal purposes. And last but not least: insurance. As an aside, privacy enforcement for medical data is much easier through fingerprinting tethered to a NDA.
You are correct that if doctors could be more strategic about how they collect info it would make the job of the data scientist much easier. This is especially true in radiology and pathology.
But the real question should be: why doesn't de-identification work and how to make it work. It is a technical problem, and I thought it was more or less solved. But the author here things it is just placebo. If so, what is the problem exactly? Is there a fundamental problem with the very idea of de-identification? Is there a "bug" in the process? What level is required to conduct these attacks? Can in individual do it? A cybercrime gang? A nation state? Is it only theoretical?
Depending on that, the answer could be very different.
https://en.m.wikipedia.org/wiki/Data_re-identification
I suspect that the more data they have for a particular patient (many visits over multiple years) makes the re-identification process easier. The article mention financial data, if dates and amounts of charges aren't masked or altered that could be cross-referenced with another data source to deduce the person's identity.
So it’s the definition of de-id that doesn’t actually “work” in that it’s possible to reidentify small amounts from de-id data, and that adds up over time.
For example, HIPAA considers data de-id if you remove 19 fields or expert determine that it’s de-id. [0]
What’s expert determination, who’s an export? That’s up to me to decide and my lawyers to accept.
The bug is that if half a percent of people each de-id datasets can be reidentified, likely acceptable in HIPAA, then each data released adds up for reidentified people. And more datasets allow for triangulation and linking to reidentify.
The article calls this out as a risk, not a certainty as it’s unclear if anyone is doing this. But the process would be something like: 1) buy HIPAA de-identified data since it doesn’t require patient consent 2) reidentify patients using other data publicly for sale (marketing data, voter registration, etc) 3) new data is not longer HIPAA restricted and fully identified health records can be sold for whatever you like (eg, super targeted drug marketing)
[0] https://www.hhs.gov/hipaa/for-professionals/privacy/special-...
How much health data is required to distinguish a specific individual even without any "identifying" information? Surely there must be many people with unique combinations of conditions and treatments. If that is so then a data set that includes those people could only ever be pseudonymous, not truly anonymous.
What other data sets might then be combined with the health data to match a pseudonymous record with an identifiable person? Payment records? Travel records? Time off work? Insurance claims?
I don't know much about how the US healthcare system works, but maybe before asking how to make de-identification of health data work we should be asking whether it could ever work effectively at all.
The industry of de-identification has a long history in fact.
It is typical for many types of criminals to employ hard-to-identify outfit and appearance. The work of investigators is to overcome this hardship and identify the criminal. The co-evolution of methods on both sides will never end.
The only way to anonymize the data in a foolproof way is to remove so many features that it becomes useless for statistical research.
Modeling-wise, regression, classification, and clustering alike all rely on being able to construct a model that minimizes entropy. But anonymity relies upon maximizing entropy. This fundamental conflict can't really be resolved. You pick some point in the middle of the spectrum that ends up serving neither purpose well.
Noise and obfuscation can be added to the data such that the unique person it identifies doesn't exist. For example changing the dates by some limited random amount. Replacing values with roughly-similar choices (replace engineer with scientist, replace Mexico with el Salvador, etc.) The technical problem is doing that in a way that doesn't harm its research value, never allows the final set of possible similar identities to be below some lower limit, and can't be reversed of course.
When my wife almost died from an ectopic pregnancy, Enfamil was confident enough of the would-be due date based on data pieces together by a broker to FedEx us a box of formula on the due date.
Third parties get a real time feed of insurance claims, hospital admissions, prescriptions and other data. They put the puzzle together and make an assertion of valuable diagnoses. (Pregnancy, diabetes, etc)
They don’t need to be 100% accurate. The collateral damage of reminding someone who lost a child and nearly lost their life doesn’t have a cost to them.
In many cases, abusing such data will also be profitable even if the result estimate is highly noisy.
Just as an example, let's assume that an insurance company knows your given name when you apply and they have statistical data which shows that if your name belongs to one set of names, you are smoking with a probability of 51%. If you have another name, you are smoking with a probability of 49%, because the correlation is noisy. It is still more lucrative for the company to exclude applicants from the "more likely to smoke" names group, because smoking causes cost. You might never have tried a single cigarette in your life but with the information the company has, it is more profitable to not insure you.
Of course the above is an exaggerated example which would not work out quantitatively (the profit they make by insuring you might be higher than the expected value of risk from smoking) but the principle still holds; it is used for credit scoring and similar things.
If you want to avoid insuring a protected class, you might be able to identify them based on those databases, and use advertising to avoid the folks you don’t want. (This is common for student apartments)
As a technical problem, it is more or less understood to be unsolvable.
The core issue is the more information you strip away, the more sanitized the data is, the less useful it is. Once the statistics of your "representative" set are no longer that (even worse, weirdly biased) you are very limited in what you can do. This works ok if you already 100% know the data you need and don't, but otherwise it falls apart pretty quickly.
On the other hand, relatively small numbers of innocuous data points about individuals will identify them uniquely.
These two points are in fundamental contention.
Imagine you have high quality 3D scans of people. You can strip the names off of the files, but anyone can volume render the data to see what they look like (inside and out).
There are some papers on adjusting the data, while hopefully preserving anything medically important, but that's obviously a challenging problem.
It's all so tiring.
I'm not implying insurance companies are doing this out of altruism. If it resulted in a net decrease in profit, the apps wouldn't exist. But it does seem like a beneficial form of price discrimination. It seems like it's only "disgusting shit" if it makes your insurance premiums go up because you're a less safe driver.
Also the factors the app monitors are not necessarily correlated with risk of a crash in the way we'd expect. For example, if you drive the speed limit or slower on most highways in the US, you will be driving slower than the speed of most traffic, and it's more likely that someone will rear-end you.
The poster I responded to is proposing we deny medical science this data, and that we remain remain sicker and more ignorant about health, because otherwise the discoveries medical science makes might get used by insurers who somehow gain access to our purchasing data later (another thing no one here is proposing they be given).
Auto insurance could maybe do it though.
It's hard to blame smoking insurance costs on 'sin' - it kills people and increases costs for the insurance company, possibly more than any other choice people can make. If you drive in drag races and ask for car insurance, don't be surprised if it costs more.
Maybe some zipcodes drink more than others, but they don't need to figure that out because they can just look at illness rates for the zipcode directly. Same for age and tobacco.
After watching the movie "I Care A Lot" (https://www.imdb.com/title/tt9893250/) and reading up the horror stories about forced Guardianship scams in the US (https://www.newyorker.com/magazine/2017/10/09/how-the-elderl...) -- i'd be worried that corporatized Guardianship companies scan medical data en-masse to find victims (with sufficient work and profit motive, it can be de-anon'd)
So, even though each hospital may have a legitimately de-identified dataset in isolation, it is not de-identified when combined with the (also de-identified) data from another hospital. The risk of this attack increases as patients visit more hospitals. We humans are fairly long-lived and tend to move around so it may be substantial. (That being said some hospital systems are quite large like Kaiser Permanente and serve huge areas so visiting multiple hospitals doesn't necessarily create multiple tables.)
[0] https://dataprivacylab.org/dataprivacy/projects/trails/trail...
if we act to shield prescriber data from corps if that would cut down on shady pharma 'bribes' e.g. the opiate crisis and the some other direct to Dr pharma marketing.
Insys and Purdue knew which doctors prescribed insane amounts of pills. And then rewarded them with $. Insys even put it on paper as ROI and at least a few went to jail.
I'm not sure technically how well (or if legally) it would work though so maybe the answer is not at all.
mckesson would still know what pharmacies pills go to and anyone can figure out where a MD works to correlate at least target zip codes/markets. But maybe since we already have this monopolistic distribution setup could prohibit mckesson from disclosing granular shipment data.
MD5 was a reasonable password hash - until it wasn't. SHA1 was - until it wasn't. Etc.
An article like this can arguably prove that there is no reasonable means of de-identification since multiple data sources can be combined. Combine that with the fact that HIPAA puts pretty high limits on the minimum cohort size that can be associated with a unique identifier and low ROI from actually sharing this data.
Many places end up in a position where it's simply not worth the risk to share.
I'm getting my feet wet in the field, but I'm thinking something like the "this face doesn't exist" project but for medical records. I'd love to read a writeup on that!
I don’t feel this is enough to deidentify.
If timestamps or a patient # is associated with records, should be possible to combine with credit card records to discover who someone is.
I wonder if we can request what is shared.
is the price of everything far lower than what is required to sustain the system ?