Could you share which government is doing this? It is very troubling.
Edit: thank you all for good responses, I wasn't aware of the risks of de-anonymization but I will do more research into this.
Could you share which government is doing this? It is very troubling.
Edit: thank you all for good responses, I wasn't aware of the risks of de-anonymization but I will do more research into this.
It is in Germany and the law isn't active yet, but will be very soon. They basically create a central organization to collect data from all insurers. They receive info and are allowed to spread it to third parties in anonymized form.
> If the data is truly anonymized (and that's a big if), would you still have an issue?
Health data cannot be anonymized , since indicators and diagnoses can quickly identify a person. But no, this is my data and I do only share it when I explicitly consent to it.
You can opt out if you make enough money because that allows you to leave state insurance, which is kind of a real problem in our health system. But that wouldn't phase our current health minister, who seems to be an idiot and needs to be removed from office. His performance is bad enough.
Now that you bring it up, I can see how one could identify a person based on that data. We're able to identify users based on things like browser extensions; I see how the concepts map.
The only reason I have to favor this sort of data sharing is that I have a (perhaps naive) hope that it would aid diagnoses and treatment.
"Spain’s National Statistical Institute (INE) plans to track the movements of millions of Spanish citizen’s cellphones to conduct a ‘sociological study’, El Pais and EFE reported. The INE will analyze user’s movements between November 18 to 21, November 24, on December 25 and during July 20 and August 15, using data from the ‘big three’ telecom companies in Spain. Data however will be anonymized before processing, the INE stressed."
What on earth are they going to learn? People go to work, people go shopping?
Armed with that knowledge, they can prepare for the eventuality where the one-off exercise gets turned on permanently.
I can see that as useful - it is targeted and pretty simple to do (although the anti-tracking mechanisms in wifi now makes it harder)
but ... I struggle with the value of such data collected on such a scale. Cell Tower radius is on the order of kilometres - the interesting stuff to do inside cities cannot be discovered on that scale - you need metres or better.
This feels like a mass public holiday migration watch. And you learn that ... people live in different cities to their parents?
As a result cells are naturally much smaller in populated areas where you're most likely to want detailed information. In a city centre the cells may be only a hundred meters across. True, up a mountain the cellular network may have no idea where on the mountain you are, but this isn't a Safety of Life application, it's a survey.
I'm not yet convinced it can be done in a way that can't trivially be reversed by anyone combining it with other datasets.
And if it can't be anonymised, then it probably shouldn't be saleable.
See https://arxiv.org/abs/1809.02201 for the main paper from the Census.
If you had access to anonymised health data from my nation, for example, picking me personally out of the records would be extremely trivial, using just two data points:
+ I have an illness that only 1% of people have.
+ I lost my spleen whilst I was in primary school.
Both of those things are actually a matter of public record, thanks to being mentioned in various local newspapers, so it's reasonable to assume someone somewhere has that data.
I believe the general consensus on health data is that you have the age a person was at an incident, and the nature of an incident, you only need to have two or three incidents in your database to de-anonymise their records.
The only way to combat it is to not provide the detail required for the analysis that is exactly what the above organisations wish to do. No profiles, no fine-grained demographics.
And yet, people with rarer illnesses or events than mine will still stand out, so you also need to eliminate them from the dataset, even though they may well be the ones who could benefit most from this sort of widescale analysis.
For example, if they're tracking all instances of various ailments seen by some doctor in some time period. Suppose you are the only person of a given age and gender with that ailment, so that bucket would one entry for anonymous you. Fine. But then suppose the aggregating party begins to buy and correlate against credit card and location records that place you at the doctor's office in that time frame, bingo, they have a match to tie all three data.
Hell yes. The only time it's acceptable is if they get my informed consent.
Yes!
It's simply not anyone's matter to decide but the affected person what their data is used for, if only because that destroys trust. It is critical to the relationship between doctors and patients that the patient doesn't feel any need to hide information because they think that it might be shared and used against them, even when that fear might not actually be justified by the actual facts of the situation.
Also, that big if of anonymization is an important reason why the decision should be for the individual to make, if it is allowed at all: If you distrust the anonymization, then you should be able to refuse, without any requirement on your part to prove that it is unreliable.
One key one imo is simply ownership. Why do they get to sell it, regardless of who they are. I find the "terms of service" argument to be disingenuous.
I really think we need to completely change the way we do IP, and "data" has effectively become a new type of IP.
The default should be open/public domain, at lwats for anonyzable data. There are lots of reasons for it, economic moral and liberal.
Tesla's data enables better driving software. Medical data lost motives medical tech, etc. If more people had access to it, the technology would improve more.
The conventional argument for patents, is/was (1) to encourage/reward innovation and (2) avoid secrecy. Imo, #2 failed, for the most part.
On innovation/etc... Ultimately data needs volume to be useful for ml applications. Opening all the datasets would also combine them, and improve their usefulness.
After all, the data is more valuable the more bits there are per patient record, for researchers to find correlations - and the more bits there are, the easier it is to de-anonymise. Especially as you never know how many bits of data are already public about a given individual.
And given that they're releasing the data to make a profit, or to advance research, anyone doing this has an incentive to release more bits - and no incentive to release fewer bits.
If applying complete attribution to our data was the norm, for at least some markets places, entities empowered by our data would be much less likely to just "share" it willy nilly.
Dark markets will always exist, but I think attribution of data would incentives certain corporations (health, banking, etc) to act in their users better interests, at least compared to what we have now.
It also reminds me of a recent post-Cambridge Analytica story about FB trying to get patient records, for which they said they "wouldn't get the names of the patients".
Okay, but the only reason Facebook would even get that data is to match it with its users' real names through some of their algorithmic reverse engineering of the data. That data is only useful to FB if it can associate various illnesses with real users on its platforms so that then advertisers can market to those people directly.
Don't fall for the "anonymized data" bullcrap. Google was caught doing something similarly recently, too.