You can either re-identify the people behind the data or you alter it that strong that it becomes useless for meaningful applications.
You can either re-identify the people behind the data or you alter it that strong that it becomes useless for meaningful applications.
People don't yet have intuition about this. I still don't even know how to articulate it. Here's a stab:
"With enough data collected, you can uniquely identify people by ruling out everyone else."
Mid 2000s, Seisent was helping law enforcement solve cold cases using big data. In layperson's terms, they'd narrow the list of suspects by ruling out everyone who has a solid alibi.
At the time, that was achieved by building profiles of everyone, living and dead, simply by compiling 1600 publicly available datasets. Like court records and mortgages and whatnot.
Today you'd include location tracking, social media, all financial tracking, etc.
It's remarkable that any crime goes unsolved. Like the back log of rape test kits, it's only because no one cares enough to bother to look.
It's a pretty neat idea overall, reminds me of those link-following competitions using Wikipedia where you have to start from a topic and get to another completely unrelated topic solely by recursively chasing links. They both exploit the interconnected nature of our culture, a question (or an article) about a recent TV show providing a clue (or a link) about a pretty unrelated historical character.
The game desperately needs some notion of statistical proximity though, it kept asking whether the character was an actor even after knowing he is a president, it doesn't do _any_ kind of deduction on what it already knows, just mad slash-and-dash till it gets the identifying info.
Practically, this will likely yield the same results.
Remember I'm still struggling with phrasing:
I have no idea how akinator is implemented. 20 questions is traditionally a drill down decision tree. From their blurb, I infer that akinator divines its own decision tree, vs user supplied questions. Clever.
Instead of finding a needle in a haystack, the big data strategy is to progressively remove all the hay until the needle is revealed.
Like the big data equiv of Sherlock Holmes' deduction (logic). "When you have eliminated the impossible, whatever left, however improbable, must be the truth."
Again, am not an academic, philosopher. I don't know what to call this strategy.
It's not logic, reasoning. It's brute force. Test every single possibility to find the matches.
In other words, it's a database query.
Ok, maybe I'm just confusing myself here.
I recently had a crazy notion for losslessly scrambling the sequence as well. Mostly for protecting voter privacy (order in which ballots are cast). One of the major blockers to fully digital voting.
I haven't found any hits using terms like "cryptographic timestamps." Surely I can't be the first.
Sadly, others will continue to push their harebrained ideas.
One thing I learned as an activist is that offense beats defense. Meaning it's easier to promote a correct solution than oppose all the bad solutions.
So if a digital equivalent of the Australian Ballot system (private voting, public counting) exists, I better find it.
There are other concept in research to still use this data in research like differential privacy or using KI to synthesize data according to training data provided. But so far all concepts that tried to alter and then publish a datasets directly for research failed miserably.
GP said sex and year of birth that are usually perfectly fine for deidentification assuming some basic k-anonymity and l-diversity protections.
It’s frustrating when people bring up examples of reidentification out of data that were not properly reidentified in the first place.
Hopefully no one competent would think that having unique records based on sex, dob, and zip code are makes data deidentified. This is usually the case of someone not actually deidentified.
The bigger risk, I think, is when you have some threshold of at least 10 records sharing sex, year of birth and still iding individuals.
Sex and birth year only creates 100 total buckets for people born in the last 50 years. There is no way you could deanonymize that.
As has been said, I don't think there is a way to anonymize data, that is useful and not reversible when combined with other data. Even if you only have 2 buckets.
For example, perhaps you take data on favorite movies and strip away every bit of PII except sex and birth year (as you stated).
Now let’s say I take the stripped data and target some facebook ads to people of a specific sex and birth year and have the ad be for people who love a certain obscure movie, doubling down on a second obscure movie and so on. Eventually I could have a reasonable chance of determining which other movies a unique individual may like based on the stripped data.
Obviously irrelevant for movies, but more relevant for prescription drug uses, sexual preferences etc.
The point is that with enough outside data available, even data stripped of PII can be de-anonimized.