"Deidentification" seems really murky and imprecise at best.
But I think this is a place where we should apply bounded distrust: there are lots of places where we should distrust Google, but reidentifying people in an explicitly deidentified dataset isn't one of them.
For example, I think a pretty worrying outcome here is that deidentification is imperfect (not surprising), the data is fed into model training, and then the model makes identity-dependent inferences. Since no one tried to break deidentification, it's within what I'd expect from the company. (And, to be clear, is bad.)
On the other hand, intentional reidentification to work around contractual deidentification to "a sell a product to the airlines that offered to keep annoying people like me from purchasing flights" is the kind of thing that would make Google's lawyers terrified, so we should not expect it.
If you look at how this worked with DoubleClick, Fitbit, etc there were initially barriers to linking data but the mechanism for unlinking was updated agreements with people who had ongoing interaction with the continuing entity. That's not the situation with the Spirit data.
The closest I'd expect to see for a "keep annoying people from purchasing flights" situation is not a list of troublemakers but a model that's very good at scoring future communications from customers, and has learned to distinguish profitable vs unprofitable customers. This is well within what I'd expect from companies, and doesn't require any reidentification.
I don't trust google to do anything out of goodness, they'll do anything they can get away with. Same as all the other big tech firms. Once a firm gets too big it stops having morals.
I can ask "Tell me about Person A's experience with airlines" and it will tell me. Or I can ask more generally, "knowing your knowledge, derive a no-fly list", and then it's likely Person A will be on it.
Maybe, a little. Top comment on an old HN stylometry tool:
“Wow. This gives a lot of false positives, but it found all ~10 of my old accounts over the years.”
https://news.ycombinator.com/item?id=33755016There’s a new tool too I saw the other day. And a slightly older one ( https://antirez.com/hnstyle ). Anyway imagine Google’s funding + data + new techniques—shouldn’t be terribly murky these days.
E.g. the parent wrote that he fears, he could be identified by his writing style, which is totally plausible. How would you "deidentify" this?
Of course, takes a lot more effort than not doing proper deindetification in the first place but if they wanted to appear like caring about data privacy they still have enough data points to correlate the sets later on (and/or over time).
Remember, Google = Ads. Their only focus and only care. Their mission statement, rendered accurately, is “Ads ads ads ads. Effective ads. Ads worth paying a lot for. Ads ads ads. Advertising and ads.”
If they choose to be evil in some additional way, (1) remember, they would only do that if in some way it serves their advertising needs — not to offer innovative new black-hat databroker services to airlines, and (2) this little dataset will not need to be re-identified. They’ll just use the 20 years of email and search data they already have on like half the world’s population.
Do I believe they have incentives to do it now? No, as you point it out for their advertisement cash-cow they can already just rely on their own data (GMail, Search, Google Flights) but nothing stops them from the potential later on, the data is now theirs.
Likely its value is just to train LLMs but the funny thing about data is that you can always try to find ways to extract more value out of it. I'd prefer there was no possibility for that without requiring me to trust Google (or any corporation).
In an ideal world my data would be mine to control, not to be traded in deals among 3rd parties, it's valuable and I've spent time generating it so in a sense I've done free work to be extracted by these corpos.