"Deidentification" seems really murky and imprecise at best.
"Deidentification" seems really murky and imprecise at best.
But I think this is a place where we should apply bounded distrust: there are lots of places where we should distrust Google, but reidentifying people in an explicitly deidentified dataset isn't one of them.
For example, I think a pretty worrying outcome here is that deidentification is imperfect (not surprising), the data is fed into model training, and then the model makes identity-dependent inferences. Since no one tried to break deidentification, it's within what I'd expect from the company. (And, to be clear, is bad.)
On the other hand, intentional reidentification to work around contractual deidentification to "a sell a product to the airlines that offered to keep annoying people like me from purchasing flights" is the kind of thing that would make Google's lawyers terrified, so we should not expect it.
If you look at how this worked with DoubleClick, Fitbit, etc there were initially barriers to linking data but the mechanism for unlinking was updated agreements with people who had ongoing interaction with the continuing entity. That's not the situation with the Spirit data.
The closest I'd expect to see for a "keep annoying people from purchasing flights" situation is not a list of troublemakers but a model that's very good at scoring future communications from customers, and has learned to distinguish profitable vs unprofitable customers. This is well within what I'd expect from companies, and doesn't require any reidentification.
I don't trust google to do anything out of goodness, they'll do anything they can get away with. Same as all the other big tech firms. Once a firm gets too big it stops having morals.
I can ask "Tell me about Person A's experience with airlines" and it will tell me. Or I can ask more generally, "knowing your knowledge, derive a no-fly list", and then it's likely Person A will be on it.
Maybe, a little. Top comment on an old HN stylometry tool:
“Wow. This gives a lot of false positives, but it found all ~10 of my old accounts over the years.”
https://news.ycombinator.com/item?id=33755016There’s a new tool too I saw the other day. And a slightly older one ( https://antirez.com/hnstyle ). Anyway imagine Google’s funding + data + new techniques—shouldn’t be terribly murky these days.