E.g. the parent wrote that he fears, he could be identified by his writing style, which is totally plausible. How would you "deidentify" this?
E.g. the parent wrote that he fears, he could be identified by his writing style, which is totally plausible. How would you "deidentify" this?
Of course, takes a lot more effort than not doing proper deindetification in the first place but if they wanted to appear like caring about data privacy they still have enough data points to correlate the sets later on (and/or over time).
Remember, Google = Ads. Their only focus and only care. Their mission statement, rendered accurately, is “Ads ads ads ads. Effective ads. Ads worth paying a lot for. Ads ads ads. Advertising and ads.”
If they choose to be evil in some additional way, (1) remember, they would only do that if in some way it serves their advertising needs — not to offer innovative new black-hat databroker services to airlines, and (2) this little dataset will not need to be re-identified. They’ll just use the 20 years of email and search data they already have on like half the world’s population.
Do I believe they have incentives to do it now? No, as you point it out for their advertisement cash-cow they can already just rely on their own data (GMail, Search, Google Flights) but nothing stops them from the potential later on, the data is now theirs.
Likely its value is just to train LLMs but the funny thing about data is that you can always try to find ways to extract more value out of it. I'd prefer there was no possibility for that without requiring me to trust Google (or any corporation).
In an ideal world my data would be mine to control, not to be traded in deals among 3rd parties, it's valuable and I've spent time generating it so in a sense I've done free work to be extracted by these corpos.