That said, there's also data without PII (anonymized) that is not attached to any specific user, that of course is not subject to this data wipeout as Google wouldn't know what to delete. Such data can be used for general ml models.
Not sure how it's done at Google but anonymization is tricky and it's easier to gather two forms of information from the start, one with PII and one without PII. I think most large companies are doing it like that.
EDIT: Just remembered one thing. At Google, even when building general models data that is rare cannot be used because it could be used to identify specific people. For example: If search query is used by less than X people on a given day then these queries cannot be used for model building.
You can also, of course, read in the data, rewrite it with the deleted portion removed, and write it out again.
(Disclosure: I work at Google, though I don't know anything about how Google handles this)
Edit: Nm, I see you're a Google engineer in this domain.
Maybe it is just just the writer's POV, but I do worry they didn't explicitly say it is deleted as far as Google's storing of such data...
Unless they're audited by a source that can be trusted and have the findings made public, I will not believe it either.
Developers could have:
sql = "Select * from Foo Where FirstName = '" + firstname + "'";
All over their code and no one in "compliance" would be any the wiser.Does Google have multiple "Deletion policies" such that deleting data from i.e. your GCP bucket follows one policy, and the "scrubbing" described in this article follows an entirely different policy? If so, do different deletion policies have different processes and different audit trails such that the end "deleted" state is subjective and controlled by the engineering and managerial oversight of the engineering/leadership team of that given product(s)?
From my (naive) opinion, it must be really, really, hard to for example, retrain every ML model that a now deleted datapoint ever touched. Its hard too to believe that, at some high level in Alphabet's org, there is no motivation to have the positive PR of feature(s) like this, but still at essence not delete the parts of the data trail that significantly drive Google's revenue. Do these datapoints significantly impact Google's revenue?
So with that caveat in mind, let me see what I can help answer.
I'm not entirely sure what nuance you're implying when you say "different deletion policies" - while for instance Cloud might have a different timeline or set of triggers for when and what data is deleted, when it happens "deleted" still generally means "deleted". Some products like GSuite have the ability for administrators to say, disable accounts, which removes them from use but doesn't delete the account, but that's transparent to the domain administrator.
It's definitely nontrivial to track data propagation within large systems, but standardizing infrastructure, having central documentation of data handling plans, and having comprehensive privacy reviews for any new functionality that launches helps keep people on the same page.
Edit: oh, and regarding "retraining every model a data point touched" - the easiest way to do this is to just always be regenerating your models on a frequent basis. If you retrain your models once a day or once a week on a fresh snapshot of your data, they'll only ever be that stale.
No offense to you, but I remember when Amazon was releasing their home devices, and many people rang alarm bells in these forums only to be answered by supposed Amazon employees or friends thereof explaining why these devices couldn't possibly been sending data. Well low and behold, they are sending all sorts of data to Amazon. Were those commenters just lying? Were they misinformed? Were they trying to spread disinformation for whatever reason? Perhaps all three..
Here in America, the gig is up. Everyone, even our grandma's, understands that security and privacy is always going to take a back seat to profit. Always.
So again, no offense to you personally, but everything you are saying must be taken with a ginormous grain of salt.
The only answers are either open source or objective third party auditing, or preferably some combination of both. Words from google employees mean nothing.
https://www.usenix.org/conference/srecon18asia/presentation/...
You can see there's a section on privacy and deleted data as well.
Each team has its own policies, because each product is different: at a bare minimum they might be using different storage systems, but it's very likely that their data pipelines are quite different, too. In any case, each team's targets are at least as strict as any published ones, of course.
That's different from your second question about model training/retraining -- there is an answer to that questions elsewhere (but it is taken seriously and training data is also deleted upon deletion of the source data where it contains PII). I don't know this next part, but I suspect models use such vast amounts of data that any arbitrarily small deletion wouldn't have much impact.
I had a similar concern where I was unsure if Google deleted my data and then anonymized it for helping their models anyways. In my opinion, that would be materially different than deleting all of it.
Oh we're all good then!
The entire burden of proof would be on you, vs. a behemoth corp whose bottom line depends on maximizing data collection and retention.
Go to court on your own dime, prove it, and yes, the penalty might have some teeth. Might.
https://www.law.com/therecorder/2018/12/19/facebook-is-being...
this is just naive
https://www.washingtonpost.com/business/technology/google-en...
But I mean if your position is that the NSA has cracked modern encryption technologies, then I guess you better get off the internet. Whether you use Google or not, you're screwed.
Edit: Stray somewhat-related question just in case you'd know, why is it that if I open YouTube in a private browsing window on a computer with freshly cleared cache, it asks me which of my two Gmail accounts I want to log in with? Is it just IP address based plus some browser fingerprinting?
Besides, for that document to work, the authors have to be trust worthy.. in this case the authors are not even close to trustworthy it might as well have been written by the hamburgler.
I bet it's not free to run, but it's cheaper and easier than elsewhere, because Google's infrastructure is built in-house and mostly integrated. I don't envy other companies that want to do the same.
If you're not willing to put even that must trust in your various cloud providers, you aren't the target market. Surely you aren't sharing this information with Google already, right? So the feature isn't for you.
What is worse:
- losing some data of the 0.1% of users who actually care, or
- the regulatory and PR nightmare that would happen if they don't and it is brought to light?
(Disclosure: I work for Google, though not on anything like this. Speaking only for myself, not the company.)
I understand your skepticism but at least in this case, I think your incentives are aligned with Google's. If we don't delete it (within ~30 days I think?) then I'm pretty sure we'd be in violation of GDPR.
Is that to say it's definitely deleted? I can't say that for sure since there bugs are always possible but at least I'm pretty confident the intention is to delete it.
Google may disassociate your data from your account, but your data lives on.