I had a similar concern where I was unsure if Google deleted my data and then anonymized it for helping their models anyways. In my opinion, that would be materially different than deleting all of it.
Does Google have multiple "Deletion policies" such that deleting data from i.e. your GCP bucket follows one policy, and the "scrubbing" described in this article follows an entirely different policy? If so, do different deletion policies have different processes and different audit trails such that the end "deleted" state is subjective and controlled by the engineering and managerial oversight of the engineering/leadership team of that given product(s)?
From my (naive) opinion, it must be really, really, hard to for example, retrain every ML model that a now deleted datapoint ever touched. Its hard too to believe that, at some high level in Alphabet's org, there is no motivation to have the positive PR of feature(s) like this, but still at essence not delete the parts of the data trail that significantly drive Google's revenue. Do these datapoints significantly impact Google's revenue?
So with that caveat in mind, let me see what I can help answer.
I'm not entirely sure what nuance you're implying when you say "different deletion policies" - while for instance Cloud might have a different timeline or set of triggers for when and what data is deleted, when it happens "deleted" still generally means "deleted". Some products like GSuite have the ability for administrators to say, disable accounts, which removes them from use but doesn't delete the account, but that's transparent to the domain administrator.
It's definitely nontrivial to track data propagation within large systems, but standardizing infrastructure, having central documentation of data handling plans, and having comprehensive privacy reviews for any new functionality that launches helps keep people on the same page.
Edit: oh, and regarding "retraining every model a data point touched" - the easiest way to do this is to just always be regenerating your models on a frequent basis. If you retrain your models once a day or once a week on a fresh snapshot of your data, they'll only ever be that stale.
No offense to you, but I remember when Amazon was releasing their home devices, and many people rang alarm bells in these forums only to be answered by supposed Amazon employees or friends thereof explaining why these devices couldn't possibly been sending data. Well low and behold, they are sending all sorts of data to Amazon. Were those commenters just lying? Were they misinformed? Were they trying to spread disinformation for whatever reason? Perhaps all three..
Here in America, the gig is up. Everyone, even our grandma's, understands that security and privacy is always going to take a back seat to profit. Always.
So again, no offense to you personally, but everything you are saying must be taken with a ginormous grain of salt.
The only answers are either open source or objective third party auditing, or preferably some combination of both. Words from google employees mean nothing.
https://www.usenix.org/conference/srecon18asia/presentation/...
You can see there's a section on privacy and deleted data as well.
Each team has its own policies, because each product is different: at a bare minimum they might be using different storage systems, but it's very likely that their data pipelines are quite different, too. In any case, each team's targets are at least as strict as any published ones, of course.
That's different from your second question about model training/retraining -- there is an answer to that questions elsewhere (but it is taken seriously and training data is also deleted upon deletion of the source data where it contains PII). I don't know this next part, but I suspect models use such vast amounts of data that any arbitrarily small deletion wouldn't have much impact.
Edit: Stray somewhat-related question just in case you'd know, why is it that if I open YouTube in a private browsing window on a computer with freshly cleared cache, it asks me which of my two Gmail accounts I want to log in with? Is it just IP address based plus some browser fingerprinting?
Oh we're all good then!
The entire burden of proof would be on you, vs. a behemoth corp whose bottom line depends on maximizing data collection and retention.
Go to court on your own dime, prove it, and yes, the penalty might have some teeth. Might.
https://www.law.com/therecorder/2018/12/19/facebook-is-being...
Developers could have:
sql = "Select * from Foo Where FirstName = '" + firstname + "'";
All over their code and no one in "compliance" would be any the wiser.this is just naive
https://www.washingtonpost.com/business/technology/google-en...
But I mean if your position is that the NSA has cracked modern encryption technologies, then I guess you better get off the internet. Whether you use Google or not, you're screwed.