We have very strict and very well engineered data retention systems. When we say data is deleted, we mean it. Various levels of automation ensure the data is purged, and all data is tracked meticulously for violations across every datastore.
It's one of those systems I wish we talked about more -- it's a marvel to behold just how much work goes into retention policies and the automation that drives it.
It does NOT say that every copy Google has is deleted. It just says that the user-visible copy is gone.
What about, say, Sensorvault? Or whatever other internal systems I don't know about?
... in response, Google took measures to prevent that category of insider-knowledge attack. With audited builds, zero-trust internal model (xref BeyondCorp), and infrastructure cross-checked by multiple human beings to guard against hardware-level attacks, such an insider attack is infinitesimally probable now.
Ok, how would you know that is true?
How many ways can you think of would there be for your statement to be false?
If someone "higher clearance" than you decided to make you believe the above, but actually retain it somewhere in someway you weren't allowed to see. Are the number of ways more than one? How valuable could deleted data be in the case of blackmail or espionage? Can you actually be confident that someone above or before you didn't write a false delete function?
I'm not implying, I'm suggesting that "things we know to be true" is a smaller list than people think.
You suspect WIPEOUT is real, but can't actually know, and you are inside. Why would I believe for even a second?
Also all the source code (for all systems) is visible to all googlers.
This kind of conspiracy theory is really boring.
"NSA taps into Google, Yahoo clouds, can collect data 'at will,' says Post"
https://www.cnet.com/tech/tech-industry/nsa-taps-into-google...
"National Security Letters"
https://www.eff.org/issues/national-security-letters
There is a reason behind the usual in court phrase of "..tell the truth, the whole truth, and nothing but the truth...". So, if a third party would get copies of the data, it would be true Google deleted it...It just not be "the whole truth".
Not as many as you might think.
The systems at Google may seem incredibly complicated--and they are--but when I worked there, the scenarios where somebody intercepts and exfiltrates data without your knowledge are extreme.
> If someone "higher clearance" than you decided to make you believe the above, but actually retain it somewhere in someway you weren't allowed to see.
The way this data is stored, it is designed so that access to the data is logged and the logs have various alerts / auditing procedures to catch exfiltration attempts. SREs will periodically create user data and try out clever ways of destroying or exfiltrating it to test that these controls work. The Snowden leaks also cast a long shadow over work at Google, and since then, basically, all the traffic and data in storage has been encrypted in ways that make it difficult for state level actors to surreptitiously intercept it. These systems are a bit nightmarish to design, because there are competing legal/compliance reasons why data must be retained or must be purged. For example, certain data must be retained for SOX compliance, data may be flagged as part of an ongoing investigation, data may be selected for deletion for GDPR compliance, etc.
Obviously, it is POSSIBLE that someone is still exfiltrating data, but you have hundreds or thousands of smart engineers who are trying to prevent "insider risk" and "state level actors". People within the company are a big part of the threat model, and agencies like the CIA, Mossad, KGB, etc. are also part of the threat model.
The stack may be complicated, but it's also designed with defense-in-depth to prevent people at lower levels in the stack from subverting controls at higher levels in the stack. For example, people who work on storage systems may be completely unable to decrypt the data that their storage systems contain.
If you're going to get pissy about it, it's obviously true that we are not 100% certain that data is destroyed when we say it is. But this invokes a standard for "knowing" that precludes knowing the truth of any statement which is not an analytic statement.
You don't have to believe, even for a second, if you didn't work with the wipeout systems. That's fine. I'm not trying to convince that wipeout works as intended, because I know that I can't provide the evidence to you.
However, you seem to be arguing that other people don't know that the wipeout systems work--that it's somehow impossible to know.
I can't know, lots of people have opinions, so I should just side with the one (avoid Google) that gives me the highest likelyhood of happiness.
The cloud documentation suggests a deletion period of 180 days; so for cloud data, at least, when it says it is "deleted" it seems to mean it will be [fully] deleted within half-a-year. https://cloud.google.com/docs/security/deletion
At least 4 meaningfully different qualifiers about the situation for entirely separate parts of Google.
I was only suggesting that it might be indicative of the period deleted data ordinarily takes to leave the backup cycle.
More sensitive user data is likely to be handled differently - both for privacy reasons and because it's honestly just not as important to keep hold of it (a user's cloud data gets lost? That's a big deal. A user's location history data gets lost? Meh), so it's unlikely to end up in long-term backup storage.
(to downvoters, that's something that happens and I have a specific and recent case in mind)
Are you saying "Someone carries out an arson attack, they (the attacker) leaks clues to their (the attacker's) identity when gloating about it on social media, and those gloat-posts find their way to law enforcement?"
How does that scenario relate to Google data retention? Google data retention has nothing to do with Twitter policies.
It relates to Google data retention because law enforcement's next move might be to ask Google for geofenced location data from the 72 hours preceding the attack in hopes of confirming the arsonist's identity.
i guess we’ll take your word for it. after all, you have no motivation to lie.
So that includes all online and offline + offsite backup systems, presumably? And hopefully any such data is "de-trained" from all applicable ml models and systems, of course.
I don't know if Google can do de-training (depends on how the training data is generated), but generally if the trained data can't be tagged for removal it also can't be reversed from the output of the training.
Look, I agree Google has tons of flaws, but these types of conspiracy theory "bet they're not going to really delete it" missives don't even make sense if you consider Google's incentives. It would be a gigantic, massive blow to Google if they said they were deleting data and they didn't. There is literally 0 reason for Google to do this. When Snowden's revelations came out, Google was furious and they rearchitected basically all of their systems to encrypt everything at rest and in transit.
I don't think the competency of any organization of people is ever assured and beyond reproach, but you're welcome to disagree.
And while I agree that all organizations can make mistakes, I think it's more important to look at the structure and incentives of any company to see how reasonable those risks are:
1. Google knows they collect a ton of data on people, and they have giant, giant financial incentives to keep that data secure.
2. Google is profitable enough that they can have huge teams focused on data security and integrity (a smaller company may have the same financial incentives to keep data secure, but not the resources to back it up)
3. Google is large and established enough that they can ensure real rigor in their processes (again, as opposed to some rando company that did just enough to pass their SOC 2 audit).
Again, mistakes can be made but we all pretty much take risks every day (banking, getting in a car, getting on a plane, getting in an elevator, etc. etc.), and I think the risk of Google fucking this up is lower than failure of a lot of those other systems.
Because that's what we're talking about here in some instances. You're asking people to bet their life on the fact that a theocratic government won't be able to compel Google to give up location data for the purposes of punishing people for what the state says is murder.
So would you bet your life on it?
You have to opt-in to location history, and it seems to me Google is trying to act in good faith by automatically deleting information that might be sensitive. Who are we asking to bet their lives? People who don't feel like turning off their phone before going to the clinic and have opted-in to location history?
And I just accepted a software update on my Pixel phone a couple of days ago and found that it had reset my default apps for browsing and music playing -- what are the odds that Google will surreptitiously do the same to my location tracking consent?
I think the real risk to my life is about 1000 times worse in a car than some diaphanous threat of a future evil government, and I trust Google about 1000 times more than I trust some Uber (or Lyft or taxi) driver.
They might be deleting the data but are they deleting the metadata? There are always some bread crumbs left.
Source: I've worked on many large storage systems, including the largest (by bytes at rest) at Facebook. They were subject to the exact same kinds of requirements as Google, and as the code continued to evolve gaps would still appear from time to time.