Google calls Drive data loss "fixed," locks forum threads saying otherwise
arstechnica.com
arstechnica.com
Write a blog post, explain what happened, explain who's affected and to what extent, explain if it can be fixed and what you're doing, and explain what you'll do to make sure it doesn't happen again.
Putting out a supposed hidden fix in the Drive for Desktop client, to see if it can recover files locally (?!) when the entire issue is files disappearing from the cloud, doesn't seem like it makes any sense.
Or if the problem really is solely with files that should have been uploaded but weren't, and nothing from the cloud ever actually got deleted, then explain and justify that as well -- because that's not what people are saying.
I don't understand what's going on at Google. If actual data loss occurred, trying to pretend it didn't happen is never the answer. As the saying goes, "the cover-up is worse than the crime". Why Google is not fully and transparently acknowledging this issue baffles me. The corporate playbook for these types of situations is well known, and it involves being transparent and accountable.
To be honest, I have gotten pretty good support from GSuite, though it was fixing Google's own incompetence. It turns out that if you bought GSuite when prompted while registering your business' domain w/ Google Domains, that was a special GSuite locked out of certain features. Because reasons. It took contacting support to fix this. In a very Googley way, it was clear that many companies had hit this problem, and they'd built tools to fix it. They of course didn't stop the problem, but they did have a well-orchestrated process to fix it...
We are still dealing with stupid Google issues like any time our mail relay sends mail to Gmail servers it's 50/50 whether they decide to block it saying we aren't authenticated properly even though our SPF record is indeed valid.
Tell that to Comcast.
Google has alternatives on all their services except arguably YouTube and search
YouTube is uncontested. Let me throw in a couple of others: for desktop maps, AFAIK Google is still tops (on my phone I use Apple Maps and it's...fine). For free email, AFAIK GMail is still the standard (despite various UI changes over the years that have made it worse).
It survives solely on network effect. I can't wait for a competitor, and several are waiting in the wing.
I use Apple Maps on desktop as well. In my area (Denver) they seem to have about the same number of unique issues, but
1. Mobile Google Maps is so, so terrible in terms of screen-space utilization, look and feel, and also its behavior doing turn-by-turn. I've never been anywhere close to as angry at my iPhone then when Google Maps confused me into a parking lot and the low-quality synthetic voice got nearly a minute behind in micromanaging my way out
2. Desktop Google Maps is a lot less usable if you refuse to give the entirety of Google.com precise location access. The move from maps.google.com -> google.com/maps marked the last time I asked it for directions.
However, I'll also say that sometimes the alternatives to Google aren't very good, though they do of course exist. Anything that requires an Apple device is a no-go for me, for instance, so things like Apple Maps are out. Google Maps for me is a must-have, for instance; I use it for navigating Tokyo on a daily basis and something like OsmAnd isn't going to work here at all (Google Maps finds businesses, tells me when they're open, and tells me exactly how to get there on public transit.) Google Docs is pretty useful for some things too; I don't use it for anything too important (I use LibroOffice at home for that), but for something I want to be able to access from my phone or work computer, it's great. The competition seems to be MS Office 365, and I'm not going to use that: I hate MS and I see no reason to pay for a subscription service here when Google's free offering is fine. Google Calendar is really useful too, and lets me share my calendar with others easily.
Comcast is truly evil and horrible, but since the US loves monopolies and lets companies like Comcast establish local/regional monopolies, people there are stuck with them. In better-run countries like where I now live, this doesn't happen, because there's tons of competition for internet service.
YouTube really is pretty close to a monopoly though; it's not like you can just go to Vimeo and find the same videos.
Supply side economics says this is what we supply, you buy. Thats how it works.
They're not pretending it didn't happen though? As per the article they acknowledged it and published a help center article on it. They named the software versions affected (notifying the affected users seems impossible, since the entire problem was that the data had not been synced). Following the links in the help center article, during the incident they posted in a pinned article in the support forum (multiple times) on how to avoid triggering the bug and how to avoid making it worse.
That's pretty much what you wanted to see except for a blog post with an RCA, no?
> Or if the problem really is solely with files that should have been uploaded but weren't, and nothing from the cloud ever actually got deleted, then explain and justify that as well -- because that's not what people are saying.
So the suggestion is that in addition to the bug that they acknowledged, there's a totally different one that appeared at the same time affecting totally different functionality and with different symptoms, and that they're covering up despite not covering up the other bug? That seems like a complicated explanation when there's an obvious and simpler explanation around.
That's also the kind of thing that's pretty much impossible to prove categorically, let alone communicate the proof in a way that's understandable to the average user. What are you going to say? "We've checked really hard and can't confirm the reports"?
(I mean, I guess it's possible to do it. Collect 100 credible reports of files going missing that can reliably identify the supposedly missing file by name and creation date rather than say that it was probably a .doc file sometime in March. Then do an analysis on e.g. audit logs on what the reality is. How many files were never there at all? How many were explicitly deleted by the user? How many were accidentally uploaded to a user's work account rather than personal account? How many were still in the drive, and the user just couldn't find them? And yes, once you've exhausted all the possibilities, how many disappeared without a trace? Then publish the statistics. But while doing such an investigation privately to make sure whether there is a problem makes sense, publishing the results seems like a stunningly bad PR strategy even if no data was indeed lost.)
From that help center article: "If you're among the small subset of Drive for desktop users on version 84 who experienced issues accessing local files that had yet to be synced to Drive"
Is "issues accessing local files" how anyone would describe deleting a user's local files?
That said, I have been in similar situations with large scale customers. It is hard. Some percentage of customers are pathological, and even after you fix their problem refuse to stop continuing the rumors.
Once it’s fixed, I want all communication forward looking. Some percent of people are flat out insane, incompetent, or just assholes. Sometimes you have to lock the thread in order to stop a conversation about something that is already fixed.
Large scale customer bases are just a different beast. Once you experience it, you know what I mean. That doesn’t mean Google took the right path - only people with a comprehensive perspective can evaluate that, and I’m just some idiot on a forum who knows nothing about the specifics.
> I have been in similar situations with large scale customers. It is hard. Some percentage of customers are pathological, and even after you fix their problem refuse to stop continuing the rumors.
>Once it’s fixed, I want all communication forward looking. Some percent of people are flat out insane, incompetent, or just assholes. Sometimes you have to lock the thread in order to stop a conversation about something that is already fixed.
>Large scale customer bases are just a different beast. Once you experience it, you know what I mean. That doesn’t mean Google took the right path - only people with a comprehensive perspective can evaluate that, and I’m just some idiot on a forum who knows nothing about the specifics.
If I run a business and someone communicates to me that they are pathological and cannot be satisfied unless I invent a Time Machine, I am not going to be particularly concerned about their outcomes. They’re just not worth it - fire your shitty customers, for the sake of your business and employees.
My point is that any time you have >1M customers, you will have many pathological people whom you don’t want to do business with in the first place. The right amount of “firing your customers” is nonzero. Anyone who has worked in a customer service role has experienced this.
You fucked over your customers and some of those who were harmed will be rightly furious with you. The solution here is to try and do good by them, not put your head in the sand and treat the them as a percentage!
No wonder people are no longer willing to assume good faith when having to deal with corporations. It's because of people who think like you.
To me, the primary fact seems to be that Google lost some customers data. We can all agree on that - they should have kept it, but they didn't. They sold a product that, for some number of customers, was defective.
What is their ethical and financial culpability here? To me, if they did their best - if they have industry leading backup/replication technology (which I think they probably do), there really isn't much that CAN be done.
On the customer side - your data has been lost. What should you do in this situation?
My experience leads me to believe that some people, as upset as they are, understand that sometimes shit happens. The other set of these impacted customers do not accept/understand that - they want you to invent a time machine a reverse reality. Barring that, they want your first born plus 10%.
To me, it is perfectly acceptable to tell that second group of people something like this: "I am sorry this happened, despite our planning and efforts. It sucks. We cannot fix it. However, you are also a toxic customer - moving forward you should look to another company to fill those needs."
At the very least, firing those customers will help with your line-employees' quality of life. Yes, shit happens and it sucks - but there are a lot of assholes who only make a situation worse. No matter what you do, you will never make them happy, and trying to make them happy will have great cost.
Those "pathological" customers, I have no problem telling to pound sand.
So it's never OK to intentionally screw customers. But when bad things happen, are people looking for the best available resolution? If not, let them go be jagoffs somewhere else.
Google should definitely refund them some amount of money (both good and bad customers). For the pathological ones, it's a nice way of saying "it's worth paying you to go away".
It sucks to lose your data. It might suck more for the Google employees who lost the data. Have a little empathy for both sides - those who don’t can eat rocks.
...than for people who lost months worth of work because they trusted, in good faith, that the platform Google promotes as being a great place to keep your data safe, would keep their data safe?
Of course. Doing the right thing at the moment is also hard. But that's the right thing. Google is famously under-communicating and opaque, locking a thread is par for the course.
Again, of course, their reputation loss doesn't show on their bottom line. (How would it? They let loose the whole CFO army, and we don't really have the convenience of a randomized trial.) But incidences like these are accumulating the kindling to slowly but surely chip users away from the behemoth.
I also think there's a long tail of Beavis' out there that you need to lock things down to stop the rumors.
Genuinely, I would love an answer from someone that believes in both "never talk to the cops" and "corporations should be open about their fuck-ups" to articulate how they reconcile both concepts. For me they're the same side of the coin, but I'd enjoy to be convince otherwise.
Whereas in business, your public and private statements determine your entire company image.
Statements made by companies in public places cannot “only be held against them”. It’s completely different.
EDIT: found it, not decades but almost 2, he left after working there 18 years.
https://ln.hixie.ch/?start=1700627373&count=1
See the HN comments
Is it a monopolistic behaviour by the book?
Since they struggle to innovate they use the strategy of cost cutting. What means moving development and support (if any) to low cost countries.
Low cost countries mean low cost country standards in both code quality and handling of the disaster after the code breaks.
In general, I have 4 copies of any piece of data I care about: the machine, dropbox, backblaze, and a local external hard drive.
You have zero backups if they're untested.
If they're tested and separate by vendor and/or distance, if they're in an accessible medium or format, then they are different.
I lost a decade of my life data when a main hd failed, and a tb backup drive failed at the same time.
Life is lossy
I also keep my important passwords written down on a piece of paper in a fire safe. This includes my borg and tarsnap keys.
We used to run an ad on reddit that read something like:
Your infra is on Amazon AWS and your backups are on Amazon AWS ... you're doing it wrong ..."
... and we had to stop because it made people angry.
They were quite irate and combative at the very notion that there was any non-zero risk whatsoever* at AWS.
They are not doing it wrong. What’s the threat model you’re saying people are not accounting for? eu-west-1 getting nuked?
“AWS” also isn’t a monolithic entity: for all intents and purposes AWS Backup is a separate vendor from AWS RDS, just with a unified billing and management pane.
I’d rather use that vendor, integrated with my AWS resources and managed with the same access controls, encryption, billing, etc that I use for everything else than ship it off to a random third party and maintain that connection.
Because the risk factor of multiple, isolated and separate AWS teams running different products with different infra having simultaneous large data loss incidents boils down to “nukes”.
So maybe people get irate in the same way as they might do with people who say stuff like “the cloud is a scam, why use it when you can host things on servers in a closet?”
Especially because your credentials to AWS are likely stored somewhere that would also store your credentials to your separate backup vendor -- so why would two cloud vendors provide any more protection than two products within one cloud vendor?
It's clear to me how to model hardware failure, or accidental data loss. It's not clear how to model "hackers gain access to one set of credentials but not another" or "provider closes your account and won't give you your data".
https://docs.aws.amazon.com/AmazonS3/latest/userguide/batch-...
Also while I have heard many horror stories about suspicious account closures for many vendors, AWS at least currently seems to be on top of their game.
I fully agree with your argument, just adding colour to it
At that point, if you're only in a single region, you're stuffed. Networking may be affected and you want to spin up in another region, but RDS APIs fail so you can't copy over your backups, for some reasons AMIs won't copy between regions and the R53 control plane APIs fail too, so even if you could bring up a replacement you can't update DNS anyway...
These can be planned for and mitigations put in place in advance, but that involves similar reasoning as might lead you to decide multi-cloud to be a safer option.
Support just repeats the same things back that I've had too many tries and my account is restricted and I can't get a refund.
It's woeful how bad the support is to get such a simple thing sorted out. Don't miss that email if setting up a developer play account!
(If anyone can fix this my developer play account ID is: 7827257533299144892)
> Hey NoScript peeps, (or other users without Javascript), I dig the way you roll. That's why this page is mostly static and doesn't generate the list dynamically. The only things you're missing are a progressive descent into darkness and a happy face who gets sicker and sicker as you go on. Oh, and there's a total at the bottom, but anyone who uses NoScript can surely count for themselves.
> Rock on!
A lot of the time I go looking for shows or movies it’s no longer offered on the same service if quite literally at all. Many of my liked YouTube videos are now just [deleted].
Not to mention any data you store in the cloud when engineer(s) experience career altering events.
1) Western Digital (MyBook I think was the name?) with built-in HD -- notoriously buggly and unreliable pieces of junk
2) Then just... no NAS as cloud seemed to be the answer
3) Now, Synology NAS enclosures and their custom OS (DSM) exists which is just a dream of ease-of-use
If there is a new era of NAS, I'd say it's single-handedly being enabled by Synology, which is really quite remarkable.
Synology is cool but too limited in what it can do for a homserver, IMO.
So why weren't they able to recover from tape? Is the tape backup more limited than people reported, and this data wasn't backed up? Was it just too difficult and expensive to scan the tapes and decide which was the canonical version of each file?
Charitably, the whole system is on tape and piecemeal restoration is impossible.
This is why professionals practice their restoration protocols regularly.
Not to imply that it's reasonable to treat customers of commodity SaaS as professionals.
I wonder if they might have suffered some invisible data corruption issue in Colossus or whatever they use now, and the effects on Drive just happen to be the most visible. Though presumably whatever broke wasn't part of GCP or we would have noticed by now, right?
Seems much more plausible that there's something wrong with the backend code for google drive (the product).
Lost a half semester of work because a stupid sys admin(s) for the university nuked the primary blob storage (instead of months old records) in an attempt to save $$$.
If it's only in the cloud it's not backed up, its stored.
If its on the cloud and on your local PC it's not backed up, it's copied.
If it's on your local PC and two independent clouds, or on your pc, on your local NAS and on the cloud it is backed up.
3 copies, two different mediums, at least one off site. In most interpretations cloud is considered a different mediu..
Shitty branding. Shitty UI. Shitty functional design. And most of all, shitty attitude.
I had the luxury of telling a Google recruiter thanks, but NOPE. I had friends who are highly competent programmers with FAANG resumes treated like trash by Google interviewers, and another who actually worked there who gave detailed accounts of a toxic culture that pitted peers against each other.
Google sucks. They're coasting on their entrenchment and that's that.
Idk seems way too off to me. Maybe we should have other organizations step in and smack the big G around until the call uncle.
Simply putting varying degrees of closed-source Syncing App in place long term, then expecting them to just keep working perfectly in all circumstances, through all possible future PEBKAC (either by the user or the corporation) - to the extent that they're reliable as a persons long term and only data backup - is a false expectation that the masses have bought into as part of the whole mobile/cloud world hype.
On a separate note, really hope the industry either makes cloud compute cheaper or starts using better priced competitors/on premisis soon, the big 3 clouds are crazy expensive IMO.
Then pick a threshold p value you like and that X% will be a reasonably confident estimate
If other words if you think you have 99.999999999 but you really have 99.9, Nature will correct you rather quickly. If reality persistently lets you get away with it, you can counterfactually deduce you must have more nines.
11 nines is such a big number, you start modelling natural catastrophes, floods, solar flares, very unlikely discrete events as well.
> Designed to provide 99.999999999% durability and 99.99% availability of objects over a given year.
https://docs.aws.amazon.com/AmazonS3/latest/userguide/DataDu...
This isn’t in their SLA but they do post it.
Don't rely on "the cloud" unless you absolutely have to.
Managers would get mad at me for not delivering things on time, when I could say "the cat ate my homework" seems like people sabotaged my career got to the rest of Google.
The worst thing knowing while my code changes were being sabotaged everybody else's seemed to be fine.
You can't make an accusation like that without furnishing some sort of evidence/documentation. Even if you're not trying to post this for internet points, it'd be still worth the effort into collecting the evidence, because you can use it later to convince your coworkers/managers that you're not at fault.
What should I have done take screenshots every 5 minutes?
Your response seems so inept. I’m posting on the internet in case anyone has to deal with something similar there’s a trail of evidence. Downvoting would be the same as agreeing with the Nazis.
I don’t expect people to believe me, just putting my experience into the space.
1. OBS allows you to take 60 screenshots a second (hint: a video is a sequence of images) with zero effort
2. Can't you take a single screenshot when you're done and sending it off for review? Or at the end of each day if you're working on a massive change?
>Your response seems so inept. I’m posting on the internet in case anyone has to deal with something similar there’s a trail of evidence. Downvoting would be the same as agreeing with the Nazis.
A random off topic comment in a random hacker news thread isn't going to help anyone. "My dog ate my homework and btw there's this anonymous commenter on this thread about google drive data loss that says the same thing" is only marginally more credible than "my dog ate my homework". If anything the former makes you look worse. Why were you wasting time surfing hacker news rather than coding? Or at the very least gathering actual evidence that spooky things are happening with the VCS?