In 25+ years of working in tech, I can honestly say I've never worked anywhere where there haven't been one or more serious issues where one or more parts of the cause was something everyone knew was a bad idea, but that slipped because of time constraints, or a mistaken belief it'd get fixed before it'd come back and bite people.
That's ranged from 5 people startups to 10,000 people companies.
Most of the time customers and people in the company outside of the immediate team only gets a very sanitized version of what happened, so it's easy to assume it doesn't happen very often.
Gitlab doesn't seem like the best ever at operating these services, but they also doesn't look any worse than average to me; which is in itself an achievement, as most of the best companies in this respect tends to be companies with more resources and that have had a lot more time to run into and fix more issues. For a company their age, they seem to be doing fairly well to me.
Also what’s the point of transparency if you’re not getting critical feedback from it and learning?
I know every company makes stupid mistakes, but all of the ones Gitlab made are public, and there’s comparatively few.
Most places with decent devopss hygiene have defense-in-depth around their backups.
I've heard of people dropping production databases in big companies (but saved by backups).
There are some stories around the bitlocker blackmail thing that had similar impact, but that was with a malicious opponent.
The only thing similar I've heard for the notorious self modifying MIT program (for geo-political coding) in the 1990s which destroyed itself without backups.
If a big company lost a ton of user data, I'd absolutely know about it, whether they have Apple-level secrecy or not.
> This incident caused the GitLab.com service to be unavailable for many hours. We also lost some production data that we were eventually unable to recover. Specifically, we lost modifications to database data such as projects, comments, user accounts, issues and snippets, that took place between 17:20 and 00:00 UTC on January 31. Our best estimate is that it affected roughly 5,000 projects, 5,000 comments and 700 new user accounts.
https://about.gitlab.com/blog/2017/02/10/postmortem-of-datab...
Yes, most incidents from most companies don’t result in this kind of data loss, which is why GitLab stood out.
[1] https://github.blog/2010-11-15-today-s-outage/
[2] https://github.blog/2010-07-25-one-million-repositories/
Probably a few hundred TB or so. Maybe nearly a petabyte?
Of course there are people that avoid this, but I've seen very few places where their processes are sufficient to fully protect against it - a lot of people get by more on luck that proper planning. Often these incidents are down to cold hard risk calculations and people know they're taking risks with customer data and have deemed them acceptable.