How is team-member-1 doing?
about.gitlab.com
about.gitlab.com
He thought he was smart. And sometimes, he was.
One day, he overheard a team leader talk to his programmers about the newly-minted database, sitting there in front of them on the table, on a brand new .. amazing .. 640Meg hard drive.
This database had consumed the disk. It had cost the company a cool million dollars to create. It was vital that we backed it up.
So, the new 640 Meg disk was on its way, onto which we'd back the database up. The first thing we'll do, the leader said, is copy the database, sector for sector.
"And only then, will we re-index the database!", he claimed. "Until then, the indexes will remain un-sorted!"
Well, the kid overheard all of this, but only heard "the indexes will remain un-sorted!".
Later that night, this kid thought he'd prove himself.
He re-indexed the database.
He didn't tell anyone.
The next day, a not-so-junior programmer came in, saw the database disk attached to the operator machine, and thought that the backup had been done. For reasons we shall not explain, he disconnected the disk from the operator machine.
The index had not been done.
The database was gone.
The new disk arrived, but nobody could mount the old database disk. Much panic ensued!
Operator logs were consulted. The computer room security cam tapes were spooled.
Oh shit!
Epilogue: I made a lot of money from those kids, writing a tool to recover a corrupted database, whose power had been removed mid re-indexing ..
i hope one day i can tell a story like you. that intro is a work of art. :)
When you make a big mistake, it is easy to place yourself in a mindset where you feel like a disaster even though everyone is accepting. I call it the "disappointing your parents" mindset, because it can feel a lot like people are just being supportive because they love you, and what you did was indeed inexcusable to a certain degree.
The feeling is made somewhat worse when you are an employee, because your livelihood and your future are dependent on how other people perceive you. To that point, I'm really impressed by the fact that Gitlab addressed the fact that this employee was still being promoted, and that the mistake hadn't affected that. In my mind that is as at least as important as all of the rah-rah stuff.
I think at the very least it'll turn the heads of a few employees that are living in fear of being fired for a typo by showing them there are still decent employers left in the field.
I'm sure they've been working on that since the deletion
Detailed posts on that is how you begin to restore confidence.
No one is just going to take their word that "stuff are in place now".
I think it's great that they are being completely transparent about this.
That said, it's true that it's been almost two months and it seems that the some important issues there are still open and don't look especially active.
1. Update PS1 across all hosts to more clearly differentiate between hosts and environments https://gitlab.com/gitlab-com/infrastructure/issues/1094
2. Set PostgreSQL's max_connections to a sane value https://gitlab.com/gitlab-com/infrastructure/issues/1096
3. Move staging to the ARM environment https://gitlab.com/gitlab-com/infrastructure/issues/1100
4. Improve PostgreSQL replication documentation/runbooks https://gitlab.com/gitlab-com/infrastructure/issues/1103
5. Build Streaming Database Backup https://gitlab.com/gitlab-com/infrastructure/issues/1152
6. Assign an owner for data durability https://gitlab.com/gitlab-com/infrastructure/issues/1163
It also brought to my attention how much they've progressed as a platform since I used them last.
That doesn't sound logical to me. The chances of such incidents happening isn't related to how they announce that incident.
A company culture of openly admitting mistakes, knowing that your team-mates will not play the blame game, means that problems will be reported quickly.
In a "conventional" culture, an engineer who makes a mistake (and who will be fired for making that mistake) is incentivised to cover up the mistake and hope that the blame falls elsewhere.
In an open culture, the engineer who makes a mistake is incentivised to immediately raise the alarm.
So while the chances of a mistake happening are the same (we're all human), the chances of it being caught quickly and dealt with quickly are better in an open culture.
...and that's the point I was making in my comment
>while the chances of a mistake happening are the same [at transparent companies like gitlab and at opaque companies like github]
>the chances of it being dealt with quickly are better in an open culture [like gitlab compared to a closed culture like github]
Gitlab.com is a repository hosting product i.e. it's in data storage business. Of course there are other important aspects to the product, like the web interface and such, but they are primarily into data storage. They literally lost data of their users [1] (some of them might have been paid customers too), not because someone accidentally deleted the data (which happens and is understandable), but because they did not have the basic things, that you expect of data storage products, functional or tested.
They have been offering gitlab.com as a SaaS product, since 5 years [2] and yet the very basic thing, like backups, wasn't tested or functional. If that's what increases the faith in a data storage company, then I don't know what to make of it.
[1] https://about.gitlab.com/2017/02/01/gitlab-dot-com-database-...
I had it all loaded up on the test environment, then I'd delete the import, make some changes, and re-run it. Wash, rinse, repeat, trying to track down the issue. As I'm sure you can guess, at one point I executed my scirpt in the wrong window and deleted last nights import from Production.
I immediately told my boss, who was very understanding with a "everyone does this kind of thing at some point" kind of shrug and went over to our DBA's office to ask him to re-load the last snapshot. But it turned out the snapshots had been broken for two weeks and no one had noticed.
And it wasn't a simple issue of re-running the import. After all the orders had been imported, humans had manually assigned orders to trucks and dispatched them in the wee hours of the morning. Now there was no way to know what packages were on what trucks.
I think the DBA ended up buying ApexSQL Log out of pocket to roll back the deletion with the transaction log.
The result was that for several hours the delivery drivers for a national office supply company in a certain state were completely unable to use their handhelds or access their truck's inventory. That was my team-member-1 moment.
I imagine that getting the snapshots working properly would be quite an important factor, otherwise the company would always be a typing error away from chaos.
Our DBA was someone who'd been with the company 15 years. I think most of the blame fell on him, but he didn't lose his job or anything like that. I think the snapshot issue was pretty trivial to fix. The only real problem was that there were no notifications going to anyone that they failed.
The customer was pissed, and it was a new customer at that. They basically lost a day of deliveries out of it. But they kind of had vendor lock-in with us (the office supply company they were delivering for liked us and had basically told them to use us because we already knew how to process their data the way they wanted). Switching wouldn't have been trivial, but we nevertheless fell over ourselves to keep them happy for a few months until it blew over.
Here's what impresses me about Gitlab: They not only say they're committed to honesty and transparency, they actually practice it.
It's easy to see this as some cynical PR move, but to me it's refreshing that they have addressed specifically what happened to the employee who made the error. It makes me believe that they are working very hard on fixing their practices to ensure this kind of failure won't happen again, and I trust they will share (as @syste said in this thread) what those changes are once they have it sorted out.
"Oh, how naive you are!" some may say, to which I respond, "Oh, how cynical and inexperienced you are!" Human failure is inevitable. Designing systems (whether in code or in management practices) that tolerate this inevitable failure is very difficult.
This sort of event can be the catalyst for tearing out what didn't work and creating a much stronger foundation for the future, but only if blame is set aside and honesty is allowed to prevail in the "after action" analysis. Call it PR if you like, but I see a healthy desire to deal with what actually happened and fix it rather than falling into the trap of pointlessly assigning blame.
Consider that it took congressional hearings and someone with the chutzpah of Richard Feynman for NASA to own up to the shuttle explosion. A far worse event, with far worse consequences, but the aftermath of those events and NASA's complete unwillingness to hold itself accountable and deal with reality cost it a lot of credibility.
Good on you, Gitlab.
A few tweets or comments are more than enough to prove your point.
They did something similar with the storage post, which was full of Hackernews opinions.
So, this, or at least post my comment on your blog, too :D
My boss covered my arse. I love that man, and i've never made a serious mistake since, as it's made me risk averse.
Gitlab did the right thing here by owning the situation and making it public.
I was a programmer in my first IT job in 1992 for a large retailer in the UK. I was working on some stock related code for the branches, of which they had thousands. They sold a lot of local goods like books which were only sold in a couple of stores each - think autobiographies of local politicians, local charity calendars, that sort of thing. Problem with a lot of these items was that they were not on the central database. This caused a problem with books especially as you don't pay VAT on books, but if you can't identify the book then the company had to pay it. This makes sense because some books or magazines you DID pay VAT on, because they came with other stuff - think computer magazines with a CD on the front. So my code looked at different databases and historical info to work out the actual VAT portion payable, which was usually nil.
I wrote the code (COBOL, kill me now), the testers tested it, all went OK until when they deployed, on a Friday night. The first I knew was coming in Monday morning. All the ops had been working throughout the weekend as the entire stock status for each branch had been wiped. They had to pull a previous weeks backup from storage, this didn't work as they didn't have the space for both copies to merge so IBM had to motorcycle courier some hardware from Amsterdam, etc etc. As this was a IBM mainframe with batch jobs we also had to stop subsequent jobs in case it made the fuckup worse, so none of the stock/finance stuff could run at all.
The branches were royally fucked on Monday as, without any stock status to know what to order, they got nothing - no newspapers, books, anything. We even made it to the Daily Mail, I think it took at least 3 weeks before ordering was automatic again. Cost the company literally millions in overtime, not being able to sell stuff, consultants and reputational damage - it was big news in the national newspapers.
The root cause? I processed data on a run per-branch. I'd copy the branch data to a separate area, delete the main data, then stream it back. My SQL however deleted the main data for ALL branches. It didn't get picked up in QA as, like me, they only tested with a single branch dataset at a time.
I actually appreciate their attitude towards errors by employees.
Unfortunately, the appearance this spectacle creates is that the same sort of attitude should apply to them as a company, i. e. "Don't fire Gitlab! You've just invested 200MB of data into their education"
It's a very smart method to protect not just team-member-1, but also employee-1.
I've been a bit wishy-washy on GitLab for awhile, but honestly, I'm thinking I might give them a shot sometime soon.
Love it.
Like I said in another post a while back, it's fine being transparent but gitlab just have taken this to an extreme. It's important to be private about certain details and just get real work done and be known for that.
Note: It was neat that much of the community was supportive. I see the article as really a thank you to them.
You do realize that it probably wasn't the engineering team writing this blog post, right?
Here's a script I use at home for this:
#!/bin/bash
tar -cf - ~/Documents/ > /dev/nullgiven a sufficiently complex dependency chain for presented problems, anybody can be a 'team-member-1'.
mistakes happen at all levels.