It's not your fault
kostaharlan.net
kostaharlan.net
In general, I'm in favor of the approach. I don't think singling people out and bullying or shaming them for their mistakes ever works. I think most well-intentioned engineers will already beat themselves up plenty for making a serious mistake, and they don't need any encouragement to do so. I know I do.
On the other hand, there is a red line. At a place I worked, a DBA was let go after he repeatedly brought production down for 45 minutes to an hour at a time by running intensive queries of his own design for data-gathering, in some cases, after being explicitly told not to do that against the prod database. This was a person whose job description required him to have access to prod.
There were process problems, maybe - being allowed to run whatever queries you want on production under your own authority, sure - but his cavalier attitude towards a production environment was still unacceptable. Process can only help when people are well-intentioned and doing their best; if people are malicious or negligent or just not good at their jobs, adding more process to get around that only makes things worse.
That said, even when there is obvious negligence, having the postmortem process look at the issue with blamelessness is important to build up tooling/changes that could prevent it from happening again. For example, maybe you could revoke individuals having direct access to the production database without multi-party authentication.
That doesn't make sense. The moment that you look back at a postmortem for use in penalizing someone via performance management, the postmortem is no longer blameless.
Additionally, if someone is going up for promotion and uses a number of launches in their packet that all resulted in regressions and didn’t have good rollback plans, I don’t think the committee needs to be blind to that fact.
Surely the first occurrence led to a post-mortem which documented and forbed the practices that became known to be dangerous for production.
That's why there is a hiring and firing process.
Trying to have some sympathy: Was he given an alternative? Or was it a "stop doing that important thing -- I don't know how else to do it, figure it out" situation?
I agree if somebody decides to keep doing the same actions after being told not do to them, because their actions would bring down production, and their actions do bring down production, then they should be held accountable.
> "Hero Programmer" is a derogatory name for a programmer who chooses to fix problems in epic, caffeine-fueled 36-hour coding sessions that frequently just kick the can down the road to the next heroic 36-hour coding blitz. Hero programmers would rather react than plan. Projects with hero programmers working on them often make a lot of progress initially, but never arrive at a stable state of completion
Maybe there are workplaces where people get together to collaborate on a design and then break the design down into tasks and assign those tasks to programmers to implement. Maybe this process is performed until the project is done. Maybe. But I've never seen it. I see people taking responsibility for small and large tasks, and the large ones sometimes involve a single person re-implementing entire systems spread across thousands of files (though not necessarily in "36-hour coding blitzes").
> Projects with hero programmers working on them often make a lot of progress initially, but never arrive at a stable state of completion
No project is ever really finished except the ones nobody cares about. Probably because they stopped being maintained by the hero programmer.
We moved on other projects.
The failure in that case is not having a more senior developer mentor the kid.
During postmortems, we would often decide something like "Chris made an erroneous assumption that the fix introduced no bugs." (That's a classic "oldtimer" mistake, BTW. I make it all the time -I'm a slow learner).
Absolutely no blame would be affixed. It was really important for Chris (that's me) to assume Responsibility for the error, and the team would develop a solution.
This being a Japanese company, of course, said "solution" usually ended up being another punchlist item, like "Perform complete regression tests for even the smallest bug fix release," etc.
I'm not thrilled with people using "hero programmer syndrome," or "bus factor" as an excuse to write naive or deliberately dumbed-down code, though.
Sometimes, a program needs to be maintained by skilled, experienced, well-paid, and motivated people. If a company insists on developing code, using advanced techniques, then turning over maintenance to junior staff, or do a bad job, writing a program, because they want it to be maintained by the absolute cheapest programmers possible, that's a problem.
I've been training a JR lately and I will always say "Here's all the things I did wrong when I did this, so don't do this things"
If I did it, he could easily do it, and if we can all avoid my mistakes, so much the better.
It WAS my fault quite a few times, and I'm ok with that. And luckily, everyone else around here is too. I'd hate to work somewhere that punishes honest mistakes. (there are limits, of course)
I'll tell you what, I always make sure to keep the updated_at column unchanged when I'm doing a manual update query. I also always double check and make sure I run DB changes before merging code, and I do a double check on all my code changes in a PR before I merge it in case I left in something behind.
All of these are rooted in screw ups on my end that my company was very understanding of and made me grow as an engineer. Would it have been better to never have made these mistakes at all? Yes, but that's probably unrealistic. I've seen the hires after me make mistakes and I made sure to let them know of my first prod bug too, the way my boss did for me.
It focuses your mind on what you could do to avoid those situations in the first place.
If it’s your fault prod broke you can fix the process or you can look for a new job where the process is already fixed. Or you can find a role that doesn’t involve pushing code to prod, maybe in R&D, etc.
I say Mike is cutting releases because he’s now the one person I trust on the team to not fuck it up.
If you need it in writing so you can fire me if mike fucks it up, let me know.
Manager mike and I all cut the next release at Mikes workstation with him knowing my ass was on the line if we shipped another debug build.
Mike never shipped another debug build.
If you mean hold yourself responsible for doing the best you can and learning from mistakes, then I fully agree.
The issue is fault can also mean carrying guilt with you and continuing to be blamed for it. This is not helpful once you’ve learned the lessons you needed.
In many large organizations, much of the success comes from the foresight, insight, and hard "work" comes from a few benefitting many. The reality is, it is individuals and not some collective group or "teams".
"Just culture is a concept related to systems thinking which emphasizes that mistakes are generally a product of faulty organizational cultures, rather than solely brought about by the person or persons directly involved."
Reading a bunch of these documents in a row made the failure mode for the corporate culture clear and obvious. The difference between effective and ineffective teams is whether they decided to fix the underlying problem or not. At $firstCompany I asked my bosses to take a dev team and fix enough stuff so that we could run a full automated test suite reliably before every deployment, was told it wasn't a good idea, and left the company shortly afterwards. At $secondCompany they tried things to get that to work, but had major cultural problems with getting individuals on each team to have the discipline to backfill the hardware and actually keep prod and dev-test identical. At $thirdCompany they looked at the same problem of individual discipline and created a team of professional non-developers who were full time focused on controlling software roll-outs so that every new roll-out went to the right level at the right time. We devs gratefully handed over this complex process to that team, and the professionals did a great job.
Be like $thirdCompany: figure out your larger internal pain points and fix those; you actually need to be hit over the head with the data on this to make it clear how your culture and tech stack are failing.
In Good to Great this is an aspect of Confronting the Brutal Facts. https://www.jimcollins.com/concepts/confront-the-brutal-fact...
If you introduce a bug and let's say in part due to insufficient documentation in the code, code reviews that didn't catch it, your change passed all tests and you even made some new tests to exactly test corner cases of the change you were making, it ran fine in beta testing. Then it is still (in part) your fault!
Everybody makes mistakes, and systems should be resilient to mistakes. That doesn't mean nobody is at fault for anything. It means everybody has faults, and if somebody doesn't think they have any faults then why should they think they have anything to learn or improve?
Fault finding should be done. People who wrote the documentation should be given notice and opportunity to improve. People who reviewed the change should as well. And certainly the person who introduced the bug. If people can not cope with accepting the blame for a problem they had responsibility for introducing, they should not be in that line of work.
And that it is normal for humans to make mistakes every once and awhile. If the system can't deal with that, then it's not a robust system.
Can’t help but notice that the people who throw around accusations like this are… still blaming the individual.
and if we’re honest it’s because the engineer’s name is Yorick.
"No, you're wrong, it's Systems!"
"No, it's Great Men!"
"No, it's Systems!"
The author found out there's another way to think about events other than "someone did it". Now he thinks he can apply that thought to everything. Sometimes it's the system, sometimes the person. You can't generalize easily.
A person's experiences and circumstances are an inexorable part of that person. They are part of what makes them great (or not). You simply can't remove the person from their environment.
False dichotomy.
Instead systems should be designed such that you can’t easily take them down due to operator errors.
Most of the tools I write require explicit --nodryrun if they're doing something dangerous or irreversible, and beyond that the really scary ones require some kind of acknowledgement that "THIS WILL TURN 1073 MACHINES OFF, ARE YOU CERTAIN, TYPE 1073 TO CONFIRM YOU WANT TO DISABLE 1073 MACHINES."
I want to meet your principal engineers who can fully understand a command's complete set of side effects solely from a bash snippet telling you how to invoke it!
Like yes, it is the principal's (well really everyone's, but sure insofar as someone is ultimately responsible, we can make the principal engineer responsible) responsibility to ensure processes are safe. But you can't do that by magically "refusing to engage in an unsafe process". Because it's difficult to know a priori if a process is inherently unsafe.
Usually you figure this out when something goes wrong, and then you take that learning from a postmortem/incident report and try to apply it to other similar tools. But no, I wouldn't blame anyone, not even a principal engineer whose job it is to ensure safe systems for accidentally fat fingering a dangerous command that they didn't know about.
They perhaps bear some responsibility for not having previously fixed it, but they bear approximately none for misusing a tool that makes itself easy to misuse.