Imagine you're built on a public cloud in a single Data Centre. That DC has an issue. Your 5 Whys shouldn't end with 'the Data Centre had an issue' It should continue with 'my service was only in one DC.... Why?' and continue from there.
Our service was down.
Why?
Because the platform it runs on went down.
Why only one platform?
Because management doesn't want us to spend several years engineering for multiple platforms.
---
I find it hard to imagine a realistic path of 5 whys that leads us back to a cause within our own team.
Yes and that's fine. It's okay to end up with an RCA like that and note that as a known risk. If a cost-benefit analysis shows that the risk is dwarfed by the solution, then it's not worth fixing. Knowing the existing risks of your system and reevaluating these risks in the face of changes (more people available, inefficient context switching, etc) is a key component of making progress.
However, most real life situations I've seen show that there's always room to improve. For example, how quickly did your service recover once the underlying platform recovered. Rare events (like DC failure) can be a bad thing for operational training. Teams forget how to recover. No runbooks, no training, or tooling that's infrequently used no longer works.
You don't just shruggie and move in "that ship has sailed" at least if you're doing it right.
Our service was down.
Why?
Because the platform it runs on went down.
Why only one platform?
Because management declined to approve purchase orders for sufficient hardware to allow redundancy.
---
And your career just hit a wall. Passing the buck up the chain tends to do that.
If your management is bad, then there are a lot more risks to your career than trying to work with them to improve the process. It is more likely that a given employee has dead-ended their career because they don't know how to explain a problem without blaming people and having the humility to see from other people's point of view; both useful skills for getting a promotion.
Is your superior's superior also incompetent? What about his superior? Does your superior have someone else on his level that isn't incompetent? If you genuinely believe the company is top to bottom incompetent then what are you still doing there? Otherwise identify the competent people and start working that angle, either to get yourself transferred under someone more competent or to move critical decision making power from your superior to someone more competent (perhaps yourself)?
Or, if you're that sort of person, find a way to use their incompetence to mask your lack of giving a fuck and just slack and collect a pay check until they fire you.
[0] https://twitter.com/trevorsumner/status/1106934369158078470
Management only comes into play when causes like "under-funding", "constantly changing requirements" or "too tight deadlines" are identified.
> We should want and expect software engineers to have the agency and bigger-picture thinking to do more than just design for inadequate inputs.
Saying software engineers should have bigger-picture thinking doesn't help if the software engineers aren't empowered by management to make the necessary changes.
For example, "I trusted this dependency which broke" doesn't mean you Did Something Bad, it just means that the cause of the failure was the lack of a way to validate the dependency and fall back on a known-good version. A good example of this is build pipelines relying on the ability to download npm/nuget/debian packages from a remote server instead of vendoring them locally - this is an intrinsically unreliable choice to make. The remote server being unreachable is not your fault, but there are steps you can take to fix this problem in the future. Avoiding external blame in your 5Ys (without ignoring the external factors) encourages you to make local changes.
As always you have to apply reason here and not adhere to dogma, but it's a useful dogma.
For some products a AWS S3 [1] outage might be a influential enough that they add a redundant second storage at another provider, for other it might be "this is too expensive and time consuming to work around for the problems it caused, let's hope Amazon gets their shit together"
Getting back to the question at hand. "5 whys = 1 how". Because if we do the root cause analysis, we will definitely find a solution to our problem. What???? Because there is always a solution to our problem??? It's that same skipping over of the most important detail: you aren't doing the most important problems first because that way there will be space for the lesser important problems. You are doing the most important problem first because then when you run out of time, you've done the most important thing! Similarly, you aren't finding a solution when you are doing root cause analysis, your are finding problems. 5 whys is a tool for helping search the problem space. That is all.
The "5 whys = 1 how" quote is really unfortunate and I think it's taken completely out of context. You don't just do 5 whys once and say, "Oh now we know how to fix our problem". You keep doing the 5 whys over and over and over again so you can identify your problems. And at that point you can start to get a handle on your solutions. "5 whys = 1 how" doesn't mean there is only "1 how" that you need.
And just to sum up, if one of your 5 whys ends up with "An external entity failed us", then you might need to find a way that you aren't relying on that entity. However, run some other 5 whys to help you search the problem space a bit more. You may find another "how" that will help you better.
On the big picture, I agree with GP - different people will come to very different conclusions with the "Why" game - because the world is complex, and many factors come into play when something goes wrong. You don't need to fix all of them, but it does become a game of figuring out the most convenient aspect to fix (and no, another level of "Why" won't fix that).
Somewhat related: I had a 2nd level manager who in meetings kept saying "In an ideal world, ..." and using that as a starting point for requirements, problem solving, etc. At one point, I interrupted him and said "Nope. In an ideal world, I would not need to work for a living. Your ideal is already far from it, so let's not make distinctions between ideal and non-ideal."[1]
It's easy to use the "whys" to come up with answers like "Because X needs to make a living" and "Because the incentives with the manager/department/org/company create a situation where testing is devalued (despite all claims by management to the contrary)" or "Because we cannot retain talent due to our compensation policies" or "Because people in the team are unwilling to learn version control systems newer than cvs."
These are not facetious answers - they should be treated on an equal footing with technical solutions. In my experience, management that does not want to deal with solutions that involve other teams/people/policies are ones where the work has been miserable. It's also been my experience that SW people seem to prefer technical solutions over alternatives, which is tragic. When you study things like negotiation, it's almost inverted. They have their own push for "Why", but the push is usually in the motives direction. The person you're dealing with is presenting a concrete demand that may seem objective, but behind it is almost always a fairly human reason (not looking bad, higher status, etc). And they always emphasize trying to root cause to those and addressing them.
Many here know the saying: In the top tech companies, all problems are social ones. They have the talent to achieve anything. If any piece of SW fails, it's not because they didn't have technically capable people. It's because they failed in the social/organizational domain. So when I hear management say they don't want to address that in the 5 Why's, I see management that is solving the wrong problem.
I took a systems engineering class once. The thing they emphasized throughout is "If every aspect of a product is designed perfectly as its own unit, you'll get a crappy product that will not win in the marketplace." The idea is that a perfect widget in isolation may not integrate well, and there are budgetary issues as well. There has to be a certain amount of compromise to integrate with other widgets in the product, and the teams involved need to talk to one another and have dependencies on one another. The role of the system engineer is to oversee that this is happening.
Insisting that your SW or service must be bulletproof of other services, when both are related and in the same product, is highly suboptimal (as is the other extreme). Often the optimal solution is to say "Our service will work, provided service Y works".
[1] I was on good terms with him so I got away with it.
Seems like 5 Whys in that scenario would lead you to look at things like insufficient test coverage, no code review, code reviewers missed the bug because...
Those paths of inquiry lead to your team changing things under your direct control that could make it less likely for similar bugs to reach production in the future.
For instance, when ext4 gained traction, some users found a number of zeroed-out files after a crash.
Why? Because the programs managing those files updated them by truncating them and then writing them.
Why? Because fopen() has no "replace file" mode, even though the "w" mode seems to do just that as long as the system doesn't crash (it does O_WRONLY|O_CREAT|O_TRUNC).
Why? Because file systems used to not have delayed allocation when fopen() was created. ext4 introduced the concept to the ext family.
Why? Because ext4 was designed to improve performance compared to ext3, and that change does that without breaking POSIX.
Why? Because, even with ext3, the only way to correctly implement "replace file" on POSIX is to write the replacement file on the same directory, fsync() the file, close the file descriptor, fsync() the directory, and rename the file.
At that point, the whys have shifted the blame through userland programmers, POSIX fopen(), ext4, and back to userland programmers. Depending on where you stop, the solution can be:
1. fix all userland programs,
2. change ext4 to force sync upon closing a truncated file's descriptor (which got implemented),
3. upgrade the POSIX standard to have a dedicated "replace file" function call.
The truth is, userland is buggy, and the POSIX standard is arcane, so they both share blame. (But ext4, which was correctly implemented, is the only thing that got patched.)
http://dream.thunk.org/tytso/blog/2009/03/12/delayed-allocat...
Implementing it in ext4 increased the risk of zeroing out a file from “having all files updated in the past 20ms be zeroed out” to “having all files updated in the past 4 seconds be zeroed out” in the event of a system crash.
Quite often when I'm doing 5 Why's I find myself socially engineering Why 3-5 to lead us to an action item we can do something about. There are any number of Why #5's someone can generate but a lot of them do not represent progress the team can get behind.
The fact that this happens with so many different teams and managers makes me wonder if it's me or there's just an obvious pattern of misuse going on with 5 Whys.
This is the big issue. Ideology, bias and self-interest creep in and you can easily end up blaming whatever it is that you personally disagree with.
And that's to not even mention the "unknown unknowns", which by their very nature cannot be factored in to a "5 Whys" process, no matter their real world relevance to the problem at hand
How is this better than "Just So Stories?"