Say that you are implementing a role-based access control system.
And then ask yourself how can my RBAC system fail?
Well one way it could fail is if someone without the necessary role can operate on something that they ought not have access to.
So then create roles A, B, and C, and create a resource and say that group A can rw, and group B can r this resource. Then create a user and give that user role C. And then try to use this user in that role and try to read the resource. If it can, fail the test. Then another test where you try to write using such a user. If the resource was modified, fail the test.
This is as opposed to only thinking about what you want to happen, in which case you might be writing tests that only ensure that those that should be able to read/write can do so. And of course you want to test that too. But ensuring that those that should not be able to access a resource cannot is the more important thing to be sure of, and also the type of thing that might slip by unnoticed.
A system where those that should be able to access resources cannot, this will be detected in normal operation anyway, by way of users performing their usual tasks. But accessing resources you should not is the most critical and could go by unnoticed for a long time.
And that is why the tests that are written in the backwards thinking fashion are the most important ones.
Thinking about larger systems is much more amenable to this approach. Taking your example if we extend RBAC to an auth/authz service we can say the following.
I want my service to be unreliable. In what ways?
I want it to be occasionally inaccessible or non-responsive, produce results that are non-deterministic and inaccurate.
Taking one of those, how would I produce non-deterministic results? I'd make every operation tied to a PRNG, or be a function of something external to the system I don't have control over. Or maybe it could look non-deterministic if I make my logic dependent on time.
How would flesh out an unreliable function based on time?
I'd offer functionality proposed to tie RBAC to the local clock of the user. This way I'd have to deal with timezones, differences in user localities, relativity, leap years/seconds etc. Functions built on that would likely give the impression of being non-deterministic across many users. (Even assuming correct implementation).
Ok what about being inaccessible vs non-responsive?
The easiest thing to do is just shut the system down of course to maximize this objective, but we want the worst possible system so that's one that may or may not be there and may or may not respond. I could set up my load balancer to include nodes that don't exist, I could run jobs on the service boxes at regular intervals that consumed 100% CPU preventing their responses. I could ...
...
And you can go on and on trying to find all the design and config choices that would make for a truly maddening service. Then you say OK from the product feature level down how can I avoid doing any of those things.
Basically, at the project planning stage, I get everyone together and ask how this shit is going to blown up in our faces. What are all the scenarios that are complete failures.
Then a few weeks before a release, I do the same thing. It's amazing the issues this catches from the cross team approach. For instance, something that the product manager was worried about, they never brought up because they assumed engineering knew about it was covered, but in reality, they had no idea that was important.
Or with two engineering teams, or DevOps etc. Normally we have more than a dozen action items out of these meetings.
I think of it less of reverse the problem and more changing my frame of reference or perspective of the problem.
Also similar in nature is safety engineering. You think about how to make the system fail easily and unexpectedly and you avoid those.
You also try to prioritize the highest risk/lowest cost issues in these. These all involve risk management, which is also critical in investing and is a major reason why Munger has been so successful - he only goes for low risk/high reward plays and bets big on the few opportunities that he gets.
I started with what actions people take that make the bathroom as disgusting (pee everywhere, vomit, shoving food in the walls, etc) and then determining defensive mechanisms from there.
By the way, I think toilets which are attached to the wall instead of to the floor are great because you can easily mop under them.
- Does X ALWAYS cause Y? - Can Y happen WITHOUT X? - Could X cause OTHER things that cause Y?