> You can, of course, write test cases that check your entire database for unintended state changes, but I struggle to find that a reasonable amount of effort.
I think you're overthinking this. I'm not suggesting they should have looked for any unintentional state changes, we both agree that is overkill (until you start doing FP).
> Though for this specific bug, that scenario would've failed too, and might've triggered an extra look at the code.
Yes, exactly. For this specific bug even the simplest test would have failed, which would have caused someone to take a look at what was happening. You are correct that if the bug had been more complicated, such as causing both the proper user AND an additional random user to be charged, then it's unlikely a reasonable level of testing would have caught it.
> The adequate monitoring and quick response they did here is probably a very good trade-off.
We agree that there is a trade-off here, and if sacrificing some correctness is what it takes to win you a much higher velocity then they probably made the correct trade off; nobody died as a result of this bug.
But... surely you see there are some cheap steps they could have taken which would have caught this bug? Not all bugs, but this specific bug.
- Write integration tests for important behaviors, such as charging users!
- Make sure those integration tests run in an environment which closely simulates production.