The beauty of distributed systems is that you spread your services across lots of different computers, so that when one of them goes down, they all go down.
But that’s just a distraction. We hear about those outages once in a blue moon, because many rely on them. What we don’t hear about is that any given colo, managed service, or CSP customer’s apps go down on their own, all the time, not because of the colo or CSP.
Such outages are banal, so we forget how much more likely they are, and fail to risk-weight our engineering efforts accordingly.
Many distributed systems try to be reliable by retrying failed nodes' work, but designs aren't always sound.