On a distributed system the user can only try again if the platform has remained stable, the failure is transient (*) and they have (crucially) have been given the information to retry.
The platform that provides a stable environment for the user to just try again has been built on these principles.
(*) there is one administrator assumes it is within the user’s power to resolve the issue
Later, when users are confused at failures and weird states. >ok now lets build a new system that tries to gather all this information on updates in "weird states" and let users fix them!
simplified example, but nightmare.
Either way, it’s no excuses for shipping slop, which is what you’ve done it your software only works under limited idealised circumstances
TFA is for you
These FTP sessions were running over WANs connecting Pennsylvania, Iowa, and Tennessee.
I ended up writing him an "until curl ftp://...; do echo it failed again; done" loop which calmed that particular issue down.
I don't miss that guy, not even 1%. Good riddance.