You are 100% correct. It's crucial to understand and manage the failure modes of your system around transient network failures and permanent bottlenecks.
Retries are a must in most systems, but need to be planned, otherwise you DoS your own network or services.