HNHacker News
TopNewBestAskShowJobs

throwaway012282

12 karma · joined October 17, 2022

submissionscomments
throwaway012282··on How to build software like an SRE
You are 100% correct. It's crucial to understand and manage the failure modes of your system around transient network failures and permanent bottlenecks.

Retries are a must in most systems, but need to be planned, otherwise you DoS your own network or services.

throwaway012282··on How to build software like an SRE
Sorry for the throwaway.

> I’ve been doing this “reliability” stuff for a little while now (~5 years), at companies ranging from about 20 developers to over 2,000

I've been doing it for 20+ years including running critical services in FAANGs.

> Use Docker > Use Kubernetes.

Hard pass. There's a reason why most FAANGs developed their own packages management and deployment system and keep using them. They are simpler, less bloated, easier to debug.

> Deploy everything all the time

On your local testbed, maybe, unless you already know a commit is broken. On production, absolutely not.