> Company said “their site was slow” and they didn’t know why. Turns out they had two database clusters: one for production and one for research. The research cluster had 8 instances costing $5,000 per month total. The production cluster had 2 instances costing $500 per month total. The research cluster hadn’t been used in two years. The non-technical company owners had just accepted their system was slow for the past couple years without ever looking into possible fixes because, once again, “the cloud means we never have to manage anything. only agile story point product features matter.”
I'm not sure why we should try to apologize for or further normalize this level of negligence/incompetence? Of course things slip if you're tracking the wrong metrics, and if your approach to cost-management ignores huge actual waste while you make the problem worse by doubling down on hiring newbies, bloating do-nothing middle management or product at the expense of engineering, over-working what seniors you decide to keep, etc.
Fixing honest mistakes is, of course, part of the job. Fixing other people's negligence/incompetence/indifference should not be part of the job, nor compensating for other people's greed when they fail to think through their race-to-the-bottom well enough.
And if shoveling shit actually is the job, then just interview for that. If we're interviewing for 10-20 years of experience and a CS degree, that creates an expectation that the work that needs to be done has some relation to those criteria.