The older I get, the less I'm willing to help disaster scenarios.
"A lack of planning on your behalf, does not constitute an emergency upon mine."
The older I get, the less I'm willing to help disaster scenarios.
"A lack of planning on your behalf, does not constitute an emergency upon mine."
Agreed. Imo, having "on call" rotations is a clear red flag and bad software practice that has somehow made it in the mainstream.
I understand the idea of "mission critical" when you're running software on the Moon or something, but if your Earthbound software has so many potential bugs that you need to have dedicated people being on call every weekend to fix bugs or restart servers, you just built it poorly.
Most software engineers just happen to be bad engineers. On-call rotations are a band-aid for poor planning and development (to be fair, often imposed by tech-delinquent middle managers or executives).
Let’s all plan our emergencies to 8am to 5pm, Monday to Friday. And don’t forget the scheduled lunch break at 1pm!
It just requires spending more money hiring more people.
In the old days operations tended to be very isolated in much the way you are proposing. The problem with this is that stability depends very much on the software, so over time operations folks would be extremely defensive and impose all kinds of constraints on what software could do, and the software engineers would be frustrated that they couldn't do things efficiently. Imagine how firefighters would feel if construction workers had a tendency to randomly leave explosives and gas cans hidden throughout new construction and then waltzed off to the next job while the firefighters had to deal with the consequences.
At the end of the day, devs need to have some skin in the game or it's a recipe for disaster.
In mature industries, there absolutely are plenty of regulations in place to make sure that builders don't make responders' life harder. That doesn't mean that the responders aren't needed, but the fact that the software industry as a whole decided to go all "response is the only thing we need for most things" is evidence that it is not mature.
The nature of software and physical construction is different.
But, as we know, useful software interfaces are difficult to define well and, once they exist, they tend to be the most inflexible part of a fast-changing system. It is always better (though of course more expensive) to control both sides of an interface for this reason.
The "skin in the game" argument elides this fundamental reason and substitutes one that implies all of this is the fault of lazy devs, which isn't (generally) true IME.
Edit: I missed the part where you say the Heroku customer has their own on-call team. IME this is not true. The whole reason to use a PaaS like this is to avoid having an Ops team. Sometimes orgs outgrow their PaaS and keep using it anyway, but this isn't necessary, only a historical artifact. These orgs would likely save more money and get a better result by going with normal IaaS or racking their own servers and, in fact, are probably actively switching to that model.
Oncall rotations are part of defense-in-depth against bugs and unforeseen circumstances: Most of the companies that survive without a formal one only do so by outsourcing this for the most common cases; to Cloudflare, to Amazon, etc. -- if there's an opportunity cost to being down someone needs to be able to pick up the phone when there's an outage or critical issue.
Nobody engineers software like spacecraft companies- that doesn’t mean that no one else gives a shit, it just means that cost constraints are real things.
Also, the demands on spacecraft software are trivial (“move the camera once a month”, “do a correction burn after a planetary encounter”, “watch this sensor and do this if it changes”) compared to a modern web application at a Fortune 500 company.
I have worked at firms of various sizes and there is always a point where things usually become complex at some point. Software is as much or more so about managing the humans than it is the software. This becomes especially true as the firm grows in size. Like all fields there are certainly some individuals that perform better/worse than others but even for the best engineers out there, mistakes happen, edge cases pop up especially as the potential complexity grows. Of course these mistakes can pop up more frequently depending on the imposed deadlines. Deadlines to me are a healthy balancing act between the different parts of the business. Sometimes they are arbitrary but I think in a healthy relationship it helps to have that pushback/friction to figure out how much effort is required.
That was a long way of saying I think its a pretty naive and dismissive view to just hand wave and say this is both due to bad engineers and tech-delinquent middle managers. You are not asking for it either but I think this also comes down to social ability/skills. If your worldview is that most software engineers I can only imagine this shows up in the workplace.
If you have different constraints, you get a different result.
Shipping things has the highest risk of breaking something. I worked with an SRE team responsible for managing incidents for a few months, they told me that ~80% of incidents are caused by bad code being shipped, and I saw that happen as well.
Modern software applications are complex and interconnected. It's pretty easy to unintentionally break something in a different part of the application, or ship subtly bad code because you aren't intimately familiar with that part of the codebase.
When you're in the middle of harvesting and your tractor breaks down, you want it fixed now, not at somebody else's convenience.
Mechanics and tow trucks also form an oncall for broken down cars.
There's oncalls all over the place for tech of all kinds. I think the biggest difference with software is intellectual property - we've made it so nobody else is allowed to fix whatever's broken, so of course, we need oncalls to fix the problems instead of letting customers go to their preferred mechanic
Mechanical watches have been around for roughly 500 years (about 4x-5x tbe time since the first program, depending on how you count), which is a substantial amount of time to iterate on core functionality. Even then, watches until the 1970s (when quartz was introduced) were often imprecise enough to lose 15 minutes/day.
The Voyager probes have both lost several instruments, were built with substantial amounts of redundancy, all to the adjusted for inflation cost of about $3.94B US dollars. Maintenance per year is estimated to be about $5M, including the occasional software update.
Do you think this company had good CI/CD and automated tests? They did not. There was fortunately a lot of monitoring so at least you knew when the service was in a bad state, but absolutely nothing else other than a ticket and an angry customer.
I would much rather have extensive test coverage, very good CI/CD, make sure not to do releases on weekends and holidays, and have a few people whose job is to do the monitoring and escalate to the right people rather than just putting a target on a random engineer's back and hope they can fix things quickly.
[0] https://www.theverge.com/2021/9/27/22696097/hospital-ransomw...
By far, the most likely thing to kill you in a hospital is not the IT system but errors by the doctors and nurses. Or that you're too sick to save no matter what they do.
Everyone being motivated to develop in a way that isn’t resulting in brittle software and breaking and maybe even use boring tech for stability so even if an on call is a real thing, it’s relatively benign. It’s not a bad thing to rely on the genius of smart people who have been at it for decades and are decades ahead in some realizations.
Having software that self reports and logs multiple occurrences of the same errors and escalating errors that start in the app are one great way to stay ahead of issues.
By the time someone reaches out it’s easy to find the error, session, user, and say ok we see it and are on it. Acknowledgement at this depth upfront quite often let’s the customers to say it’s ok take a look on Monday. It’s reassuring. Also easy to forward such an issue to a distributed team.
With that said, there are some types of software that exist to deal with complexity external to itself. External systems (integrations, screen scraping systems, etc.) can break without fault internally, but need to be fixed ASAP.
One might say, "You built it poorly by deciding to build this at all." And sometimes that's true. But I've been in plenty of situations where the options were 1) don't solve the customer problem, don't pass go, don't collect $200. Or, 2) build finicky software that works but needs to be babied... and hire engineers that don't mind doing this type of work.
That last bit is critical. Some engineers like the heroics and drama of it. Some people can't help but run headfirst into the fire... and there's jobs for them there :-)
So lending a hand with emergencies that even could have been avoided is ok to help someone new to it.
Caring via software or infrastructure for users is often about effort as you have described.
"Just throw more humans at it!"
I thought so too until I gigged for a SaaS company that absolutely needed it. No, it was not moon-shot rocket control software, but it was a money printer with literally thousands of clients around the globe paying six figures for the service. The on-call SWE duty was not onerous: about 2 hour shifts every 2 months, and your job was to attempt to handle it and then escalate it to staff if you couldn't. There were so many CI tests on the way to production that nothing ever happened, until it did, and then you needed to make sure that money printer go brrr.
The best on-call rotations are the ones where you’re rarely paged, but when those pages happen they’re for important urgent work where your involvement is vital (even if your role only needs to be channeling the work to the right person). Ideally you’re also compensated for being available, even when nothing happens (my current team gets a day off [in addition to normal PTO] for every week of on-call time).
My experience is that a dedicated team running a follow-the-sun rotation (so the team doesn't need much if any "out of hours" because it's always in hours for at least one of them) actually leads to flakier services and more alerts. Without the visibility of the impact of the flaky services, fixes are lower priority and also more difficult to validate.
In this manner, on-call is not a lack of planning, nor an emergency, but deliberately ensuring we have a plan in place to triage out-of-hours incidents, mitigate those that need mitigating, and leaving anything else to be fixed in-hours.
On the other hand, unpaid on-call is unacceptable.
Depends on your relationship with the person who planned poorly: if you truly don't care, then you run the risk of becoming known as the jerk engineer who isn't a team player. The reality is that employees are generally expected to cover for each other to provide the (paying) customer with consistent support.
If the party who planned poorly was the customer, you have to decide whether they're paying you enough to scramble for them, and whether you want their accolades or their ire. My M.O. is to support my customers through thick and thin (within reason, of course), which tends to get me repeat business and referrals.
Yes, going out of your way for your customers is good for business. Give a shit.
It can be made more or less direct to great resonance.
“If you want my help, plan for it and understand I’m not able to help last minute when there’s no planning”