Most software engineers just happen to be bad engineers. On-call rotations are a band-aid for poor planning and development (to be fair, often imposed by tech-delinquent middle managers or executives).
Most software engineers just happen to be bad engineers. On-call rotations are a band-aid for poor planning and development (to be fair, often imposed by tech-delinquent middle managers or executives).
I have worked at firms of various sizes and there is always a point where things usually become complex at some point. Software is as much or more so about managing the humans than it is the software. This becomes especially true as the firm grows in size. Like all fields there are certainly some individuals that perform better/worse than others but even for the best engineers out there, mistakes happen, edge cases pop up especially as the potential complexity grows. Of course these mistakes can pop up more frequently depending on the imposed deadlines. Deadlines to me are a healthy balancing act between the different parts of the business. Sometimes they are arbitrary but I think in a healthy relationship it helps to have that pushback/friction to figure out how much effort is required.
That was a long way of saying I think its a pretty naive and dismissive view to just hand wave and say this is both due to bad engineers and tech-delinquent middle managers. You are not asking for it either but I think this also comes down to social ability/skills. If your worldview is that most software engineers I can only imagine this shows up in the workplace.
Nobody engineers software like spacecraft companies- that doesn’t mean that no one else gives a shit, it just means that cost constraints are real things.
Also, the demands on spacecraft software are trivial (“move the camera once a month”, “do a correction burn after a planetary encounter”, “watch this sensor and do this if it changes”) compared to a modern web application at a Fortune 500 company.
Oncall rotations are part of defense-in-depth against bugs and unforeseen circumstances: Most of the companies that survive without a formal one only do so by outsourcing this for the most common cases; to Cloudflare, to Amazon, etc. -- if there's an opportunity cost to being down someone needs to be able to pick up the phone when there's an outage or critical issue.
Let’s all plan our emergencies to 8am to 5pm, Monday to Friday. And don’t forget the scheduled lunch break at 1pm!
It just requires spending more money hiring more people.
In the old days operations tended to be very isolated in much the way you are proposing. The problem with this is that stability depends very much on the software, so over time operations folks would be extremely defensive and impose all kinds of constraints on what software could do, and the software engineers would be frustrated that they couldn't do things efficiently. Imagine how firefighters would feel if construction workers had a tendency to randomly leave explosives and gas cans hidden throughout new construction and then waltzed off to the next job while the firefighters had to deal with the consequences.
At the end of the day, devs need to have some skin in the game or it's a recipe for disaster.
In mature industries, there absolutely are plenty of regulations in place to make sure that builders don't make responders' life harder. That doesn't mean that the responders aren't needed, but the fact that the software industry as a whole decided to go all "response is the only thing we need for most things" is evidence that it is not mature.
The nature of software and physical construction is different.
But, as we know, useful software interfaces are difficult to define well and, once they exist, they tend to be the most inflexible part of a fast-changing system. It is always better (though of course more expensive) to control both sides of an interface for this reason.
The "skin in the game" argument elides this fundamental reason and substitutes one that implies all of this is the fault of lazy devs, which isn't (generally) true IME.
Edit: I missed the part where you say the Heroku customer has their own on-call team. IME this is not true. The whole reason to use a PaaS like this is to avoid having an Ops team. Sometimes orgs outgrow their PaaS and keep using it anyway, but this isn't necessary, only a historical artifact. These orgs would likely save more money and get a better result by going with normal IaaS or racking their own servers and, in fact, are probably actively switching to that model.
When you're in the middle of harvesting and your tractor breaks down, you want it fixed now, not at somebody else's convenience.
Mechanics and tow trucks also form an oncall for broken down cars.
There's oncalls all over the place for tech of all kinds. I think the biggest difference with software is intellectual property - we've made it so nobody else is allowed to fix whatever's broken, so of course, we need oncalls to fix the problems instead of letting customers go to their preferred mechanic
If you have different constraints, you get a different result.
Shipping things has the highest risk of breaking something. I worked with an SRE team responsible for managing incidents for a few months, they told me that ~80% of incidents are caused by bad code being shipped, and I saw that happen as well.
Modern software applications are complex and interconnected. It's pretty easy to unintentionally break something in a different part of the application, or ship subtly bad code because you aren't intimately familiar with that part of the codebase.
Mechanical watches have been around for roughly 500 years (about 4x-5x tbe time since the first program, depending on how you count), which is a substantial amount of time to iterate on core functionality. Even then, watches until the 1970s (when quartz was introduced) were often imprecise enough to lose 15 minutes/day.
The Voyager probes have both lost several instruments, were built with substantial amounts of redundancy, all to the adjusted for inflation cost of about $3.94B US dollars. Maintenance per year is estimated to be about $5M, including the occasional software update.