At HBO Max every incident had a full writeup and then real solutions were put in place to make the service more stable.
My team had around 3 incidents in 2 years.
If the cultural expectation is that the on call buzzer will never go off, and that it going off is a Bad Thing, then on call itself isn't a problem.
Or as I was fond of saying "my number one design criteria (for software) is that everyone gets to sleep through the night."
The customers win (stable service) and the engineers win (sleep).
Was there any cost considerations associated with prioritizing the implementation or even limiting the scope of the solution?
That typically didn't happen because engineering reviews had to occur first.
A single command created a new repo, setup ingress/egress configs in AWS, and setup all the boilerplate to handle secrets management, environment configs, and the like.
All that was left to do was business logic.
If the issue impacts tens of millions of customers, then yes, get it fixed right now. Extended outages can be front page news. Too many in a row and people leave the service.
Ideally monitoring catches outages when they first get started and run books have steps to quickly restore service even if a full fix cannot be put into place immediately.
Speaking as a technician whose seen 3 AM at work many a time.
One possible bonus, being on call operating your own software also gives you a solid incentive to not wake yourself up in the morning by writing bad code, and fixing those issues that do arise quickly.
Unfortunately, my software interacts over network with software written by other people; if something goes wrong at 3 AM the users don't know which part caused the problem, so they wake up a random person.
Now I see young engineers from top-tier school working "on call" without complaint. I've found ways to avoid such roles, but it always seemed ridiculous and completely unnecessary in a world where there are software engineers around the globe that could easily work full time support positions.
When you're making real money, they own your ass.
Both salary and hourly gigs have income hanging by a thread with plenty of work, yet only one can get overtime.
Be given responsibility/salary for something (aka hired) by a particularly needy manager/org and be 'undependable'. Read: not at their call. See how it turns out.
The worst/eventual outcome: bye-bye money. Hopefully one has a more reasonable environment. Workers have little on their side.
As someone who does SRE (not AWS, elsewhere)... I would absolutely prefer pay as an hourly rate over salary. I don't like putting in more hours/making less money because Developer Kelly had a bad launch... but I have to, The 9s (and bills) Must Flow.
Fortunately, my current place takes this into account. I don't actually need bonuses or structure change... but the larger trends remain. The employer is buying you, salary opens the time box.
If your product is really so important that it can't be down, hire more engineers and pass the markup to your customers.
I'm glad I live in the country where you have to have 11 hours between end of work and start of work(except for special cases afaik).
If the company cannot afford that, then the product is not that important and can remain broken until the morning.
Even 24h fast food places hire 3 people (each working 8h)!
Manual fixes should never be done in a hurry, and if your system is that fragile, I really wonder about the competency of your senior employees and leadership.
Oncall is only crazy to anyone who also believes it's totally acceptable to have whole services down for hours throughout the night.
To those who understand what it takes to have anything available 24/7, you understand damn well that you need someone to jump on a laptop as soon as an alarm bell rings.
Keep your fancy valley salary (with the ridiculous rent prices attached), and I'll keep my European workers right's protection—including undisturbed sleep after my 8 hours workday.
That is not common in Europe. Generous compensation and additional time off is quite typical for engineers handling on-call burdens.
At one company, I was technically on call 24 hours a day 7 days a week for over ten years. Did I get called that often? No. Did I get called at the worst possible moments? Yes.
And it’s certainly industry-specific. Some doctors have this, firefighters—and software engineers. Contrary to the first two, they usually don’t save lives, but revenue though.
Mostly things went smoothly so that's a pretty nice bonus.
Something ridiculous like that is luckily impossible in (most?) EU countries.
Is the effect unformal harm?
In the EU when you are on call this is a contractual thing.
There is a cost to having on-call. Whether it's in the extra hours you are paying your engineers or other technicians, or sleep deprivation, dwindling motivation and performance, the cost is always there.
In a business, cost is always balanced with the return on that investment.
So it trivially follows that on-call only makes sense where the return is bigger than the investment. If you are having your $100/h engineers become $20/h engineers during the day because of the on-call rotation, and you lose $200 of sales over night when things are down (even your customers are asleep) — you are actually investing that $80/h difference for 8 hours ($640) to recover $200, for a net loss of $440.
Yes, there are cases where it's fully acceptable to simply have your service down for the night. Eg. imagine a service that provides the amount of energy sun is providing for a location (to combine it with solar farm production): is it really that bad if that's down at 2am? Sure, it might be nice to get it back up before the sun is up, but this is just a trivial example where an uptime of ~70% (fluctuates) is perfectly acceptable.
I don't understand your take. Every single time I had a job with an oncall rotation, that oncall was paid. I was paid a bonus for being oncall, I was paid a bonus if during oncalls a pager fired outside of office hours, I was paid a bonus if I was pulled into an incident response outside of my oncall rotation. There was always a cost, and we were paid for it. Being oncall represented loosely a pay bump of around 15%.
If that's not your case then I'm sorry but your problem is not the oncall rotation.
So, hire a team in another time zone.
This is a problem of management not prioritizing the health and wellness of their employees, simple as that.
Even if your app is critical infrastructure (it isn't, and 99% of you shaking your head and saying it is are objectively incorrect), you don't need a software engineer to fix it. You need an SRE. That's completely different.