Oncall Compensation for Software Engineers
blog.pragmaticengineer.com
blog.pragmaticengineer.com
It's one thing to be on call where you get called 2-3 times a year, because you're working on a quality system where bugs get fixed more often than they get introduced. Then the pay, if any, is mostly compensation for hurting your social life.
It's another to be on call where you get called 2-3 times a week, because the organisation has decided calling you is cheaper than fixing the underlying problems. In that case, the compensation better be worth messing up your sleep cycle and upsetting your partner.
If that's the case, you need to reconsider if you need a devops/SRE team in the first place, if you need an oncall rotation, or maybe if you need to be more proactive in implementing/releasing/deploying new features as long as you stay within SLO budget. We've had weeks and months where we just looked at our graphs and uptime budget and go "our systems are getting worse, we need to slow down releasing and tighten up the automation", and we've also had months where our load was so light that we'd consider doing large migrations or more daring experiments (for our devs) because those also improve our service and our users' experience.
Not really.
If the company loses $20,000 per minute when the system is down, the system should be well engineered, so it rarely goes down - but it's still worth paying $700/week to have someone available in 10 minutes if it does.
Specifically, see the "Motivation for Error Budgets" section of that article.
If nothing is ever abnormal in your system then yes, your error budget is probably too high. But there is also big space between "nothing is ever abnormal" and "I had to get up at 2AM twice a month".
Right, that's why you usually schedule releases over a Tuesday to Thursday (giving you ample rollout/canary/rollback time). You don't schedule on Monday (timezone) or Friday (weekend).
> But there is also big space between "nothing is ever abnormal" and "I had to get up at 2AM twice a month".
Speaking from Google principles, you will never wake up at 2AM because we don't do overnight oncall. We have a split rotation across the globe so there's always someone within waking hours to take care of pages. The real question is what happens between 6am-9am (one timezone) and 6pm-12am (other timezone, at least for my Ireland/New York split team oncall). Obviously you don't do pushes/releases during sketchy periods like I mentioned, and you usually have a "prod freeze" during holiday period (couple of weeks between christmas and new year), but stuff fails for whatever reasons anyway.
I used to work in large datacenter deployment, we'd have sketchy disks that would fail maybe once or twice a week, we'd get paged for certain machines getting stuck in repairs because our automation would fail under certain assumptions. We'd have machines that would go down and never come back up and our automation wouldn't detect that, etc etc. These are all tricky hardware issues that can be made more robust with software, or you get better hardware (some of our old hardware was REALLY bad and would randomly die with seemingly random errors and it took us months to migrate and decommission it properly), etc. These are all problems that one way or another will surface through your SLO budget and can affect how "daring" you can be during planned migrations and new releases, but it's still stuff you need to take care of even outside of work hours.
So, yes, you don't schedule big stuff outside of work hours, but that's not the whole picture either.
The rest of us have to muddle with questions like, how do we do it if we only have 20 people and they're only in two time zones and only half of those really know how to diagnose and recover a corrupted filesystem? A Google-like approach to error-budget-centric risk management just doesn't fit into that world.
Ultimately, you're resilient for what you prepare for. There's a lot of tradeoffs in spend. I get that that's an example and probably not a real pressing concern, but the point is: You shouldn't have everyone trained in everything. You should have escalation paths for everything non-obvious. You should also train your people better.
Trying to "optimize" these systems to use more of their error budget and save operational cost results in fiascos - There was a multi-day outage (not full outage, but several full days below SLA) on a minor system while I was at Google where it boiled down to "Bad Engineer tries to justify their job and management lets them implement a poor design over the objections of everyone".
That's one problem with a fixed on call rate that some organisations offer. It's a hefty chunk of cash and sounds generous to the engineers. But the cost is already known and sunk up front and not proportional to the amount of call outs so the business sees it as a fixed operational expenditure rather than an appraisal of how fucked things are.
The performance metric quickly becomes how many people you still have on cover who haven't quit to work somewhere else because they are burned out.
For partial on-call weeks, the standby comp was adjusted. For good or bad, it was a literal pager, so we could easily adjust things within the team and file paperwork afterwards, as the NOC just called the pager. The downside of that was that it required being in physical presence to hand the pager over.
#1 and #2 are pretty obvious. These things actually eat into your life because you’re actually working.
#3 is hard to calculate how much that’s worth. If I’m required to be logged into vpn and starting to dig in within let’s say 10 minutes. That means I cannot reasonably leave my house to: eat dinner with my family, go to lowes to grab some pvc because the pipe for my sump pump started leaking, walk my dogs, take my kids outside to teach them how to ride a bike, etc.
I feel I should be compensated for having to be ready to go and not having the freedom to live my life. That in itself is an interruption.
When my spouse was on call for a hospital, they had to sit with their phone in their lap when we went to the movies. I had to be prepared to Uber home because they’d need the car. It’s not fun!
Being on call for a hospital is absolutely a different experience IMO than being on-call for a software system. But at the least when it comes to healthcare having had several folks in my family tree being in healthcare in different functions it’s absolutely clear to me that a big reason for the US having a buckling healthcare system is lack of supply of doctors and nurses at the very least to cover each other and to provide better care per patient.
If the rotation is spread only on 2 or 3 Ops in the team, well, being on-call every other week, even on reliable systems, can really suck (given you must always make yourself available).
Things can get even worse in periods with a lot of PTOs like Summer or Christmas. During these periods, if the team is small, being on-call 2 or 3 weeks in a row is not uncommon.
Someone finally said the quiet part out loud about ‘Devops Engineer’ as a job title. Only a matter of time before we wise up about SRE as well, I suppose.
I’m more focused nowadays on “what problems are you hiring me to solve?” since it feels more and more like the Venn diagram of the three job titles has nearly completely coalesced into a perfect circle.
Difference for me is I’m scrutinizing far more intentionally in job interviews about why an org is hiring for SRE/Devops before accepting any offers. Too often orgs are hiring for this talent and turning them into kitchen sinks for anything and everything the SWEs aren’t doing.
Compliance? Send to Devops.
Upcoming audit and need a pen test done in 3 days? Send to Devops.
Did a bad job prioritizing bug fixes and now shits crashing? Devops.
Etc. once you go through that a few times you start to figure out the right questions to ask in an interview and figure out if you’re about to join a company with Devops practitioners or pretenders.
- Why are you hiring Devops/SRE?
- What is a Devops/SRE going to bring that isn't/can't being done by engineers presently?
- Why isn't it being done presently? What have you tried so far?
- How many other SREs/Devops do you have? When will I get to interview with them (if applicable)
- Who is responsible for platform? Infrastructure? Deployments? How are they involved? When are they involved?
etc. As mentioned in my last comment, a lot of it comes through the baptism of working at a lot of really crummy shops to know the kind of bullshit you don't want to put up with. You gotta deal with some of it no matter where you go, but you sure ain't gotta deal with it all.
This is a lot of boilerplate stuff, sometimes you're lucky and these questions get answered before you can ask them, sometimes they're in the job description. So let me talk about that for a minute.
You really want to take your interviewing to the next step? Learn how to inquisitively, but tactfully challenge what you're reading in job descriptions. The answers I've gotten have been far more revealing than "what will I be doing day to day?" if you ask for more details about a bullet point or two and why those bullet points matter, or who they matter to. That includes, yep, on-call.
Most of my other questions are very probing questions about things in the job description; not necessarily because I'm looking for a specific answer, I want to see how the hiring managers and others describe those topics. Can they actually talk about why they're looking for someone to do x, y and z? Can they have a meaningful dialogue about what those responsibilities mean for the team or are they just parroting back what the job description says, like someone in a zoom call just reading words off a powerpoint slide?
Here's an example:
Job says they want a Devops to come in and also be responsible for security, risk and compliance in the infrastructure? Okay, here's my counter-inquiry about that: if Devops has the responsibility for security, risk and compliance, talk to me about the authority Devops has to recommend or deny certain actions in the platform if it is assessed to be too risky or costly to maintain a compliant and secure posture were we to do it anyway (if you've ever been in that unenviable position, you probably know exactly what I'm getting at with this question).
Interviews are two way streets, and in my thirties with a family where "family time" has no fungible cost, I'm driving very defensively on my side of the street.
Interviewed with dozen of companies over my career - never been able to get a straight or truthful answer to this
Put differently, I think one of the ways somebody goes from the "shiny/trendy and unstable" side to the "boring and stable stuff" side is by experiencing the operational pain of their choices. If the pain falls on others, will they still learn?
Of course, the way you talk about your job makes me wonder if you are already experiencing so many systemic/managerial issues that there the feedback loops are already pretty broken, so this one may not make a ton of practical difference.
Oncall is a scourge not because of the experience of technical problems, but because people already working full time have to arrange their lives outside of work around a second "oncall job". A job which occurs after hours, one out of every X weeks.
A dedicated, pure "Ops" night shift (perhaps in another time zone) would be more humane.
In my experience, this leads to design that pushes problems to outside of working hours.
"We don't need to fix that edge case, just have the off-hours ops team do a manual workaround every now and then."
Or "What does it matter that the deployment is error-prone? We can just schedule it with the off-hours ops team."
I think that depends on the seniority of the individual/team. In my experience, of course one can still learn.
To give you a real example: years ago one of our systems went down on a Sunday morning and our team had no oncall people. The infrastructure team was the one who fixed the issue (don't remember the exact underlaying issue, but it did make clear one aspect of our service we didn't properly: signal handling). Next morning the team wrote down a Jira issue to improve the way we handle signals. Ticket got prioritized very high and was fixed the very same day.
Now, what would have happened if the issue that Sunday morning was due to a bug in the software our team wrote? The same thing. The difference is that infra team would have no clue on how to fix the thing and would have to revert the service to a previous stable version. Would the business be fine with it? In our case, yeah. As a matter of fact, they didn't want to spend the extra money hiring ops people for each team to be on call. You see, if the business really cared, they would immediately have hired a software engineer willing to be on call... They just didn't care that much (and they couldn't force the current team to be oncall because our contracts didn't specify so and the average age in our team was around 35, and nobody wanted to be on call).
But how did the senior engineer learn to handle those situations in the first place?
Personally, if I get paged at 3am due to a bug, I'm going to fix it regardless of what the 'backlog' and 'prioritisation' and 'sprint goals' and 'feature roadmap' and 'product owner' say I should be doing.
But some would say I should not be bypassing the process in that way, and that the feedback loop of external stakeholders making requests to the product owner is more than sufficient.
My company goes with option 3 from the list, “It’s not part of the job outside business hours, but we might still try to reach you during those times.” and it's working fantastically for us.
One out of eight weeks my only job is handle alerts, incidents, and questions from other teams. I and my coworkers dedicate this time to burn down technical debt and to add documentation to the codebase. This keeps error rates quite low on its own but we combine this with a scheduled release cadence (3 times a week) and a reasonably sized QA team that tests the major feature flows of our product for each release (~1 QA per/ 25 engineers) and an ops team whose job is built around being available to rollback to a previous release.
Even though we're a large company with millions of DAU most teams get pinged out of working hours less than once a year. My experience here has really pushed me to the opinion that continuous delivery has been hugely destructive for most of the industry, eroding our pool of experienced engineers who can't be bothered to do on-call.
Work has had three out of hours pages in the last two years, all self resolved within a few minutes.
I don’t know why we need to have a job title treadmill for this; I hate not knowing what your definition of “devops” or “SRE” is when interviewing. (Both as a person who interviews others and is interviewed by others).
Before anyone says it: Sysadmins could code (not to the same level as feature folk), shitty operators pretending to be sysadmins couldn’t.
Or the company wants every team to add time-series DB-backed monitoring to everything now, and you have to use the same tech stack (no matter how good/bad) every other team is using, and you have to add a bunch of stuff even if there's no actionable thing you can do if a number crosses some threshold or what have you, because again, you don't actually own any of your deployment infrastructure. At best it can complement your regular application logs for debugging some issues and noticing trends (good or bad).
It doesn't have to be all bad though. When specialization works well, it works well, and there's at least a minimum level of service you can expect (even if it's not the best) without having to work for it yourself like you would if you owned all that extra stuff.
There is a legal requirement (regulatory, but carrying force of law) for some industries to implement ITSM practices (and similar, don't quote me on specifics) . There is a requirement in those practices that Developers not have access to production, and that Operations have access to the code. That's incredibly wrong. It's misguided in the worst possible way - The point is to make sure the two audit each other, but it requires black box auditing, when you actually want whitebox auditing. (Note that allowbox and denybox are not acceptable substitutes here).
SRE is called SRE because of a difference in those practices. DevOps is an inexpert redevelopment of those practices. Sysadmin practices evolved into both, but what's modernly called Sysadmin is descended from the AD and Exchange people, and have bad practices. You can't walk back the evolution of words, you can fix them through evolution as well, but it's as slow or slower than getting there, because the ecological niche is already "filled"
DevOps itself as a concept was born in nebulous circumstances (“dev-ops days” being where the verbiage comes from but the founder of that conference called the job “agile systems administration; and the concepts espoused by the devops movement being almost exclusively borne out of the “10+ deploys a day” talk from Flickr).
Anyway, SRE is not materially different than Sysadmins except in three dimensions:
1. Hire only programmers, none of those operators who click buttons.
2. Treat reliability as if it is its own feature.
3. Solidify the contract between feature folks and people focusing on reliability.
I’d like Ben Treynor-Sloss to weigh in here as he likely knows best, but that’s the most condense version of what I understood
You’re right about the exchange people, but they too suffered title inflation, the exchange folks used to be called IT technicians.
The people automating AD deployments across sites and managing reliability were sysadmins, and they programmed in the most ugliest of languages to achieve that, autounattended.xml and bat files for days.
The tools are better now, but the work that devops/SRE’s do in most companies today is why sysadmins used to do in 2008-
Yeah, that's a really good point.
> SRE will not be the title for much longer, and DevOps was a bad title from the beginning and a flash in the pan but it was the title from 2010-2015.
I hope the terms become better defined, not less, though. What do you think the next title will be.
IMO "software reliability" shouldn't and can't be thought of as its own feature. Reliability is part of parcel of every feature. You can't make bad software reliable and good software doesn't need the "reliability" tacked on. This mindset (similar to thinking of "quality" as an add-on from the "QA team") seems very problematic to me. At the very least it's an unnatural way of getting there, at worst it just doesn't work. Where I would draw the line though is the reliability of the infrastructure, that's not the domain of the software developer, so it's totally OK to delegate that portion and draw a contract (e.g. in your item #3).
It's also "Treat reliability like a software engineering problem not a process/operations problem".
Though the majority of efforts were around making the initial designs robust and with as few moving parts as possible; sometimes automation efforts caused more outages than the dead simple operational problems.
(see also: split-brain with pacemaker/corosync on replicated databases)
Hmm...opaquebox and transparentbox? clearbox?
DevOps started as an idea that the development team should be responsible for operations. Before this, most dev teams created artifacts that got handed to an ops team to deploy and be on call for. That idea went to corporations that wanted to modernize, but you can't just disappear an entire workforce of admins used to doing things differently. It's a similar situation to where graphic designers started being UX designers. These people didn't magically develop a different set of skills...just a different set of expectations.
SRE-SE and SRE-SWE's are responsible for application code and often embed on application teams to bolster either code or system performance or both.
Please do not take companies bastardizing these practices as truth to what they are. There are companies who do this right and we should champion them above the garbage.
Having also worked at Google, I found this situation ridiculous. Facebook treats oncall as something you're just expected to do on top of everything else you're meant to do. So if you have 50 alerts fire, 40 tasks create, 5 UBNs (UBN = Unblock Now, which should be responded to immediatley and will probably be a SEV) and 3 SEVs, well you just have to do all that and your job.
Google oncalls (IME) tended to be fairly light. You'd often do releases too but there tended to be a lot of automated processes around this (ie building binaries, packaging MPMs, release to staging, release to canary, regression detection, push to production).
Facebook's releases (other than Web) were (again, IME) a dumpster fire.
Web was a special case because of continuous push. Push a commit and automated processes would build the (very large) www binary and handle the push to C1/C2/C3 (these are sort of analogous to internal testing, canary aka 1% and prod). Automated processes would verify a commit by deciding what tests to run. This wasn't explicit and would miss relevant tests for various reasons. This could (and often did) break trunk. This could back up pushes for hours. First thing in the morning it may take as little as 2 hours to push to prod. Later in the day it might take 8+ hours.
Facebook works around this by using conditional code, like... a lot, meaning certain code would only run if you're in right set of GKs (gatekeepers) and QEs (quick experiments). Behaviour would be flipped on by a separate GK/QE push, which is much quicker.
But this means when something of yours breaks (which it often does) you have no idea why. Is it a bad code push? A bad GK/QE push? By you? Or some infra you depend on?
I mention this because you had to deal with this sort of thing oncall a lot.
The problem with not giving oncall compensation is that the burden is never shared equally. The person or persons who do more than their fair share are never going to do it for the money because it is annoying but at least the money is some form of recogniation or, dare I say it, compensation.
Disclaimer: Xoogler, Ex-Facebooker.
For tier 1 oncall (5m response time), for each hour oncall outside of working hours, you are compensated for 40 minutes, which you can either take as time off or at your current pay rate (i.e. you are compensated at 2/3 your usual pay).
For tier 2 oncall (30m response time), the compensation is 20 minutes per hour outside of working hours.
For a tier 1 rotation, the team has a staffing requirement of 12 people, split between two sites. There's a max of 80h oncall, outside of working hours, per person per quarter. Because oncall is split between sites, you are never oncall overnight.
The saving grace is a lot of teams aren’t really doing anything that critical, so the on call is more a formality bc that’s what real teams do. Still pointlessly stressful but less serious.
* For every hour you're available for on-call you get your regular hourly rate.
* For every incident, which involves you working, you receive twice your hourly salary for the duration.
So companies based her tend to have that as standard, though I'm sure some companies would pay more to stand out.
(Or is there a 40 h/week limit or night work rules or something that prevents that?)
You have your team of 20+ developers/sysadmins working Mon-Fri, 9am-5pm, and then you have a single individual responding to on-call events outside those hours.
If you're working 7.5 hours a day, you're on-call for 16.5 hours Mon-Fri, and 24 hours for Sat/Sun. That gives a total of (( 16.55 + 48 ) 10) = €1350.
So yes, you get paid a lot. (Of course the hourly salary there was low to simplify the numbers, but you can see the income for working a week on-call is about the same as working three normal weeks.)
I worked for AWS, and our service was critical for like half of the internet. Very hard oncalls.
Doesn’t this only apply to SRE rotations? The dev teams I know of are definitely oncall overnight.
Tier 2 is still "proper" oncall and we were split between sites. I think devs would be tier 3 oncall instead which has no compensation and no expectation of a certain prompt response time (also SLAs and other metrics might be different since they usually aren't an SRE team). In that case tier 3 rotations wouldn't be handled by SREs and would likely not be split across sites (since dev teams aren't usually).
There may be exceptions, and I've never been in a dev team myself although I interacted with many so I might be wrong, but I can't recall ever seeing a tier 2 oncall SRE team not having a site split and having people 24/7 oncall overnight. Just my personal experience.
Yes, you cannot go on a hike, which is why you get paid. You don't get paid more than you do for your normal time working though.
https://physiciansthrive.com/physician-compensation/on-call-...
Otherwise life as normal. Laptop in the car, 4g hotspot on my phone. Only actually had to fix an issue remotely a few times (once with only my phone in a restaurant after a martial arts class).
So they are paid for being available at a diminished rate, and if their availability is needed, then they are paid overtime.
A 5 minute response time means to respond to the call out and start working on it. If you're on call, you should have a suitable WFH setup and it should be on standby, so 5 minutes is ample time. It doesn't means you have to have it resolved within 5 minutes of being called out, that would be absurd.
Empirically, it's more likely my internet connection gets overloaded and dies on a Friday night than I get paged.
However, according to recent ECJ decisions[1][2][3], "Standby Duty" is not reserved exclusively for when the employee is required to remain on-premises, and it also depends on the degree to which the freedom of the employee is curtailed, specifically stating in one ruling[2]:
> ...
> 32 In the third place, and as regards more specifically periods of stand-by time, it is apparent from the case-law of the Court that a period during which no actual activity is carried out by the worker for the benefit of his or her employer does not necessarily constitute a ‘rest period’ for the application of Directive 2003/88.
> ...
> 36 Second, the Court has held that a period of stand-by time according to a stand-by system must also be classified, in its entirety, as ‘working time’ within the meaning of Directive 2003/88, even if a worker is not required to remain at his or her workplace, where, having regard to the impact, which is objective and very significant, that the constraints imposed on the worker have on the latter’s opportunities to pursue his or her personal and social interests, it differs from a period during which a worker is required simply to be at his or her employer’s disposal inasmuch as it must be possible for the employer to contact him or her (see, to that effect, judgment of 21 February 2018, Matzak, C‑518/15, EU:C:2018:82, paragraphs 63 to 66).
And while I'm very definitely not a lawyer, I think it's possible (likely, even) that having to be at a computer and working within 5 minutes of a page, even at 3AM, would constitute significant constraints on the worker and turn it from "On Call" to "Standby Duty", although the exact implications of that will vary from country to country.
All of that to say that I think that 5 minutes is absolutely bonkers as an expected response time. If I were subject to that, I wouldn't be able to leave my apartment for the duration I was on call - it takes me a lot more than 5 minutes to get to and from the supermarket or even the coffee place just outside. Even taking out the trash could take > 5 minutes (and with no cell reception, due to being underground).
[1] https://home.kpmg/xx/en/home/insights/2021/03/flash-alert-20...
[2] https://curia.europa.eu/juris/document/document.jsf;jsession...
[3] https://eur-lex.europa.eu/legal-content/EN/TXT/HTML/?uri=CEL...
[4] (WARNING: auto-download PDF) https://ec.europa.eu/social/BlobServlet?docId=6474&langId=en
That pretty much implies you cannot leave your home while on call.
I've never _quite_ had that demanding an on call requirement. For me the only "5 min response time" requirement has been to acknowledge the notification (mostly so it doesn't get sent to the escalation on call staff), and the requirement to be "on tools" has never been shorter than 30mins. That means I can at least head to a nearby cafe for breakfast, or go do some grocery shopping, or even head out for lunch somewhere nearby with friends. I meant I couldn't do things like go to movies or concerts or events more than 20-ish minutes from home (unless they were events I could reasonably take a laptop to and assume there'd be somewhere quiet for me to disappear to for as long as it took.)
This is why we have secondaries. If you need to leave your house and expect to not have internet access, you inform your secondary oncaller to cover for you for the time you're not available. You need to go to the store? You need to take a shower? You need to pick up your kid from school? You want to have a lunch break with friends? You want to go for a walk to mentally recover? You ping your secondary and ask them to cover you. That's literally what they are there for.
Secondaries are not your primaries. Secondaries are not supposed to be sitting there waiting for your call to cover them. That's not how escalation works.
Your secondary is someone who's not oncall but is available in case you need help or you become unable to acknowledge pages for a limited amount of time. You get into a car accident? You have a fever? You find yourself in a family emergency? Your secondary should be available to take over (it's not an escalation). I would regularly organize my commute time in the morning with my secondary because I'd have spotty internet (although later on we stopped doing that because my oncall response time was long enough for it to not be a problem), I'd tell them "hey I'll be unavailable between 9:30 and 10:00 am, can you cover me?" and they'd turn on their pager and take over the oncall duties while I commuted.
For people with stricter oncall response times (like google ads or google search SRE), you'd often communicate/coordinate with your secondary for everyday things like "going to the store" or "taking a shower". My friends in search-sre would just tell their secondary "Hey I'm planning to take a shower, can you cover me?" and they'd turn on their pager.
Maybe other companies do it differently, but that's how Google does it.
Most of us work at places that don a lot of things differently to Google I suspect.
There's a _huge_ difference between how on-call works in a dozen or so person startup, and hundred or two person single timezone business, and a thousands of engineers across almost all timezones.
I _dream_ of working at a place that has follow-the-sun teams of SDEs and SRDs across 3 or 4 timezones. I have not yet worked at a place large enough to have on call secondaries, I've only worked places where the only on call escalation is that the on call person's manager gets paged (and angry) if the on call person hasn't responded within the SLA. (And I've been both the on call person and the manager in that scenario in several different organisations...)
If it is important for the application to be up 24/7, the company needs to pay for it at the usual rate!
I’m not expected to “pull” responsibilities from a chat room; pages are pushed to me. If someone needs to get ahold of me they are supposed to page me, not message me.
Edit: that being said, oncall outside of business hours does limit my activities, such as hiking, biking, camping, traveling, and I would of course appreciate 100% time or time and a half.
No, it requires that you be able to stop whatever it is you're doing and be working on a problem within some latency tolerance (5m and 20m are cited upthread, for example). For most modern datacenter workloads, that can be as simple as "carry your laptop and stay within reliable coverage". While sure, that rules out a lot of activites, most of our lives are spent in that regime already.
There are plenty of things I do in my own time that can be interrupted: books, movies, HN, housework, etcetera.
Software developers aren't special, were not even operations staff. Just build stuff that fails gracefully and deal with it on Monday.
And be sober, and somewhere quiet enough you'll reliably hear/feel the notification, and be somewhere you can get that laptop out and type away at it for a while.
I get _much_ less enjoyment from many social activities when I'm on call. I enjoy gigs way less. I enjoy parties way less. I pretty much wont go to movies. I hate being "on call" while out at dinner with friends. I will not go on a date while on call (at least not with somebody I don't know well enough for them to understand the on call obligations).
> most of our lives are spent in that regime already
But not all hours in my life are of equal "value" to me. A lot of the "most valuable and enjoyable times" get disproportionally affected by being on call. I care way less about potentially being paged at 2:30am on a Tuesday morning than I do about having to curtail social events on a Friday evening or a weekend. You _will_ need to pay me handsomely to do that, and guarantee it only rarely becomes my responsibility. Been there, done that, am perfectly happy to turn down job offers that don't understand that (or to walk away from companies who try and spring that on me later I've accepted).
Technically this is not a requirement. I've definitely known people going oncall while tipsy or at the pub, as long as you're not shitfaced drunk and physically unable to answer the page. Not that it's something I'd ever do or recommend doing, but it's technically not forbidden.
Every time I was oncall during weekends or holidays (or outside work hours) it was just a normal day with the occasional "phone call". As long as I had my laptop with me and I had some kind of network coverage (which I did unless I went trekking in the non-existing mountains of Ireland, which I didn't during oncall days) it was fine.
My coworkers in search or ads were a bit more stressed out on that though, I agree, having to ask their secondary to cover just for the 5-10 minutes they wanted to take a shower because they could not miss a single alert, but for us on a secondary service that was not a problem. I've had days where I commuted by train (40 minutes ride) with spotty internet and 0 problems because having a 15-30 minutes response time meant that I had enough buffer to get off the train and find some place with wifi with plenty of time to spare.
> I definitely consider every hour of the day I'm on call (all 24 of them) as a working hour
You'd be incorrect. We also don't do 24 hours shifts.
And not every tech company has Google's on-call policy. The company I work for has team-defined shifts, generally these are one or two week rotations where the person on call is on call 24/7 during their rotation.
"being available" is not in any definition of labor I've ever read. Reading a piece of fiction on my couch is not labor under any reasonable definition, because I am not working.
Like if the trade off is Google's policy (2/3 time but freedom) or time and a half but you have to actually work the full weekend and you're expected to write code when not responding to incidents, which do you pick?
If you're a firefighter, is it labor to be at the station playing cards, just because there aren't any calls coming in right now? If you're an ER physician, is it not labor to be waiting for patients on a quiet night?
> Reading a piece of fiction on my couch is not labor under any reasonable definition, because I am not working.
If it's a Saturday and being on-call is preventing you from buying groceries or going to the movies, then being on your couch reading a piece of fiction is labor. If it's the Fourth of July and being on-call is preventing you from having a beer at the barbecue, then that's labor.
"Labor" isn't just the activities for which you are actively producing value for somebody else. Labor is any time your allowed options are restricted as a result of your employer's decisions. Sometimes, those restrictions dictate only a single option of being on-site working on a specific task. Sometimes, those restrictions allow multiple options have some flexibility to them, but the existence of those restrictions at all means that it is still labor being required of you.
Like I said, this is an abnormal definition of labor. It would mean, for example, that I am laboring 24/7, because there are some thing that my employment agreement does not allow me to ever do.
If you'd like to work under that definition of labor, that's fine, but then you cannot square it with an hourly-wage based definition of compensation for labor, so "time and a half for additional hour beyond 40" makes no sense in such a context.
I fully support people being compensated for such inconvenience. I don't think it makes sense to expect a greater-than-normal-work-time compensation for a lesser-than-normal-work-time inconvenience.
The relative difference may very well not exist between the two scenarios. If I can't just go to the beach with my wife, if I can't go walk the dog in the farther-away park, if I can't play an online game that lasts over 40 minutes per match, if I can't schedule a music lesson - then if the above is my definition of free time, then it's going to be difficult to convince me that there's a difference between "you can't do this because you're working" and "you can't do this because you're on-call". All it takes is that I take "can't do it" seriously enough.
If your typical day-off is filled with "short" activities, if you being on-call doesn't affect plans of other people close to you, if you plan your month so that you do all the housework & chores on your on-call days, then you'll probably be OK and will testify to the huge difference between the two.
The perception of this difference will thus vary from person to person, from circumstances to circumstances, from lifestyle to lifestyle.
But there *is* a difference, and that difference is exactly why you're paid 2/3 of your normal rate instead of 100% (or 150% as some people are saying). You aren't working, but you aren't entirely free either, so you are compensated for that by being paid something that is not quite your full rate. *OR* (at least by Google guidelines) you can accrue enough time to be able to fully take a day off later to make up for that time lost.
By the way depending on the day, requirements, oncall shift, style, etc you can definitely relax, play games, go to the beach, etc. Just because you are oncall it doesn't mean you can't categorically do any of those activities (unlike if you were *actually* working), it just means that you need to have a laptop nearby with internet access and temporarily drop whatever you are doing to be able to deal with an outage if it happens. For this reason, the company pays you, but it's not a full rate.
At some point I also romaticized the idea of being on the beach enjoying myself when the pager goes off. So I jump into a terminal, get the adrenaline rush, fix the problem, save the day, and carry on. That narrative just doesn't work for me anymore.
Hiking? Nope. Driving through dead zones? Nope. Going to the movies? Not really. Bike riding? Maybe, if you can hear your phone, haul around a heavy laptop, and stick to areas with phone reception.
Being on call is work. Call it labor or don’t, I don’t care about the semantics. Work is work.
Or do you, maybe, get paid a general smoothened out curve based on the average for your work expectations over a certain period of time?
You get bonuses, raises, promotions based on how well you perform your job as your salary gets adjusted (ideally at least) according to that (+ end of year bonuses, stock/options, etc). This all also contributes to your total compensation including oncall (which is based on your normal work rates).
Usually how it worked in my team at least, if someone had a tougher-than-usual shift (lots of alerts, large scale incidents, etc) we'd get some extra "rest time" (unofficially) or we'd be told to just take some time off in lieu, etc (on top of your oncall pay already) at discretion of your manager. On the other hand if your team's oncall stats (pager alerts, SLOs metrics, etc) were bad over a long period of time with a lowering trend, you'd have to restructure the way you approach/monitor your system and deal with releases and change management practices because something clearly isn't working. This is all encoded in the principles[0] of what it means to be a good SRE and design good systems and is already taken in consideration as part of your stipend.
you know what this sounds like, a wonderful opportunity to exercise some privilege at the managers discretion.
Same as when I’m stuck on a bit of code and I’m looking through the window or taking a walk to think the problem through: I’m working and get paid for it.
Why should being on call and it’s mental + physical (being sober, within arm reach of your computer) burdens be any different?
Correct
> you’re working.
Incorrect
Regardless, people *are* getting paid for their oncall availability. Just not at a full rate (or 150% which would be even larger).
Useless expectation as it still steals your freedom.
I have many friends "in the trades." These are usually unionized, and the compensation can be jaw-dropping. Many of my friends deliberately try to get overtime, which can include "on-call."
But the work can be tough, and the salaries -although good- are usually less than most SWEs.
On-call #1. Averaged one page every week or two. Pages typically resulted from a failed automated process, which was scheduled to avoid running in the wee hours of the night. Most pages could be handled remotely, in about 10-15 minutes. For issues that required going in, even if the root cause couldn't be determined, the system could be brought to a safe state with further troubleshooting done the next day. If there was an overnight issue, you were not expected in until the afternoon, and supervisors would actively tell you to go home and sleep if you were there in the morning.
On-call #2. Averaged 3-5 pages per day. Pages occurred at random times during the day or night, with little to no predictability. All pages could be handled remotely, but typically took 1-3 hours to resolve. Issues frequently required creative problem solving, which was expected regardless of time of day. If there was an overnight issue, you were still expected to be on-site for the daily 7:30 AM meeting.
There was a drastic different in quality of life between the two on-calls. The first was as you say, an on call with the expectation that most of the time nothing will go wrong. The second would be more accurately described as a "working nights and weekends rotation", rather than an "on call rotation"
With the pay structure described above I assume this is applied outside your normal working hours, where you're not doing anything other than being on call.
Someone has conned you into accepting less. I'm sorry.
I agree with you fully, on call time should be compensated at the usual rates, including overtime.
Why would I not charge less for this than real work? It involves much less actual work.
Why I believe it should be at the full-rate: because I don't trust the company culture to stay the same over my tenure there. My expectations for a "shit company" have to be the same as my expectations for a "good company", because a good one can turn to shit quickly.
> But why? Why do you think oncall should be paid the same as full work? Perhaps you have a different definition of oncall than me, where you expect to be paged once or twice a week, and spend maybe an hour or so fixing it each time?
When I'm oncall, I need to cancel all my social engagements for that week and delegate all my errands and such to my partner. Also not drink or take any mind altering substances. I must be 'ready' at any time of day or night. I (as well as others) sleep in the same bed with my partner. If my phone rings due to an alert, my partner is also woken up. So I need to sleep in the living room for a week. From the start, this affects my personal life to the extent that it would be unfair NOT to compensate me extra. It also affects my family way more than a regular desk job should.
You're mentioning the expectation to be paged once or twice a week. If those pages come at odd hours and you need to fix them on the spot, no exceptions, failure is not an option, etc.. it's still very disturbing to your personal life. Additionally, that's a parameter which is well outside of your control. I've seen oncall shifts which turned from '1-2 pages a week' to '5-10 pages a day' after the product finally got in the hands of regular users or after the team grows in size and code contributions grow suddenly. Or even better, when you're doing such a great job that your boss promotes you in the oncall tier and now you also get to do triage for alerts coming for the whole organization.
The volume of the alerts don't and shouldn't matter. If you're oncall, you're oncall, you have a responsibility to be available at all times, rain or snow, night or day. This deserves compensation. It's the same as with regular work. Do you get paid extra when you merge more PRs? Nope. You're paid relative to the value you add to the company. Even if you have weeks in which you barely do anything. You're paid for your 'availability' first, then your work.
Some companies (some I've been lucky to work at) implement some sort of follow-the-sun oncall shift and you at least get to have your sleep and generally minimal impact on your personal life. That is great and does not deserve extra compensation, because your work hours aren't altered at all.
I'm sad that labor in the US don't consider paying extra for oncall a norm. But it's not surprising, considering we did have dedicated engineers at one time who were paid to watch and maintain the health of the livesite 24/7. But then we figured we'd make regular engineers fuck their sleep cycles by adding oncall to the list of responsibilities, because it would be cheaper this way. And everybody agreed, because 'full-service ownership' and we're already paid way more than in other fields. When the latter changes (and it will), we'll still not get paid oncall and I'd love to see the discussion when that happens.
> From the start, this affects my personal life to the extent that it would be unfair NOT to compensate me extra.
I don't think anyone is arguing that people oncall shouldn't be compensated extra. It's obvious that you should be compensated for being oncall, it would be criminal not to do so in my opinion.
The difference is that it's not full time employment compensation, because you're not working your normal work expectations.
No, you're right that it's not your normal work expectations.
It's working Overtime. Because it's availability ON TOP of your normal working expectations.
Overtime would be if you were actually sitting in front of your computer actively working on your project (coding, answering emails, bugs, feature requests, etc). Just being available counts as a remunerable activity but I don't think you'd be able to convince anyone that it counts as actual overtime duties like you would if you were actually overtime. It's "doing something" more than it is "doing nothing" but it's not as involved as actually "doing work" like you normally would. Hence, you are being paid for it, but it's not your full rate.
"failure is not an option" is not something I recognise, in the same way that sometimes features cannot be implemented as quickly as wanted, and systems are not as bugfree as I would like. But I am expected to put in a professional level of effort.
In my experience of oncall, it means carrying my laptop to social events, not drinking, and apologising if my alarm goes off in the night. For that, I accept the deal that is offered, which is less than my normal hourly rate, but still substantial given the number of hours.
If the volume of pages increased, or the required response time was lowered, I would reconsider.
On-call outside of working hours is simply a second job, so the above argument still applies.
Except you did. There are pretty specific legal definitions of "on call", what it means and when you get paid for it in almost every jurisidiction. I've never seen one that pays you time and a half for being "on call". This is not the same if you get called and actually work overtime; that's regular rules. How a company entices (or doesn't) for taking a shift is up to them.
The Kool-Aid was really good though! XD
I feel like this is a very absolutist statement that does not look at the actual nuance of the situation. I could maybe agree that a 5min response time (like Google Search or Google Ads SREs go through) could be argued to be "work" (although I honestly don't think so), but I don't quite agree with the definition you are using to define "full employment/utilization".
Assuming I have to show up at the office every morning at 8am, this is basically saying that my employee is restricting my "movement and activities" outside work hours because if I can't get to the office in time by 8am then it means I am not free to do whatever I want off work. If I wanted to go to Hawaii the same morning as I'm expected to show up at work, and have no ability to get back to the office in time for my shift, does that mean that my employer is restricting my freedom of movement hence I should be compensated for it?
No, obviously not, that would be ridiculous.
Er, yes, that's exactly how it works? You can't take a vacation in the middle of the week and expect no reprimand. Thus the same should apply to oncall.
There is no point in a company paying someone 1x for 8 hours of work and another 3x for 16 hours of oncall when they can just hire 3 engineers and work 8 hour shifts. That way they only pay 3x (instead of 4x) AND have the engineers do engineering work 24 hours (instead of 8 hours + 16 hours of oncall).
Companies need to stop squeezing by on free or under-compensated labour from their workers and instead hire sufficient numbers of people to cover the work they want to be done.
What a shocking concept.
Unless you are working half the weekend, every weekend you are oncall, the tier-1 OCC policy wins over time-and-a-half for time worked.
Time and a half for hours worked is only > that 2/3 for time not worked if you're working 50% of the time, which you aren't, at least not regularly.
What if the call comes right then? Now dinner is fucked. They'd better pay a lot to go messing with my outside life.
This is like the equivalent of saying dinner (or your day) is ruined if someone knocks on your front door unexpectedly. No it's not.
And yes, you're getting paid 2/3 of your (large) salary for the possibility of this inconvenience.
The alternative is that your company expects you to actually work full time for the weekend, since that's what they're paying you for. Is that really what you want?
That's a strawman. Surely there are more alternatives, so I question whether or not you're acting in good faith. There's some dissonance here bc I see you getting viciously defensive (is that really what you want?) over something you're presumably happy about?
Like unless you believe that there is truly no difference in the imposition of "normal" work and oncall situations, and you believe that no one else who is rational can see a difference, it follows that oncall will be compensated less, because it is a lesser imposition to the people who choose to do it.
If you believe there are other rational alternatives, present them. Don't deal in innuendo and then claim I'm acting in bad faith. Nor am I being defensive, lol. I absolutely, in good faith, do not believe you fully understand the effects of what you and others in this thread are suggesting. And given that you weren't able to present any alternatives, I still really don't think you do.
And the whole take is stupidly privileged, to boot. "I'm only going to get so many risottos in my life". Getting paid $65 an hour to not cook risotto is not an imposition to most people, or even most software engineers.
There is a difference—the on call shift is more of an imposition. During regular hours, I could just decide to go for a walk for twenty minutes and nobody cares. I can't do that if I need to be able to have a few minute response time.
Shedding the analogy, the pager went off, the whatever is down, and now the company is losing $X million a second. People who's experience is with systems where X is a large number are going to have different opinions than that of those where X is way below 1 and it's fine to take a few minutes to finish deboning the chicken.
That's what the company decided was acceptable, you aren't required to go above and beyond that. You won't be rewarded for it. There's no need to be a hero.
If they actually want a 1- or 2- minute response, then sure in with you, you're functionally chained to your computer. But we're taking about situations where that isn't a requirement.
Butt also, systems that cause your employer to lose millions per second don't really exist, at least in the steady-state sense. The highest gross profit companies in the world are on the order of $1 million per minute, across all revenue streams.
Unless you happen to be on call, of course. Are you going to advocate for not paying people who work in call centers for the time between phone calls? Being on call is working, even if you've just been told to hurry up and wait.
If you're going with the inconvenience definition of labor argument, the inconvenience of having to be somewhere specific is greater than...not.
put a different way, no rational person would choose $100k + time and a half overtime over a flat $500k to do the same amount of work.
This seems extremely unlikely to ever be the actual two options someone might have.
More likely is something like 100k + time and a half OT versus 120k flat rate.
And rational people will take the flat rate, because the company will assure them "We don't ask for much Overtime"
And by the end of the year you've given them 100k worth of free labour.
The pager should be tuned to your SLOs, and you should be incentivized to exceed those SLOs.
1. Time on-call was paid at 25% of normal hourly rate. (Maybe holiday premium boosted the base rate? I can't remember.) 2. Issue-resolution pay was the normal overtime rate, including shift premium and holiday premium, from the time the pager went off until the issue was cleared. 3. The person on call had to: a) be able to get to the plant in 20 minutes, if necessary, but remote support was perfectly fine and paid the same. Only resolving the issue mattered, not where you did it. 4. The person on call had to remain sober and work-ready the entire time on call.
It's 3 & 4 that justify the 25% pay for carrying a pager. Friend having a party? I'll have cranberry juice, thanks. Fresh snow in Tahoe? I'll have to miss it this weekend.
Restricting someone's movements and social life without compensation is simply abusive. As an industry, we need to stop.
Company policy said we had to have different devices on different networks, so the company issued cellphone was one, and the pager was on another provider.
Why would you pay for a separate pager device and service when every cell phone in the world can receive text messages? Bizarre.
They were cheap vs a phone plan.
They did have long battery life and ran on AAs.
They had clear notifications. Two devices going off? Better check.
Way simpler than ringtones for select users, or the “was that a simultaneous email and text?” game.
Mostly they were cheap. We promised we’d always have redundant providers. I’m sure this seemed like a good idea at the time but didn’t scale well.
They’re no longer in use.
Google will pay you significant extra money for oncall, with the huge caveat that your management structure has to fight for it. I have worked on three teams at Google over four years. All three had some kind of rotation, and none gave extra compensation. Usually they call it something else -- "emergency contact," "caretaking rotation," etc.
Am I doing this math right…
You get comp’ed 26 min for every 1-hour of Tier 1 Oncall?
If so, getting paid ~50% seems pretty darn good.
40/60 == 2/3
So for each hour, we can take either 40m of time off or 40m of pay.
In other words, we get paid 66%.
Take your salary, convert to hourly assuming 40h work weeks, multiply by 2/3.
That's the pay rate per hour oncall outside of normal working hours.
Google isn't even giving employees overtime rates for the work they do on call. It doesn't sound voluntary, either, and where are the handsome on call bonuses?
- if oncall is a part of the gig, you compensate _somehow_ (demonstrably above market salaries, explicit extra pay, time in lieu, etc); oncall culture (or the lack thereof) should be explicitly mentioned in any hiring process and employment contracts
- the team should be striving for 8 or more engineers in the steady state; temporary vacancies should be temporary
- primary should be handling 80+% of pages in the steady state; if this is not the case on average across the team, you are not building enough resiliency into your oncall culture, or relevant tech debt should be high priority
- relatedly, kpis/incentives should be structured such that as call gets worse, progressively more immediate investments are made to address technical root causes (a la SRE error budget)
I'm tinkering with that last one my head. It's easy to say, hard to execute
Being on call is super stressful, and if it’s causing burnout, you don’t need to keep doing it. Does this increase the burden on your teammates? Yes. But so would you burning out.
Absolutely. Being on call means we have to be ready to respond. Can't ever fully relax, can't make plans that compromise that readiness. People need to be compensated for that. Where I live doctors get paid when they're on call.
Its worth going through; the worst that can happen is they say no (or, admittedly, fabricate a reason to fire you), but you’ll know where you stand.
Two key factors are "(1) the degree to which the employee is free to engage in personal activities; and (2) the agreements between the parties." Beyond these general factors, no universal rule applies since the details matter (frequency of calls, response-time requirements and geographical limitations, etc) in the degree to which they limit personal activities, as do any agreements laid out in a contract or company policy (e.g., how specific 'on-call' requirements and compensation are defined and agreed upon in advance).
Note that these factors have only to do with the time spent _waiting_ for a call; time spent _actually working_ while responding to a call is more clearly work that should be compensated.
I really wish this wasn't stated so matter-of-factly. Neither of these is actually supposed to be true. A lot of times on-call gets stuck on these folks because they're often treated as second class citizens in the softwarescape. There is really great structure for doing these roles right that doesn't involve them making them full time on-call.
Sadly spot on. For anyone that’s not been the ops side of the fence, you are expected to do major works out of hours as it’ll impact developer productivity. You can also lose your public holidays, weekends and evenings to issues while others get to switch their phone off and forget about work until they return.
It does a real number on you. Glad to be fully on the dev side of the spectrum now but some of the attitudes of interviewers I had to endure to get there… And that was with a CS degree and a truckload of IaC/glue-ware experience!
I mean, really, what's going to happen if you can't see the score of a baseball game until tomorrow?
Well, nothing. Unless you’re MLB.com and have tens of millions of people paying you >$100/yr to have that information readily available. If that’s the case, you’re issuing credits (which is a huge time and money sink) and losing customers.
They won’t. No one will. They want their salaried employees to also be firefighters and the premise is absurd.
I don’t take calls at night; you cannot reach me. What are they gonna do, fire me?
Good joke. I run interviews and staffing someone who knows their left hand from their right is nearly impossible. Leverage works wonders.
Hate to break it to you, but starting your own contracting/consulting company means you are forever on-call. It just so happens that it's not called that explicitly, and the people you have to answer to are your customers (aka your bosses).
I get back to clients typically within one business day or I will show up to any meetings scheduled a week in advance. This has never been an issue.
I do not carry a cell phone and I make sure every client knows this. If I am outside my office I am living my life.
You essentially get a week to find bugs, fortify the application, and make it so that the next person has an easier on-call.
If everyone goes into it with this mindset, eventually on-call becomes a quasi-freebie week where you can either work on "fun stuff", or it becomes invisible.
Not to mention that no product can survive without love and support from its devs.
I don’t understand how it could tuen into extra time to work on fun stuff.
From the replies, it sounds like a lot of places have constant fires to be put out by those on call?
That doesn’t sound “well run” to me…
Not on-call engineers work on 20 story points a sprint (2 week sprints). On-call engineers (if you're on-call for a week) get 10 points of work + on-call.
That's independent from being oncall.
I prefer to spend my non 9-5 time with my wife and daughter. Sadly, many 'innovative' companies out there don't like this mindset of mine and reject people just because they don't want to do oncall rotations.
I have to give it to the companies and to the whole devops/agile movement. They have truly convinced us that being oncall is the right thing somehow. And that non oncall engineers are a somewhat inferior race.
Ok simplification of affairs here but…I mean…
If the business wants a stable and well working application, prioritise it as part of regular dev work. As a dev this is certainly not my problem.
Actually, a great way of managing this sort of stuff is implementing error budgets and SLOs. If you app isn't performant, the next sprint is dedicated to fixing issues, et cetera.
Additionally, the impact on personal live of being Oncall on the weekend is bigger. At commercetools, we recognize this by paying more for an Oncall day on the weekend (200 EUR on Fri/Sat/Sun) vs. a day during the week (150 EUR).
But when I think about arbitrary companies and people having regular jobs, I think that of course people should be compensated. It's labor and we pay people for that. And especially when it's more than just a few people, having on-call time and incidents be uncompensated means a broken feedback loop. The company should have strong incentives to make sure that on-call people don't suffer for the sloppiness of others.
Part of that is small vs big, or startup vs established. But there's some part of me that seems too reluctant to insist on proper compensation for my own on-call time. Clearly something I need to chew on before I next take a job at someplace larger than a few people.
I get paid a static sum (which is very competative) just for having the phone for the week, and then if I get a call, doesn't matter how minor, it's 4h of normal hourly pay for the call. This is only if the call is outside of normal working hours (so between 8pm and 8am)
I do really like this setup since it forces the employer to actually fix issues that are causing "wake-up". It should be expensive to make the call and thus an effort should be made not to make it.
I'm not sure how it works in other parts of Europe, but in Iceland this is considered normal.
We also try to split it quite evenly so it's usually only a week every 4-5 weeks or so.
But that means that in one case you make 200k base, have an on-call and don't get any extra on-call money. In the other case you make 140k base, have an on-call and get 10k extra on-call money.
Ultimately you end up doing the same work but one of them gets paid less.
“7. “It’s part of the job for all software engineers and not paid additional.”
This approach is common at many companies. A few which stand out:
Companies paying top of the market. Places which sit in the top tier of the trimodal nature of salaries usually pay far more without compensating for oncall, than lower-tier companies with very generous oncall compensation do.
Big Tech. Most of Big Tech don't pay for oncall with cash compensation. Google is the only exception.”
What we're discussing here is how companies encourage and reward (or don't) for the inconvenience and impact of staying near your computer, not going out of town, or being woken up in the middle of the night. None are going to pay you time and a half on regular work commitments because you might get called.
Jobs like a fire fighter are completely different. They work a scheduled shift and either respond to calls OR do other work during that time. They're not really on-call as much as prioritizing work. They also don't get 1.5x for their regular scheduled work.
In particular with regards to US law, salaried computer employees (and highly-compensated employees) are 'exempt' from the minimum wage and overtime rules of the Fair Labor Standards Act, so while it's true that you need to pay someone for doing work on an on-call shift, I don't think a US employer is necessarily legally required to pay them any _extra_ if they're an exempt employee.
First, I suspect (but I'm not certain) most companies do it as X% of salary. So I have no idea if I'm looking at truly different on-call policies or rather salary spreads.
Second, there's no associated estimation of how much work "being on-call" is. For us, a small team with SWEs doing voluntary on-call, any out-of-hours page is immediately top priority for work the next day. The person on-call also gets the final say over risky deployments after lunch / on Friday. I know that's not universally true, and we've worked with companies that consider a page a week or even more normal (still without a separate SRE/OpsEng team). If any of us was getting paged once a week, we'd refuse.
Your second point is definitely a major concern, though - the author talks about it (calling out Amazon and Twilio as particularly bad), but doesn't provide any sort of hard data on what the workload is like, possibly because it varies heavily even between teams or groups within the same company.
Obviously bad places to work, but there are many of them.
Any company that I've seen put SRE/"DevOps"/... as the sole primary on-call rotation basically just created a glorified operations team.
Unless you have shared pain for botched releases, you will never get rid of these problems.