Eliminating Toil
sre.google
sre.google
It's fun the the engineering at Google is so great at recognizing things, while the product/"human" teams (like whoever came up with the account reviews and other parts) seems to suck so much.
If YouTube applied the same view of what should/shouldn't be automated, they could solve the problem of peoples YouTube channels being locked in front of them, even if they don't break any ToS's.
You can't expect a free service to have highly trained human judgement whenever you want.
They need it to run fully autonomous. And does run flawlessly for 99% of the users, which is impressive.
I wish they offered a paid option in the 1% cases. Like an arbitration.
But that would be a cost center for them, and they don't want it.
If you can't run a service without arbitrarily banning people (and inadvertently affecting their livelihood [ignoring the question if people should put their livelihood up to a free service in the first place]), maybe you need to adjust the service you're offering people?
There is no requirement for them to run it fully autonomously, because obviously, they are unable to do so without punishing some of their users without any recurse. That's probably a sign it's really hard to run it autonomously, and they need to do something else.
If you don't want your free service to be a cost center, then maybe change it to not be?
They don't even manage to provide adequate content reviews for the .1% (probably an even smaller percentile) of YouTube creators that bring in the big advertisment revenues for the platform. Providing good support for them would easily pay for itself, and is the single biggest issues creators have been raising for years but nothing has substantially changed.
Partly due to a dominant market share. If there was more competition, they'd more about user experience.
However, I really feel for the author, because the distance between this philosophically ambitious position and the reality of using Big Tech products - which in my life are the primary source of pointless makework activity - is sad and frustrating.
Now, one cannot blame the tools for stupid policies that lead their misuse in constructing makework processes, but (to take the Heideggerian stance) they carry certain exacerbating values along with them. Technologies can be 'seductive'.
The point of departure for me, was the definition of "overheads" as justifiable makework according to some value set. Whose values, exactly? And if, in Weber's sense, bureaucracy is an unavoidable side effect of process, then the possibility to design "toil-free" systems is really about complexity management, not post-facto eliminating toil through automation, because that will only break things and introduce more toil (Which is of course the primary theme of Gall's "Systemantics").
I strongly support the Biden Administration infrastructure program. BUT selling it as a “Jobs” program (at a time of near record low unemployment) makes me very wary.
If your question is “how do we keep a relatively uneducated workforce employed gainfully” digging a canal with shovels Amy be the way to go. Especially if you only have a fixed number of useful canals.
https://en.m.wikipedia.org/wiki/The_Myth_of_the_Machine
https://en.m.wikipedia.org/wiki/Technics_and_Civilization
https://en.wikipedia.org/wiki/Max_Weber
https://en.wikipedia.org/wiki/Bureaucracy
https://en.wikipedia.org/wiki/Systemantics
https://en.wikipedia.org/wiki/The_Question_Concerning_Techno...
https://geezmagazine.org/blogs/entry/jacques-elluls-76-reaso...
There's ~4 orders of magnitude difference in the scope of work between the two.
I think the main cause of this is that humans tend to become very wrong and misguided without accurate feedback. An engineer who is wrong is proven wrong immediately; their code fails, doesn't pass tests, breaks, etc. A people-person design-minded product-guru takes years to be proven wrong, and even then their tactical obfuscation of reality can morph who ends up being blamed.
Several times in my career I worked on projects to reduce toil. Sometimes the project would fail because the time it took to work on them went well past the cost saving estimation. Sometimes they would be completed, but the value created was far less than their cost. And sometimes automation wasn't even the solution, and we just needed to change our process or system, or do some other manual thing that reduced the toil cost. Sometimes we chose to automate toil because we were afraid to take on a larger project we knew would make the toil unnecessary, so we paid for the automation and then later for rebuilding everything. Or toil was used as an excuse to justify a project that didn't really have to do with toil.
One of my biggest mistakes as an engineer was assumptions I made about my work that ended up creating more waste than value. Talk to an outsider about your plans and why you're doing it, take their advice seriously. And if your automation is optional, make sure you have buy-in before you start working on it; i've sunk months on things that nobody ended up using.
A great way to automate toil is incrementally. Typically you have a runbook with step-by-step instructions, and over time you automate one step, then another, etc. The investment is minimal and gradual, it can change over time, and you can target the costliest parts of the toil, optimizing value.
Or maybe in the process of automating, you discover new things about the process itself and can improve it.
I agree that one should be careful to consider work priorities and return on investment, but there are often hidden returns to something like this that leaders don't understand and take into full account.
But that assumes the task is still valuable at doing it 10 times more a day.
One key distinction to think about might be if your task is letting you reduce cost vs increase revenue.
A task that is a "cost" - e.g. if a user wants X, we need to do Y - likely won't need to be done 10x more frequently with 10x increased value if the demand for X hasn't changed. So you make it cheaper for us to do Y when X is desired, so the margin for X is increased, which still might make a ton of business sense, but the top-line boost to profitability is limited to the original manual cost of Y.
A task that is revenue-driving - e.g. "we have to do X any time we're putting together a sales deck for a new prospect" - can have a much higher flywheel effect. Can our existing sales team potentially now bring in 4 times more clients? That could be huge, and so you've both increased margin and top line.
Imagine for an ML product, making an accuracy report. If it's slow and required lots of human time, you might do it once a quarter for releases to important customers. If it's cheap and quick then you can run it on every CI run to check for regressions before merging code. Sure, you run it maybe 1000x more and don't get 1000x the value.
But, critically, the value is not the cost savings of not having to run it manually per quarter, the value is the more stable product and avoiding spending time bisecting a quarter of engineering work to figure out where bugs were introduced. And this was enabled by automating.
Yeah, it doesn't need to continue to increase in value linearly with repeated runs, it's the summed value that matters.
The "cost" is fuzzy too, often - e.g. time and budget spent on reliability-focused engineers or active troubleshooting rarely drops to 0 if you don't automate anything. It might just make it more expensive to react to incidents!
Maybe turn it away from a "gate" question - "should this thing be automated" - and into a prioritization one - "we could automate so many things, which ones should we do first?"
Because the problem with ROI calculation is that for some areas the ideal amount of toil would be 99% and for other 1%. For both extremes, you'll bleed valuable people, so sometimes your Investment becomes "hire new team just for that" and the rest is peanuts.
To put the thing back on its feet first create a team that takes 50% toil and give them these areas that ROI-wise require approximately 50% toil. Call the team "SRE". Create a team that takes 10% toil. Give different areas. Create a team that takes 90% toil, etc.
They missed the big one : human error is a common point of failure. Some of the big outages on GCP were due to ops configuration changes. Gitlab wiped their prod DB one time. KnightCapital suffered death by config error..etc.
Like if I need to change the spelling or add a new configuration setting and I need to make sure to use the same spelling in three places because they are all "stringly(sic) typed", is that toil?
In the below paper an example given is migrating from one API to another. The paper describes a semantically-aware large-scale tool for refactoring a Google sized codebase using map-reduce.
Given the externally visible churn in Google products it isn't much of a stretch to imagine they have similar or worse internal churn. In fact I have heard from xooglers that it was common place to internally have competing systems in different states of development and adoption.
True. But it is also common to find that software automating the process didn't cover some corner case and you need human intervention. And it's worse if the process assumed that human intervention would never be necessary...
Many of the automated systems at Google were developed by geniuses. Others, not so much, and it ended up making a lot of work for other people.
https://www.usenix.org/publications/loginonline/prodspec-and...
"This plant basically runs itself, but we do have a human present for if something goes wrong".. 50 years down the line, something goes wrong and nobody has the kind of insight and familiarity with the system that they'd have had it had been manually operated.
It might have been a important plant and the damages done when it breaks down and the surrounding society discovers they both cannot do without it nor rebuild/recover from it may be far worse than the "cost" of running it at a higher degree of manual operation.
I feel like the fetishization of the "job" is detrimental to society and individuals. I've been in a situation where I could really quickly develop automation that would replace a team of 15 people hammering on spreadsheets all day. The project was canned because we didn't want to put 15 people out of work. The way I see it, there were some distinct outcomes for possible decisions:
1: The automation project is canned. Nothing changes.
2: I get the automation done, 15 people are out of jobs. They have to seek new employment, or otherwise find out what to do with their lives. They are displaced in an unpleasant way, and the lives of many of them may be heavily damaged for at least the short term. The company saves over a million dollars a year.
3: I get the automation done, the company re-tasks the workers to other parts of the company. They aren't out of jobs, but they have to adjust and re-train and learn new things. Some are happy, most are annoyed, but nobody is seriously hurt. Some people are unable to adjust and maybe eventually get let go.
4: The automation is done, and the company continues paying the employees who now are being paid to do nothing at all. They can relax, or work on hobbies, or whatever they like.
The interesting part for me is that 2 has the greatest advantage for the company, 3 is a good compromise, 4 is the best scenario for the employees without even harming the company more than the cost of my time (which isn't really all that expensive in the big term; less than a day's pay for all the employees together), and 1 is the worst case scenario for everybody, but they chose 1 because the "job" is sacred. 1 and 4 are nearly the same from the perspective of the company, so rather than improving anything, they choose to be inefficient, wasteful welfare.
People want welfare to exist, but they don't like the idea of "freeloaders", so they force people to do useless, Sisyphean work. It's extraordinarily wasteful.
Maybe life could be better for everybody if losing your job couldn't completely destroy your life. We could automate things and improve things faster without having to hold back progress because "people could lose their jobs", we could dismantle destructive industries that currently are kept afloat because "Hey, that's 80 thousand jobs!" People could leave jobs that they feel are ethically wrong, rather than being trapped into doing something they think is evil because they need to feed their families.
I'm not sure this is a great outcome. It soft-locks them into their current position without any great motivation to better themselves. It also creates resentment in those around them. When the company does eventually let them go, only the most forward-thinking will still have useful job skills and can find another job.
Sure, they could quit or request a different job, but how many people can recognize those mental problems coming ahead of time and avoid it? Most people are going to be fat and happy and do nothing to get ahead. I don't even blame them. It'd be incredibly tempting for me, too. In fact, since I've been at this company so long and basically stopped growing, I kind of already have fallen into that trap. It's a pretty comfy trap since I like my work and I get paid pretty well for it. It's just not forward-looking at all.
Company B does the automation project. They can offer their services for a million dollars a year less than Company A. Company A is eventually outcompeted.
That's market capitalism!
[1] Who, then, programs and builds the next generation of automation... innovation would have to continue. These people would still go in for the grind, I guess it would be for more money but at some point would the tiered tax rates make it worth it?
[2] This would only be for certain sectors of jobs. Waitstaff is still going to exist, and a whole realm of the service industry. This would just lead to an exploited workforce, or striated (more striated) population filled in by immigrant labor and other "invisible" labor groups.
[3] Our dependence on machine infrastructure becomes ultimately vulnerable to attack from foreign intelligence and private actors and we are far from able to defend against it.
EDIT:
for the record, I have argued for UBI and am not against it in the past. I am still for it but not on a huge scale. The UBI that I have argued for would be a $1000/month UBI. This would be a replacement/supplement for SSI, disability, childcare tax credit, school supplies, food stamps, etc. It is not enough to live on but a support system for emergencies and savings.
Tossing the hot potato over to just about anyone else is bad.
For instance, if the company is pre-product-market fit reducing toil seems like the wrong investment; doing stuff manually can be the way to go until you find what works (unless the effort investment in toil reduction is trivial).
If the company has reached something approximating product-market fit, reducing toil still ought to be weighed against the other priorities. That (as all technical debt reduction) can do wonders to productivity, but alternatives (e.g. pushing for a new feature) may as well be the better call.
Now all I need to do is learn this new Domain Specific Language and find all of the exact configuration parameters to express my specific needs. Oh except this tool has leaky abstractions under it, and those tools also have their own DSLs and configuration parameters. And the tools under those do, too. It's all turtles, all the way down.
Unless you just meant the acronym - that I'm aware of but wouldn't say it's as common as the concept. To me 'toil' is firstly an English word, secondly an SRE term, and only distantly third 'time off in lieu'.
Just a bit sad that someone at Google seems to have read this and focused on the "Automatable" part going "but that includes basically everything we do!"
cf youtube/contentId, cf account blocking, cf customer "support", ...
> The characteristics of the technical phenomenon are Autonomy, Unity, Universality, Totalization. Technique obeys a specific rationality. The characteristics of technical progress are self-augmentation, automization, absence of limits, casual progression, a tendency toward acceleration, disparity, and ambivalence. [1]
Supposing the harm Google does (e.g. ambivalence towards individuals harmed by algorithms) is a direct result of this totalizing impulse, maybe it's time to question some of the fundamental assumptions present within.
Long, satisfying careers often involve proactive, design-oriented approach rather than purely reactive.
The only way to make grunge work an entire career would be if you’re constantly doing something for the first or second time, eg artists, novelists.
Even scientists, they can initially discover something significant, but they keep repeating the work on the same topic without more depth or breadth, the work will become tool.
> Eliminating toil allows people to focus on the inherent complexity of the difficult, interesting problems at hand, rather than the incidental complexity caused by choices made along the way.
> Toil can be eliminated...by drawing the system boundary a bit differently. When we use an external service instead of an external library, we’re moving the code outside of our system – thereby outsourcing the entropy-fighting toil to some third party. Not our entropy, not our problem
[0] https://www.stedi.com/blog/excerpts-from-the-annual-letter
And this is after having someone who is extremely aggressive with automation and empowered to do whatever they like to reduce that surface area working on the system. I've taken codebases and hacked out 60% of the lines of code in order to remove brittle external surface area along with unnecessary requirements and contain the project better within its own boundaries and stop repetitive issues. I've taken clever ideas that someone had 5+ years ago out behind the barn and shot them in order to reduce total surface area.
But people can walk into an area with a lot of toil going on and go "oh, I know all the strategies on how to reduce this, I will explain to these people who clearly aren't as clever as me how to do it" without realizing that there's often a minimum level of toil for a project which you can't effectively reduce. There's a nonzero vacuum expectation value of toil in any project, and in some cases it can be quite large. Inherently.
I don't know how many managers I went through who would come and decide to document all the different failures we were having and spreadsheet them and look for the patterns to address them. And every week there would be 2-3 that would come up and they'd struggle with the fact that there was really no pattern, other than that the project inherently touched many different third parties, because it really HAD to, and that those third parties would change, which would then force interrupt driven toil.
There's some point where you just have to hire more people and spread it out. There's no magical incantation to manage your way out of additional headcount.
And I don't think the OP article even touched on re-enginering to reduce surface area and brittleness. Automation isn't the only answer to toil. You can automate restarting a service if it crashes, but its always better to just fix the bug (which may involve fixing architectural issues) and make it stop crashing in the first place.
But I don't have any term ready for the very familiar category of problems you've described, well maybe except CI business-as-usual :)
With 150,000 employees.
Eliminating toil costs lot of time from every engineer.