Should never happen. If it does, call the developers
stackoverflow.blog
stackoverflow.blog
Wait... What? Maybe I've been living in my own bubble, I always thought devops was the idea that instead of having devs and ops seperated as two roles you should have the devs be in charge of their own ops. Similar to a full stack developer being in charge of both front-end and backend.
If you give the same person responsability of development and operations of what it develops, there is no more need to "get along", because the two groups are the same group. You are the dev and you are the ops.
But, yes, the original idea behind devops is that separation is nonsensical. It's just been misconstrued at many places to mean that we need more dev-y people in the ops role, rather than it's the same role.
I don’t believe that’s part of the SOX regulation, which for the most part says you must define your management controls, including your segregation of duties policy, and also have processes to audit that those controls are followed, but I believe is silent on the specifics of what that policy must contain with respect to one person releasing code to prod, particularly code which does not impact financial reporting. (Your company controls may have adopted such a rule, of course, which make them a SOX requirement for you.)
So, yes, you are technically correct (the best kind of correct!), but in any applicable sense, it is indeed a requirement. As I called out, there is loads of wiggle room on how to address it (technically you probably are in compliance if you just force PR approvals, but I'm not that familiar with the actual verbage, and am definitely not a lawyer), just that the risk aversion of most companies (and the intransigence of auditors) sees those who write the code (dev) and those who deploy the code (ops) being separated quite frequently in order to be SOX compliant.
Consider the jobs board for your company, the documentation site, or some other unrelated to financials, but production tech nonetheless. You can write your controls to say that someone can push an update without separating duties.
Just as we do for PCI, one key to making SOX sensible is to keep as many systems out of scope of the strictest controls, by sensible design choices (in tech and in policy) and be strict where strictness is sensible and more free when that freedom is more sensible.
SOX is only concerned with accuracy of financials and while it’s bad law in several key regards IMO, it seeks to accomplish this mostly via transparency and compliance with elected policy rather than specifically prescribing the policy contents.
Many internal functions run on Excel sheets using “we save copies very carefully on shared drives” revision control. When those get changed, there might not be one person who knows everything that changed, let alone two.
Here's an example of compliance setting a user back for no real benefit.
I'm in an organization that did this separation before we were forced to put our socks on. What we realized is that only a dev can usefully check a dev. Gating of code into prod needs to be by a dev, and the gating of infra changes needs to be gated by ops people. But that individuals, and the role, are best as combined as dev+ops.
So we implemented the separation, but not by splitting our worked into two separate groups and putting a wall between. We had other teams so we moved the review inter-team. I develop, maintain, and manage my piece of my team's service, but also do release-review for pieces of an entirely separate project that I have no access to other than to read the code.
Ok, but that just seems like one of those useless definition switcheroos no?
In that world, what has changed from before? When dev and ops were seperate, and dev and ops are still seperate?
I think when you need to scale up, a better model is to treat shared infrastructure like it's own business. You have products that customers can buy, rather than saying "operations manages the servers" and "dev manages the code". I have worked at places that do that, and the results were very good. I have also done the "dev manages everything up to the end user", and that's also great (but requires a training investment for team members that haven't done operations before). Finally, I worked at a place that split dev and operations as hard as possible and ... I don't think in the years that I worked there we ever ran anything in production. The process for putting code into production seemed to take on the order of years, so people just ran it on their workstation. Sometimes implemented inside an Excel spreadsheet. Now that I think about it, most of my time was spent debugging that and not delivering features. I wouldn't recommend it.
For example, I've been on successful projects where a representative from "Ops" sits in on early project meetings, offering help on technologies and deployment. And then later helps with documentation, automation and deployment.
I've also been on smaller projects, we're it might be just me doing dev and Ops,or where I do ops/architecture, and a few people help with development.
The main thing I see,is that typical modern software is deployed as services, not packaged apps that the end user downloads,installs and run (or Ops download and install for them).
At such a point,you need to consider security, backup, restore, monitoring etc.
Not all services need high nines uptime, but needed uptime and how to achieve it should be factored into the requirements.
I think one of the things many people get wrong with small projects/teams and scrum is to realize there may be more roles than people. You need the dba for the db, the network person for ingress, the dev for functionality. Sometimes that's one person with many hats.
But ops is a "real job" - at some point, it will demand one or more full-time employees worth of time. For maintaining quality you might want specialization - for communication you might want to share responsibility across your team.
Odds are, either way, for any given team/product/meeting some people should/will wear dev hat, others ops hat.
But I had the impression some (non-developer) operations people would see things differently.
In a bank you are not allowed to have devs running ops (or ops doing coding). You need clear segregation of duties. Ops don't code, they just install and run. Devs don't install and run, they just code.
You may have flexibility on minor, not SOX-scoped systems, non-material systems, etc. But have a coder with privileged access to your PROD and see the auditors having a party!!
Edit: yes you can have dev with prod access for emergencies. The access should be disabled 24/7/365 and be enabled for the absolutely necessary small number of minutes, where all actions are logged/traced/etc and then reviewed and signed-off. Still an auditor will challenge you for hours.
Operations and development have different priorities and skill sets. Having those two priorities under the same person or group is a recipe for disaster (either organizational disaster, or individual burnout).
I was around, and at times very involved with, the community coming up with this stuff. The description you're quoting is pretty close to the original intent, but at this point the phrase has been basically destroyed by management fluff books and consultants. It means everything and nothing at this point and is best left on the trash heap of history.
- "we have dedicated ops people who do things that are similar to programming", but also do all classic ops toil
- "We have SREs", both in name only, or in reality
- "all the ops-work is done by the devs"
- "devs and ops are separate teams, but, we make them work together better somehow"
- "we use kubernetes / we use docker / we use a cloud hosting service"
- "the devs are on call too"
etc.
That's why there's infrastructure as code i think. Some programmers in big departements started focusing on deployment.
( At least it happened in a couple of big companies I encountered)
I have a co-worker who takes it upon himself to "discover" and "fix" things that are broken, but to him, "broken" means "I don't understand this, or it's doing it in a way I don't like". He won't even read a runbook if it exists. What he will do is create a ticket, assign it to himself, and then declare it done when he's met whatever criteria he set for himself.
That's bad enough, but once he's done changing it to suit his personal tastes, that's usually it. He may, at times, write up something that more or less just reproduces what he did, but without any motivating context or explanation of the reasoning for going from A to B.
He strikes me as one of those "I'm trying to make myself irreplaceable by being the only person who understands this" people. In three years, management has nothing but praise for his "productivity", and they didn't even raise much of an eyebrow when he claimed "closed N tickets" as one of his accomplishments at his last yearly review.
So developers at your workplace can create tickets themselves, solve the tickets themselves, and then use the number of tickets solved as a measure of productivity? Only one word suffices to describe this situation: Minivan.
It isn't an official measure of productivity, but he managed to finesse it. There are management issues. There was some skepticism over this attempt to inflate productivity, but imho the response was overly forgiving and credulous. Something along the lines of, "well, he's a good developer, we aren't going to give him full credit for this, but he probably didn't create makework for himself". If he had reported to me, that probably would have led to downgrading his review and some minor disciplinary action for shenanigans.
Sometimes good corporate behavior even looks like sabotage. An old dev said something along the lines of "We reinvent everything because our review process rewards it. If you try to use an existing library you go to a review process and people undercut you and try to convince management you should use their pet library, only available on their stack of course, and you never get anything done. Instead just rewrite the bit you need and move on. Do what rewards you, not what you think is 'right'. Let the company decide." And indeed, some companies favor this because they can't get their process under control.
Once I was in a design meeting and my product had a config file (10-50 settings, tops) which got noticed. A senior architect said I should put it in a database ... and that the DB should be Oracle. "Nothing else is as stable or respected."
I argued a bit but when it was clear he was huffing and setting himself up for a win-by-authority I folded, left, and removed the config file so I never had to speak to that guy again. Now you set any changes in environment variables.
It literally would have quadrupled the work and footprint to add Oracle, or any DB. Had it been a compliance-req and not an idiot-req I'd have used SQLite or something, but Oracle was insanity. And by contract with our customers, had we used it, we'd have had to hire a specific Oracle DBA who of course would do no useful work in any other area.
Years later though, a friend asked me what would have paid better and I had to admit that I'd probably have gotten my promotion a project earlier by pointing to the big expensive parts we used.
[1] https://www.joelonsoftware.com/2000/04/06/things-you-should-...
That's a dangerous individual who should be coached (Lesson 1: Chesterton's Fence), or fired.
It is a red flag indeed, but “in hindsight”?
From my experience as an engineering manager, in addition to what article mentions, the following things are useful:
* having some kind of error codes (e.g. PRJNAME-E0032) so that search through documentation or knowledge base is unambiguous and instantaneous
* Any incident that was not handled by support/ops routinely, according to known procedures and was escalated to someone on call, gets an agenda item on a regular root cause analysis meeting. The result must be someone’s issue report: engineers on my team, infrastructure team, vendor, downstream or upstream application team etc.
* regular check-in meeting with support/ops where the central agenda item should be going over statistics about escalations for certain period and even incidents handled routinely (e.g. it was low disk space due to excessive logging and was handled without involving developers - do we log too much, do we need bigger disks etc?)
First off, Devops doesn't mean "dev and ops cooperating". Fundamentally it means that devs do the ops. Now you may have ops specialists who write tools to make it easier on the devs (hence the Devops, operators who write code), but at the end of the day you still have the devs doing ops.
Secondly, throw away your runbooks. They are out of date the moment you hit save. Spend your time writing code to automate whatever it is you're putting in the runbook and comment it. It will always be worth your time.
Not only will it be worth your time, but the engineer that takes over maintaining that code will thank you because they don't have to learn a runbook to fix things. They can run the automation you put in place and update it if necessary, if not just fix the code itself.
Also, if a dev is writing remediation scripts, they will often realize that they could just spend that time fixing the code instead, so that situation doesn't happen anymore.
Don't use runbooks. Documentation on how things work might be ok, but it's always better to document it in the code so that if the code changes hopefully the documentation does too.
For those that work in SOX/SOC2/etc. environments, this applies to you too. Yes, you can't can't have one person push code alone, but remediation should still be automated and not in a runbook. If you have to, make the code require an auth step from another engineer/ops person, but compliance really shouldn't change any of this.
The problem with “DevOps” is that everyone has their own definition.
For some people it’s a rebranding of sysadmins that use Infrastructure-as-Code tools.
For others it’s a term meaning the breaking of Silos between development and ops.
For others it means a (primarily) developer person who can also configure Apache.
Luckily, we don’t have to guess what was meant by the term as it was clearly defined at the time of inception[0], and unfortunately the original creator of the term does not agree with you.
I suppose if you’re right then I fundamentally disagree. Dev and Ops chase different goals and it’s very hard for a single person to have a role that makes them responsible for everything.
Exactly. Break the silo. Make the devs do their own ops. Have ops specialists write tools to make it easer for devs to do their own ops.
The fundamental point is that devs don't just "throw their code over the wall".
I feel like something like this should have been the preface to this post.
(I'm imagining someone from a ~10-25 engineer company reading this and figuring that if you don't have devops runbooks you're doing it wrong.)
The developers need to be involved in writing the docs, but they don't really need to be the ones writing them. Often, it's better if the community that will be reading the docs writes them, so it has the information they'll be looking for (somewhat addressed by a template). Or, a technical writer can write the documentation (with input from the people who are likely to read it as well as the people who have the knowledge, of course), which often results in nice documentation.
As mentioned, a feedback loop and editing documents as they are used is very important. Cataloging interventions and working to reduce them is also important. Sometimes it's the right thing to keep things manual with intervention from the runbook; there are many things where it's easy for a human to determine the scope of an issue and take the appropriate action, but hard to automate. But frequent interventions should be automated away if possible.