DevOps uses a capability model, not a maturity model
octopus.com
octopus.com
Before the popularity of the term "DevOps" it was always true that any System Administrator worth hiring knew how to script for automation (the job title for people who didn't was System Operator). Unfortunately the market was flooded with a lot of people in that role who barely knew what they were doing and so resisted every change anyone else in their company wanted.
From my perspective, the greatest achievement of the DevOps movement was to push the bar higher for the expected skill level of the average SysAdmin. I see that as a good thing.
The "DevOps" job title can arguably just be a sysadmin, but I do think the main ideals of the movement are just so tightly ingrained in software now that we don't even notice.
It's easy to forget that 10-15 years ago, the most common dev/ops model was "toss it over the fence"--developers write code, do a ton of QA work on the test system, and then toss it to ops, who had no idea what metrics to look for beyond standard cpu/mem/etc. Your dev team would write up a "runbook" for operators to follow if anything broke, which was usually just restarting things until they got into a good state again. The big tech players of the 00s like AWS and Google were probably more advanced of this, but the rest of the world largely wasn't.
Now, the most common model for companies of all sizes is "if you build it, you operate it". Devs are expected to know what metrics to expose, have oncall rotations, and have direct access to their production systems. Cloud, CI/CD and git helped quite a lot in this regard, reducing time between deployments from months to weeks to days to hours and minutes.
This continues to be the lasting legacy of "DevOps" IMO, it shouldn't be understated how much of an impact it had on our industry.
Hum... My first guess is that what really changed was the ratio of developers working on places that do this to the ones working on places that integrate the jobs.
The wall was never the only model around. It still isn't. The same kind of company that practiced it still largely have walls. The IT industry just hired a lot of people.
Well sure, and that's because software companies that integrate the jobs is now so much more widespread and commonplace, thanks to "devops". Kind of my point.
15 years ago it was often still I without the C, integration that wasn't continuous. Code freeze 1 month before release, put all the bits of code together, everything breaks, try to fix it.
Many struggle with "if you build it, you operate it" because many developers don't want to be on pager duty.
Yes, this is definitely happening.
I try to frame "DevOps" roles not as "doing DevOps work", but instead "enabling DevOps work". So, for example, setting up systems to make it easier for developers to take control of their own deployments and environments.
In my perspective, DevOps is a fundamentally different job from what Systems Administration has historically been.
It’s a strange new world to me but I like it.
We have both on-prem and cloud stuff, we've ran automation (via Puppet mostly) from the very beginning, and so far biggest difference is that writing template-backed YAMLs is utter shit compared to "proper" programming language or purpose built DSLs.
Like, I complained that Puppet is just kinda "shitty half-finished programming language" compared to just having Python/Ruby as DSL but boy I'm fucking happy to use it now (and to be entirely fair, it got better over time as a language) compared to whatever the fuck tool is in vogue this time that uses the "data language + template language" (because apparently programmers deploying the code can't program or something) model for interaction.
Same kind of work sans running to DC to fix stuff but you now have black boxes you have no chances fixing/analysing yourself and your broken code might fix itself next day because you thought it was your bug, but just a given cloud API decided to return nonsensical error that looked like it was your fault (greetings to MS Graph API team here)
As a developer and sysadmin, there is no trade and only
difference is that you probably (and I'm saying probably
because probably some poor fucker had at some point) won't
need to debug NIC driver/firmware problems on "cloud"
server.
Then the words "private cloud" drop by, and you find yourself fixing idiotic purchasing decisions that somehow led you to building custom firmware ROMs for intel X520 NICsI would argue that it was the other direction, in my experience the DevOps philosiphy was very similar in at the core to the agile philosophy; however it met the same fate as the agile movement. Everyone who was an "agile" consultant or "scrum master" found a new buzzword to declare themselves experts of and then use to go around doing a whole lot of nothing and generating impressive sounding promises, before moving onto the next gig.
DevOps died because it was a handful of engineers trying to force a movement about solving business problems. We were never going to be successful. Business people need to push the movement, not us.
That said, we can continue the movement anyway, if only to improve our own work. If more true DevOps faithful become managers, then directors, then VPs, then maybe in 30 years engineering orgs won't be run as horribly as they are now.
System Administrator responsibility is limited to...adminstration/operation.
If you have "DevOps" positions that have dual responsibility to operate and develop, that it something different, no?
---
I'm not saying this is actually the case. But DevOps being a job is not necessarily just a rebranding.
I wish I agreed. As far as I can see, the title is just as closely associated with being an expert consumer of cloud services as it is with actual skills relevant to development and IT operations, or anything we'd recognise as sysadmin today.
Having a DevOps onboard also surely meant that the company didn't need a system administrator anymore, but that's not really because DevOps was a new name for sysadmin -- they automated sysadmns out of existence.
I still prefer not to touch Web and service-style products. And, in my world, DevOps doesn't really exist. People with similar set of skills are usually called "infra" or "automation". Having worked in automation department one would most likely have learned enough to apply for DevOps position in a company which needs that, and vice versa.
Hackers that fiddles with stuff still definitely exist.
The same way the DevOps movement still exists.
But your point is not invalid, hiring a "DevOps" engineer is futile; especially given how the goal of any engineer in charge of DevOps should be to render their job obsolete.
- Use Terraform to build infrastructure as code
- Get involve in containerising applications
- Run and operate Kubernetes
- Spend a lot of time on CI/CD
- Improve the development experience
- Implement service discovery
- Build and run developer platforms
- (To a lesser extent) Build and run cloud environments including Serverless components
This seems like a new set of responsibilities which don’t fit cleanly into Development or Sys Admin, and are substantial enough such that someone could specialise in this role full time.
I think that DevOps as a job title is one of the best things that ever happened to the industry.
The lineage of interesting or useful things I do are tied to the client and dies with that client or when I leave.
I just think of the thousands of CI/CD systems, build systems, attempts at parallelising builds, impressive optimisations, tooling, automation that have been written for each company over-and-over-again, and there's no cross polenation except when they are open sourced.
I suppose Kubernetes is part of the answer here, a distribution of practices that survives organisations and spreads between organisations and client-specific lineages of software evolution.
I want to work on interesting capabilities such as diagrammatic observability and live visualizations of systems.
I really need to make a idle cloud environment simulation game where you invest time in servers, capabilities to handle load and problems that occur randomly or on a schedule.
Not Kubernetes (?) unless I missed something.
The rewarding thing for me is applying this and other sensible defaults like observing iteration speed and driving delivery..
Circuit breakers, bottlenecks, IOPs, load shedding, traffic behaviours can all be visualised.
I'm not sure how you would represent latency with this visualization but that's also important. It more represents throughput.
Can also be used to represent human work/tasks itself.
In real world, permutation and combination are endless. Every org has peculiar problem either created by the engineers themselves or are result of certain business decisions.
I hear you. IMHO there's massive opportunity / unmet need here. And working on the things that interest and excite you is the surest path to (or maybe even the definition of) success. I hope you can find a way to start pursuing your ideas! Good luck!
- AWS - Postgres - Linux - Docker - Jenkins or similar - Slack - Pagerduty - Jira - Packer - Terraform - etc
If you stay on the path there will be dozens of tools, plugins, and paths to do outstanding things with minimal work. Parallelising builds for example is built in to jenkins (if you define the workflow), which will autoscale workers in a setup that takes < 1 hr to setup in aws.
If you're writing code to solve a problem that a standard tool exists for, you're the problem.
You'll probably spend at least an hour figuring out IAM permissions before you even get to deploy a VM.
Something broke ? Well, rip it up and reinstall! But what about the data ? Who cares?
Platform engineering is a term I used to use in its place to try and differentiate but it seems that's being taken over now as well.
If your writing a script that will touch 10k servers, the operation is likely already slow. if you throw an unnecessary for loop that iterates over everything and runs something, that's going to a painfully slow script and wasteful.
This is true of software dev as well. Open sourcing things is something you have to sell well and demand up front.
In my experience, a big fancy post in smart words is never actually a practical implement to change the current practices of a company.
The only question is whether pushing for working things instead of nice talk will get enthusiastic support or opposition from the management.
I don't think you should feel excluded from "Software Engineering" just because you don't have a passion for managing containers. We need people who can write great code in the world, too!
The one area where you might want to focus is HPC. All the companies building huge GPT models need highly optimized hardware.
Personally, I do care that cloud is in the order of 10x the price and do have to explain infra cost as a metric of our road to profitability to our board.
My company does work closer to the metal, especially because for us the notion of “scaling” is not that we can simply slap a load balancer or a cache in front of a bunch of servers and call it a day: when you work with HFT or AAA Games: the performance you get on a single machine really matters, as does the ability of that machine to work reliably since there is state.
People in HFT and games really bleed for people like you and I, since its not as simple as CRUD stateless HTTP stuff where performance is measured in milliseconds and the average node runs 2GHz on all cores with 14 different abstractions.
Cloud optimises for the web, when its not the web, there are major dragons- those are your people.
Why can't you, as part of the process implement a capabilty driven dynamic (as in you review and change it) maturity model?
I mean, capability driven sounds great to talk about but how do you go about implementing it? A maturity model is simple to define and measure. I worry about endless meetings with what is proposed here but maybe I misunderstood a few things.
This post reads about just like that. I wish we had better insight into how to manage processes in programming businesses... but so far I haven't seen anything that truly goes beyond the obvious stuff. The only difference is how much the author is willing to elaborate on that obvious stuff.
On a tangent, I once drew 2 dimensions and put all our political parties on them. It turned out the "left" and "right" were trending toward "up" and "right" not opposites. Upper right corner was totalitarian BTW.
Suppose I plan to assemble infrastructure for some ideal of "Deployment Maturity" - no downtime, any time of day, one-click, etc. But it turns out the development team has designed the software to be un-load-balanceable thanks to in-memory sessions, so, thud goes my big plan. That's a very common problem.
Of course many of us see a "proper" devops discipline as interdisciplinary, so I suppose those folks would tell me to get in there, gently push devs out of the way, and fix that session mgmt problem. Of course I need advanced skills, but I also need advanced permission. Somebody's gonna fight me. Now I'm turning into more of a site reliability engineer.
So the maturity model definitely applies - especially when it comes to security - but when you're in devops-just-means-ops mode, you're much more tail than dog and it seems like you have no choice but to put capability first.
1. People or organisations wanting to implement process by well-defined stages and checklists. They'll see that works in certain situations elsewhere, but not realise that such processes will fall apart quickly in the complicated or complex regions they're trying to manage. Talking through where they sit on a Cynefin diagram can help them understand which action model is the most useful, whether it's really possible to define "best practice" for any given situation, etc.
2. Products being managed as though they were projects. Big organisations tend to run on a project model by default because it seems like a way for them to manage risk - a certain amount of budget signed off for a few months to a year that ensures X, Y, Z is delivered for a certain timeframe. The true risk is that absolutely kills innovation for an early stage product looking for PMF. You don't really know what the end result is supposed to look like, but you probably do know what the process for getting there should be. Being able to talk about complicated (often, projects) and complex (often, product development) regimes being distinct areas that require different handling is a good start.
Maturity models quickly devolves into cargo cults and of course making metrics a goal makes these metrics useless.
What really matter are meaningful capabilities, which when read aloud it just feels so obvious.
Once the research picked up, they started building out a broader picture of what a DevOps organization did and whether those things made them more successful (better at delivering software, more reliable, more profitable, etc).
The Phoenix Project / The Unicorn Project explain the concept by telling a story - they kind of tell the same story, but from different perspectives. There's also Investments Unlimited which takes an even broader view by adding governance, risk, and compliance (but in a way that aligns to DevOps).
In 2023, DevOps is best described by the DORA research (The State of DevOps Report) as it covers technical, cultural, and product concerns that all amplify each other.
Basically you clean up after big-shot devs, got it.
...so, exactly like being a SysAdmin.
If you have a good maturity model (e.g. for DevOps one based on DORA metrics) then the capabilities needed to arrive at a higher level can be determined per-org.
That solves the real problem. I think the article has a fundamental "AB problem" issue.
How do metrics like deployment frequency and lead time for changes can have an impact on "the capabilities needed" by an org?
Honestly, your comment reads a bit like machine learning generated buzzword bingo.
Hitting tighter metrics is going to need more advanced capabilities.
I guess... we're pretty much in agreement - except perhaps over the definition of the two types of model :)
Nicole Forsgren, PhD, Jez Humble, Gene Kim
IT Revolution
ISBN 978-1-942788-33-1
While I think it's an interesting read, we should take it with a grain of salt.
I'd argue that if your releases take forever, or lead times are huge they are worth improving first. What would you look at first?
Any other variation is not what Patrick Debois was talking about when he created DevOps days, from which the job title takes its name.