DevOps Is Bullshit (2022)
blog.massdriver.cloud
blog.massdriver.cloud
I am entering into my third major project in a row where it’s almost the exact same situation (and same consultant I’ve bumped into)
- site starts to scale, never had an ops team
- realizes oh shit we cant manage this ourselves but who wants to hire some hacker nerd for $150k a year that “gEnErAtEs nO BuSiNesS rEvEnUe” so they hire a consultant or consultancy firm
- firm builds them a solution that will technically “scale” but no one in the organization bothered to learn how it works
- it works for a while til it doesnt and then the “oh shit” moment becomes “oh shit we need a dedicated devops team”
- devops team attempts to work with the shitty consultant solution (editing yaml) and does an okay job until something breaks horribly again or they become a massive bottleneck
- finally they allow someone to come in and clean house and you get an infra platform like the author of this blog is describing
Platform engineering is the right word. “devops” has literally no meaning. I’m only 6 years or so into my career, but I can count on a single hand the amount of devops engineers I’ve worked with that can actually “dev” - and by dev, I mean do stuff like dig into application code and suggest modifications for the infrastructure, or write their own full fledged libraries.
Let’s get rid of devops and hire platform engineers. The problem is it’s a really weird and rare combination of skillsets, in a field no one really wants to do except complete weirdos (I am one). I’ve long said that DevOps isnt a career you choose, it just happens to you.
Well said! thank you so much for this post.
Somewhere along the way "sysadmin" became a euphemism for helpdesk, and some folks thought "devops" (as a job title) was new, but it was just the same old sysadmins really, or new devs becoming new sysadmins with the fresher title, they felt themselves the vanguard of a new paradigm, I'm sure.
Now we're in the same situation again, where "devops" (as a job title) can have no development experience. So we need a new term, apparently.
Maybe gatekeeping is useful to stop this job title treadmill.
This was a totally misguided re-branding of the field. It attracted bureaucratic personalities and classist managerial types, and utterly destroyed the field.
Like you, I couch my job description: I'm a "Systems Software Developer, I develop software systems" .. because its true, thats what I do, even if its a majority of the time working on embedded projects.
And if someone then responds with "ooh, so you're in IT", I immediately go to another corner of the party and avoid them like the intolerable bores they are ..
No need to be arrogant towards people outside the field. My grandma doesn't know the difference between the random IT guy that updates your Windows printer driver and a Linux kernel dev, but I'd still not call her an "intolerable bore".
It isn’t a problem for my ego to do that, and it’s something that’s both relatable and typically stops the conversation from revealing anymore about my work (not that there’s really anything to hide, I just enjoy my privacy and no, I don’t want to meet your “computer genius” nephew or fix your phone).
If I’m in a clearly technical group, I’ll go into more detail. I don’t identify myself with the DevOps term though, or sysadmin, or developer. I don’t clearly fit into any of the -currently assumed- definitions of those.
I always default to "I work in IT". If that person is in the space I might go into specifics; but even then majority can't relate or understand the role i spend my days working in.
It also has the benefit of coming across as more blue collar and I catch less grief, as the disdaine for tech workers continues to grow.
In social situations, I usually just say "I'm into computers". Most people who ask what I do don't really care, they're just engaging in ritualistic small talk anyway. If someone does care, they'll ask clarifying questions.
Hard to convince management that a dedicated onsite IT would be required - especially when they think it is your job already.
Gotta appreciate if you're working in the former. Drawback is you can't BS your way through issues. Everybody above me in the chain (which is like 4 people) could write a quicksort blindly. If anybody tries any BS they don't have their job much longer.
Well, nobody forces you to work for a company with managers like that. I personally wouldn't.
That's kind of like saying, even slaves have a choice. Death being the option available to all.
Drawing a comparison to slavery is quite shocking, tbh.
A good example is how software releases work. Where I worked 15 years ago, releases were something we could afford to do about four times a year, at night, on the weekends, doing the rollout by logging onto machines and starting them back up, babysitting the rollout. Because rollbacks were catastrophic, an insane amount of time went into QA, so the time between a code change and a release could be easily multiple months. Now I'm at a place where releases are done every day of the week. Rollout steps are largely automated. Rollbacks are mostly drama free. Risk is mitigated through automated canarying (and automated testing before that).
And certainly the business has successfully gotten rid of paying people to log in to servers on a Saturday night to perform a release (which frankly not a lot of people are into in the first place, so everyone wins).
[Weeps] That wasn’t the point, the point was exactly the opposite of that, to get developers and ops working as a single team, building an understanding of each other’s domain in order to make things better. Sadly the point was lost long ago, and devops engineers and devops teams have taken the place sysadmins and ops teams in providing a nice little silo you can chuck your terrible code over the wall at and not have to think about how it works.
But you know what? There are tons of developers who are _perfectly fine with this_! There are tons of developers who just want to write business code in their CRUD webapp and call it a day, the more abstractions you can give them, the happier they are. And while I don't like it, I don't blame them.
Where this falls down is when development teams are building complex systems and then expecting the ops team to divine how that’s meant to work in production.
the ole “runs on local therefore it’s good to go” should be good enough.
Which database you choose and how you scale it will affect your application code.
You can't write code that makes use of the local file system (or in-memory data for anything meant to be persistent) if you have more than one container serving your app. (And yes, some people need to be explicitly taught not to do that.)
And so on.
What tends to happen is that at some point in time, an organization finds itself in a bad place. Where pagers fire regularly, and the skills you need to manage the noise are "fast context switching", "good communication", "debugging", "ability to work insane hours". Very little of this involves coding, or building software. Teams may start by assigning some engineers into this role as an elite team, but if you come back in a couple years you'll find that the elite team has burned out or forgotten how to write software and due to the impossibility of hiring new elite team members - has resorted to hiring people to carry a pager.
There are a few companies which invest in big platform orgs, but due to the lack of revenue generation - it seems that this can only be sustained with very strong leadership buy-in.
I absolutely agree with this. It seems like a it inevitable for anything that a business needs to do but doesn’t really want to do. Another place where you see it is with project management, where there have been many different titles for the same role.[0]
Definitely. often it's more about organizational changes than software and upper management is the one blessing things in most big corp.
One of the fundamental shifts in my opinion was building APIs on top of sysadmin tasks. Provisioning a VM, installing an OS, getting software onto the server, configuring ingress/egress, handling process restarts, etc.
These all were turned into tasks you could run from a developer’s laptop.
But it was still a different skillset and devops dumped that skillset on developers (“Hey! Now, instead of opening a ticket, just write a 1000 line magical incantation of poorly understood YAML and the system can do it for you!”).
Platform Engineering grew out of that pain. It’s the layer we are building on top of devops in a way. We are taking those APIs and treating them as foundations to build products that manage the SDLC lifecycle for developers.
I think at its core, deveops was about “how much sysadmin work can we turn into an API and give to developers?” The answer turned out to be almost all of it. Then Platform Engineering is “now that we have APIs, how much can we automate away for developers.” We don’t know the answer yet, but suspect we will find the answer is also “almost all of it.”
Site Reliability Engineering seems to be following a similar story arc and converging on Platform Engineering.
At my current company theres literally no coherence. Some people deploy on EKS, or ECS, or Anthos, or regular linux hosts. We use every database on earth. Things on Azure GCS and AWS. Databricks and Snowflake. The list goes on!
A lack of gatekeeping I don't think was ever the issue; the incentives were just perverse.
I spent time working around the space industry. Almost a decade of my life.
The big contractors in aerospace have a _really_ difficult time finding people who know enough about the hardware and the software to be competent engineers. They’re basically unicorns.
Not that they don’t exist, but if you’re looking for someone who is demonstrably good and has a validated educational background in both, you don’t have a huge pool of talent to go with.
So they hire them when they find them, but they also just tend to hire experts in each field and let them collaborate. They didn’t try to reinvent the wheel with some new “paradigm shift”.
The team I was on was made up of ops/engineering[1] people who knew how to code. We built tools to fill gaps or let us do things we couldn't do otherwise, but would typically use out-of-the-box functionality in third-party products by default because it was cheaper in the long term than building and supporting custom code. Maybe there's a better term for this, but I think of it as "OpsDev", and AFAIK it's more or less extinct.
DevOps has always struck me as being developers[2] getting fed up with having to deal with a separate ops team, and building tools and platforms to automate away that ops team entirely.
The two things can look superficially similar, because they achieve similar goals (automate away repetitive work, let people do more work), but the approaches are almost totally different. Kubernetes and Terraform do amazing things, but they are very obviously tools written by and for developers. I think it would be difficult and unwise to hand off first-tier support for issues involving them to a help desk, for example. It's pretty safe to write instructions for first-tier support to use a GUI or run a command but substitute the name of a server farm and a software package in the arguments, especially if the command is a script written by the engineering team to verify that the work makes sense before executing it. It's a lot less safe to try to have non-developers hand-edit source code, a script or even a YAML file and push the result to production.
Both approaches work. They just require different kinds of people, and each has their own benefits and drawbacks, and IMO it's a mistake to put them both under one label.
[1] "Engineers as in Geordi LaForge or Montgomery Scott": the people whose primary job is to keep the engine(s) running, respond to emergencies, make improvements where possible, and generally be a smart-person-of-all-trades.
[2] "Engineers as in Burt Rutan": the people whose primary job is to design and build new products (for lack of a better word) from the ground up.
https://www.linkedin.com/jobs/junior-devops-engineer-jobs/?c...
No programming experience required (only some minor scripting in shell languages & a reference to education), in fact this job ad reads identically to a sysadmin job from 2014; in the middle of the "sysadmins can't code" zeitgeist, when all developing sysadmins had moved to become devops engineers or whatever.
Copied from the ad:
----8<---- Your tasks:
* Containerization and orchestration of our products for development, testing and production environments designed for cloud and on-premise solutions
* Produce reference architectures and standardize solutions to streamline efficiencies within the business
* Mitigate against Cybersecurity threats by ensuring systems are safe, secure and compliant to best practice recommendations (e.g. CIS Benchmarks, Vulnerability scans and Penetration tests)
Your profile:
* Finished studies in Computer Science or software engineering (Bachelor, Master)
* Proficient understanding of virtualization, containerization and orchestration (e.g. Kubernetes, Podman)
* Experience working with CI/CD tools such as Azure DevOps, Jenkins
* Experience of Infrastructure as Code tools such as Ansible, Terraform, Helm
* Scripting knowledge including bash, PowerShell
* English fluently, German is an advantage;
* Improvement-driven, structured and proactive team player, high sense for quality, open minded
* Working in an agile environment (Scrum, Safe)
---->8----
The point is coding skills are not required to do the job. Sure it would be nice but if somebody ticks all the other boxes then being good at programming isn't important.
What I'm trying to drive home here is that you do need programming experience to be a good sysadmin/devops/platform; back when devops as a job title was becoming "a thing"(tm) I was told over and over again that sysadmins couldn't code and thus we need devops.
Now it seems devops don't need to be able to code (which you seem to think as well). So we will come up with a new term to cover programmers who understand systems.
Those "degree" sections are universally optional, they're always listed by I have never seen them honoured.
I'd put $20,000 down on the wager that I could get that job with only a Cisco certification (or even nothing at all).
But the criticism was that the job ad was not requiring knowledge of how to write code, and I was pointing out that the degree requirement (strict or not) indicates that that's not true.
If you don't have the required degree but can demonstrate equivalent skills, then that's fine. Kudos to the company about being flexible on the formalities side.
There is a line that is universally ignored in job ads. This is that line.
if I showed no ability to write a for loop or a linkedlist. I could still get that job.
I see it time and time again, it's like we must add that line to job descriptions, but it's never checked, never verified and ultimately it serves no purpose. There will be no difference between someone with a BSc in CompSci, an MSc in CompSci or no degree at all.
The thing that will get you that job, likely, is discussing how http works, how tcp works and perhaps something to do with exit codes.
Might not be true for this particular posting.
The only european country that takes titles even more serious is Austria. So without a M.Sc. in CS or a similar M.Sc. your chances may be quite low to even get an interview. Especially since university is basically free here so many people have a B.Sc. or M.Sc. in germany.
No BSc or MSc here.
I’m sure this is the case in some places. This can’t be farther from the truth in many others. At every organization I’ve been a part of, the best “DevOps” on the team was usually the person who was the best at coding, or at least understood it the best - and coding is one of those things that you can’t really understand well if you don’t do a lot of it.
But plenty of developers only have a vague picture of what production looks like especially in a complex environment. However they are perfectly able to do their job okay and produce working code.
Sure, maybe some of the work isn't writing software. But some of it is, and it shouldn't be written by amateurs. Also, they are building these things for people who do write software. It would be hard to produce something that works for them without understanding what they do.
Reading code, understanding building and deployment are all useful and required to some extent (especially the latter). But specific programming skills are not required any more than that would be for say a technical writer or network engineer.
Proficient understanding of virtualization, containerization and orchestration (e.g. Kubernetes, Podman) Experience working with CI/CD tools such as Azure DevOps, Jenkins
That typically entails programming experience, especially in Germany.
Most companies do. Perhaps not top Silicon Valley tech companies with cutting edge needs and paying over $150k, but overall I've found that dev ops teams have plenty of jobs for 'green' team members. You just have to accept a 'green' salary and 'green' job responsibilities.
Yes. Devops people come from a SWE background generally + they understand infra. They usually have 10+ years experience. Most specialties work this way in software: security, devops, AI, ML, ... etc. You need to be a SWE first, then specialize.
Also I work in AI/ML and hardly anybody has 10+ years of SWE experience before specialising. Most people start fresh out of a PhD or postdoc program or come from some sort of math or physics background.
Now it's actually 80:20 due to software engineers becoming frustrated and disillusioned with the current state of things. Not only due to specialization and the push towards tool standardization (you don't get to build many custom components because we k8s all the things so most of the time you're just writing yaml/hcl), but because working with regular sysadmins gets extremely frustrating.
This could be the entire problem. These are specialized fields. They are being treated like they are not.
One thing that was very different was provisioning new machines consisted of ssh’ing in and installing the required packages and binaries. The deployment was probably a simple script that used scp.
However, once the term DevOps was coined almost immediately some people began to stop doing development work and just focused on automating operations work and constructing glorious crystalline castles in the clouds to handle traffic that could have been realistically served by 1U in a shelf somewhere. Complexity begets complexity.
In our arrangement, the ability to push code to production is gated by the GitHub/Azure integration path. The QA or project person who is rotating the production deployment slots (azure functions) is not granted access in GitHub to deploy to those same functions.
So, the developers pushing code and those deploying code are mutually exclusive groups. You could still defeat this with collaboration between employees or screwing with AAD records, but that's why we have a ton of audit logging turned on too.
Absolutely, you need proper controls in the system to control and authorize deployments with appropriate responsibilities in place, but fundamentally if your deployments are fully hands-off, it shouldn't matter who clicks the 'deploy' button
DevOps is similar to agile. Originally they both described a set of principles, an ethos, not a specific role. DevOps was your development team work closely with ops as one team, using software engineering practices, config as code etc etc.
Then, it just came another name for sysadmin and later platform teams and back to where it all started, passing things over the wall from engineering to platforms.
The way I understood DevOps (back from the first DevOps Days in Ghent) is that clear separation of sysadmin and developer roles harms the process of software delivery and that developers should maintain their code up to production, and learn from the process.
It was, in a way, a critique of the project-driven business where teams ship and jump over the board.
Then it turned into this guy who can not really do anything. Can not really check in code because the devs do that (separation of duties). So you never get any real experience in the codebase. Can not make any decisions because the marketing/product owners do that. Can not really make better test cases as that is for QA. You can however help on the helpdesk they are short 3 people today. Then the meetings. Endless meetings upon meetings. Because your 'devops' you need to know everything that is going on and better have those statuses on the high priority items to be fixed. Oh and you are 'on call' so you can never really 'go home'. Then if something breaks all you can do is just call in others and make them fix it. Making you little more than tier 3 helpdesk. Useless paper pusher job with zero authority.
I enjoy making sure software works for people, the more people the better. If that involves all that sysadminning, I will enjoy that as well.
But most of my peers are oblivious to the fact whether their code will be used at all and are having fun filling their heads with complex abstractions right before spitting them on screen.
Point is, once developers start doing support, they stopped taking those learnings and going back and developing better. Instead the organization just started having developers do support. But then, why are the support people these highly paid developers, lets change the role definition again to not need to code.
Seems like on-going title shuffle, around, we don't really want to pay to develop, why can't we get by with no-code.
If you want to cut through all the BS, this is ultimately a debate about cloud vs on-prem relative to our ability to manage complexity across multiple vendors & styles of product. If you go to the cloud and actually do it entirely their way without antagonizing the process with 3rd parties and other ridiculous nonsense, you don't need leet hackers to figure things out for you and build elaborate processes that you'd require a "devops" team to support.
In all 3 major clouds, there are well-documented happy paths that a junior engineer can follow to put a perfectly functional webapp live on a TLS-secured domain in about 1 business day. As far as I am aware, "getting the site live on public internet and making sure the certs are ok" is DevOps bread & butter. I am not going to pay for a dedicated team to do this kind of babysitting in 2023.
The real tragedy is all of these companies who "went to the cloud" but just wound up in an even more complex & contorted hybrid stance with the cloud simply appended to their vendor list.
"By leveraging the capabilities of a multi-trillion dollar Death Star ..."
The objective is to get completely out of the game of managing VMs, certificates, authentication, MFA, et. al. There is no way in hell we can do a better job than any cloud provider when it comes to compliance, consistency, security, etc.
Maybe when we have 10k+ employees we can reconsider the benefits of suffering additional complexity.
That's a hell of an assumption.
> there are “easy” plug and play solutions offered but they cost a lot and come with tons of drawbacks.
Agreed. But, I'd argue that these drawbacks/constraints are exactly what you need when you are faced with a virtually unlimited # of choices. When you are building software that needs to survive audits and is sold to financial institutions with 5+ year contracts, intentionally paying to have your hands tied is a fucking blessing.
I also used to shit all over the cloud until I was put in charge of the whole product stack and had to field difficult phone calls from other CTOs. You may find your position change over time too if you are put into this kind of a situation.
First of all platform engineering is two words. second of all what is my understanding that devops = platform engineering? Do you have a dictionary definition for these? No, so it does not make sense to say that you should use X instead of Y. Btw. systems engineering and technical operations are the things we usually mean by devops, deploying, configuring and scaling applications in an IT infrastructure and dealing with different aspects of running these applications. DevOps was just an approach to remove the silos that some companies built between development and operations bringing good old tribalism to technology because we cannot exist without tribalism.
> Let’s get rid of devops and hire platform engineers.
I wish this problem was so simple.
> DevOps was just an approach to remove the silos that some companies built between development and operations
I know this is always the intention but it is never implemented this way, it becomes a “throw crap over the wall” situation again because now you’ve made a second team and management sees that stuff as “their” job and the silos develop.
I don’t profess to know the correct way to do it, but if I was managing an IT team, I’d assign a dedicated “Ops” guy to every dev team who would occasionally also write code. I dont care what you call this person. I’m in favor of platform engineer because it’s a lot more defined than “DevOps.” People have wildly different ideas about what “senior” devops is too. With a term like platform engineer, my gauge for if a person is a senior at that role is by the quality of the platform they can deliver to me. it’s vastly better, but again, I don’t give a crap what you call it as long as organizations finally understand what they’re hiring someone to do.
I like the term platform engineer or maybe just infrastructure engineer, but I don't think they really add much specificity. For instance, we don't use any of the big three cloud providers, which I suspect would be an expectation for most of these titles. I think the possible diversity of modern tech stacks is just too much to pin these titles down, and mostly I just tell people I'm an "IT guy" ; it's only when you talk to someone in the industry that titles like Devops or Platform Engineer have any meaning, and as this thread shows, there's not even that much consensus within.
To me that is simply not DevOps. That's just Operations. Businesses are throwing around the title without a care for the word
Can you be more specific?
From what I understand, "DevOps" helps you get what you want into production. Whether that be a few Docker containers of your node.js or Rust or Go or Java or C# or whatever application (probably an HTTP microservice or a batch job) as well as your infrastructure (Redis, Postgres, a message queue maybe)
Obviously scaling your microservice API from 2 to 4 isn't "all devops do". I also get that not everything is a Docker container. I'm just curious what I am missing.
I get that for example, making Postgres "scalable" is much harder but... is that really a "devops" thing or is that a Postgres-specific thing?
Same for Redis if you need multiple scalable instances.
Scaling infrastructure isn't easy but I thought devops was just responsible basically for standing up the "pipelines" to deploy whatever you want.
Why can't you just have a monorepo of application YAML that goes into ArgoCD or something like that? Where do you hit the "need devops/consulting" part specifically?
I think they might not “need” it as much as they think they do, but they want faster deployments, better automation, less downtime. Usually the devops team is the one that handles this for most organizations. ramming code into prod by scp’ing some files into a single monolithic host can get you a long way, but at a certain point, it doesn’t, IMHO.
Which part of that isn't covered by `kubectl apply / argocd sync` in a pipeline of some sort like a Jenkins or a TeamCity or a GitHub Actions?
damn, that sounds pretty good. need to look into some devops courses.
It always has been. Always!
This is why FAANGs paid their SREs top dollar.
The problem, as you noted, always starts from the top. Cultural change needs to be championed by someone with power inside of the company. However, cultural change is slow (site starts to scale) and expensive ($150k hacker nerd) and many executives still think of tech as a cost center to be optimized (see also: ChatGPT).
Platform Engineering now is where "DevOps" was in 2014 or so, but with more Kubernetes and Rust. But as long as executives who do not (or refuse to) understand the value that strong engineering teams bring forth, the Platform team will just be interchangeable with today's DevOps team.
The AWS Lambda + Serverless framework shit trend has to go. There are no savings or agility. It never handles everything. The gaps are the painful parts, not the bits it gets right.
I'll choose a simple nodejs app any day over that yaml trash fire.
I do not think my organization would call this a platform engineer, as that title makes no sense for the work I do.
TIL about platform engineers. But I'm still unsure what kind of things they produce. I know that they basically create a self-service organisation-specific infrastructure that can be used by developers inside an organisation to deploy software. But how does this self-service look like? Do developers fill in a web form to automatically get a server with the right firewall rules?
The whole idea behind platform engineering is the mindset that the "platform", i.e. the stuff that runs the things that business cares about, is a product.
Reframing the platform as a product instead of "IT" or "the place where apps go" introduces three concepts that were previously foreign to operations:
1. The platform has customers instead of consumers, and getting/acting upon their feedback is important,
2. The platform has some semblance of reliability guarantees, and copious observability is needed to uphold them, and
3. Like other products, the platform needs a dedicated, cross-functional team to maintain and improve on it.
So basically platform engineers are engineers that focus on either the "developer experience" part of the platform (i.e. the platform's frontend) or the platforms back-end (servers, persistence, networking, etc.).
> But how does this self-service look like? Do developers fill in a web form to automatically get a server with the right firewall rules?
Kind-of. In an ideal platform, developers wouldn't even know or care that their stuff is on a server. It could be in a toaster, for all they care. So it's kind of just there. They write and maintain code locally, and their build and release pipelines stores and deploys into it.
In reality I spend my day doing everything from yes infra work and living in yaml, to working in Go/TS/ruby/elixir/etc repos and building those as well.
We have dedicated infra team, and as platform engineer I have to know their job to an intermediate level in addition to having advanced level in shipping code across stacks.
I love it but most other people seem to struggle with the constant context switching.
I’ll add that platform engineering is where much of the hardcore engineering happens nowadays in modern software stacks. In the olden days, SWEs had to know the hardware/low level software stack themselves.
I think the SWE role is what is at risk of commodization. It’s something sufficiently good LLMs may even be able to do themselves at some point given the right instructions.
Platform engineering will never go away - it’s close to the hardware and so many things can go wrong. You’re building the abstractions so SWEs can sit in their cozy IDEs debating about overly complex language features and libraries (said tongue in cheek as I’m guilty of this plenty).
Here’s where it went wrong.
> “devops” has literally no meaning.
No, it does, you’ve just chosen to let vendors, consultants and managers redefine it as “whatever we were already doing”.
I'd recommend the lead takes the role.
Align incentives to create simple, reliable software.
This sounds like an incredibly toxic work environment
This is the way.
What is an "infra platform" ?
Is having a DevOps team wrong? No, having automation experts, infrastructure experts, integration experts, etc. all makes perfect sense. Bundling all those disciplines under a single term is what is wrong.
Is aiming to do "less DevOps" using "serverless" wrong? No, aiming to reduce infrastructure maintenance and design overhead is perfectly sane. Interpreting that as entirely avoiding infrastructure, or as affecting any of the other disciplines is what is wrong.
What we need is better work descriptions and team organizations. But while the buzzword DevOps is bad, it's not so bad to have a fancy term that executive management understands and can put some money behind - money that allows you to spend time or buy solutions to solve your actual process problems are.
It just became a buzzword that everyone had to have, then C-level leadership needed "invest in DevOps", "grow the DevOps team" (because buying is easier than fixing the culture) until we lost sight of the original point and DevOps just became... ops.
My last job was the first place I saw DevOps as developers working along side operators as DevOps, which was a slight improvement over when they didn’t work together, but seems unnecessarily disjoint. Unfortunately, they had compliance requirements that forced developers and operators to be separate humans with separate access. Not wanting to get back into that world, I’ve not looked at how strict those requirements are or if there are ways around it (HIPPA, FIPS, etc.), but it was an annoying distinction.
This is also exactly how i've always understood the term, and it makes perfect sense.
But the origin of the term is indeed in distinct dev and ops people working more closely together. The origin story is basically Patrick Debois watching this talk, and then turning the title into a noun:
https://www.slideshare.net/jallspaw/10-deploys-per-day-dev-a...
I think the devs-do-ops idea was also floating around at the time (i was doing in in 2008!), somehow got swept up into the movement.
Personally I view platform engineering (or SRE) as the best endpoint of the ops role and devops philosophy. A good platform engineer is a developer that knows OSes, networking and databases and can expose abstractions that meet the requirements of application developers, akin to how a good application developer knows their the business domain and can expose an interface that meets the requirements of users. Nowadays, the platform engineering is often done by AWS or Google and that's sufficient for many developers and smaller organisations, but many orgs will probably still need some people of various skillsets between what is exposed by the cloud and the business requirements.
I mainly think traditional ops is just a dead end. If something happens often enough that you feel like you need to develop a runbook and hire someone to follow it, you should just be automating the fix or fixing the underlying problem. DevOps in the sense of applying a development philosophy to ops (i.e. automation and version control rather than paper runbooks) and applying an ops philosophy to development (i.e. writing stable systems that are robust in the event of failures) is still a valuable philosophy, but the word itself is meaningless now.
Like, if your DBA is attending plannings of teams working on a database heavy project. That is an action following a devops-mindset - breaking up siloed knowledge at the right places in a good way in order to save time on both sides. Weeks of planned work can disappear with the words "Oh there's a postgres extension for this. Give me a week to check it out."
Or, if we work on some system so a consultant managing a customers system can roll versions forwards and backwards automatically with a simple UI. Again, something following the devops-spirit of removing yourself from the loop, putting trust into the automation and enabling people to work with less friction. But again, this isn't done by a "DevOps engineer", it would usually be done by one of our operational dudes with a focus on automation and pipelines in close cooperation with the dev-team.
It may be picky and pedantic, but it steers thoughts into a better place. And it gives us an answer to "What the hell is ops doing all day?" Making your devs better, and you money, good sir.
IMO there are 3 distinct roles that product gets thrown into, 4 really but the technical product manager is completely different.
Research/data, operations and UI/UX and then the 4th being the technical product manager.
The person that maps and sets up the metrics, KPIs, product intelligence and knowing the business model is very different from the person that knows the UX flow and can inform the UI.
The problem IMO lies in the middle management layer, maybe ai will be able crack that but for us ;)
It seems to be a problem all over many disciplines tbh
What we need is LESS HR.
And what youre saying is what we just refer to as an SME (Subject Matter Expert [SUBJECT])
Its stupid to worry about labels. But if youre an SME on the sites DBs, or Network, or scaling factors etc... your the SME for [SUBJECT] and you can still be an SRE, Dev, Ops, Network guy and still be on the Devops team/ops team/ IT team.
I've been at a large organization for a while now and my backlog is mostly planned out quarters ahead and I don't feel like I have nearly the same impact.
And that is without considering that a full time role would keep doing more work than that singular time saving, or the decreased operational cost from failures in manual process execution.
Just because the jobs seem menial to you does not exclude it from bringing great value to an employer.
(Above numbers are subject to crude linear extrapolation, but are sufficient to make the point.)
>> "The growing zeitgeist is that “platform engineering is the future.” And given that I co-founded a product in the space, I sure hope so!"
I am so sick of these shoesalesman trying to get you to adopt their platform which runs on another platform which runs itself on wishful thinking and startdust.
Kubernetes can be deployed with a single binary nowadays there is no magic anymore the terraform scripts do look similar because you are not doing the actual hard parts like cis compliance or even selinux.
This whole discussion is so moot, that every time it comes up i feel a little bit sad inside. Call the things i do what you want, call us the "server-clowns" who cares.
It doesn't matter if its "platform engineering" "principle engineer" "devops dude" its the whole shebang of "guy with experience builds stuff that code runs on".
You dont get people who do "DevOps" fresh out of college or people that have been barely using a computer for some time. You get the dudes with experience, but every time someone with experience makes up one of these buzzwords like "platform engineer" "lord of the box" or whatever, tons of people just assume the role without having the experience. And then you hire the cheapest ones and wonder why you need 15 of them.
None of this is hard if you have been playing with these concepts for ages and finally that's what companies pay you for.
And having your entire statup on serverless? Give me a break. How can you tell that someone doesnt follow compliance or any other regulatory body in one sentence....
If you click on the username you will see that this person tried the same post 7 month ago and therefore i will say it again: Please let al bundy, the shoesalesman complain somewhere else.
I think you misunderstood basically everything about this post. The author is not saying to adopt the author's platform, the author is saying that beyond a certain scale, every org is actually building their own internal platform whether they elect to view it that way or not: the author is arguing the opposite of what you describe in this sentence.
> You dont get people who do "DevOps" fresh out of college or people that have been barely using a computer for some time. You get the dudes with experience, but every time someone with experience makes up one of these buzzwords like "platform engineer" "lord of the box" or whatever, tons of people just assume the role without having the experience.
did you even read the post? From the post:
> We have an industry-wide shortage of expertise in the cloud space, a great IDP should have safe, dependable building blocks for designing cloud services quickly.
The ... entire argument is that you need strong foundations and reusable building blocks, and those things take expertise to build specifically because most product engineers do not have or want that expertise.
Your whole argument is describing something the post isn't saying.
Ohh so devops is not bs? quelle surprise ;)
> The ... entire argument is that you need strong foundations and reusable building blocks,
Like the terraform code he himself admitted to copy and pasting.
> because most product engineers do not have or want that expertise.
henceforth again, devops is not bs. Therefore i conclude the intend of the post is nothing more than an advertisement piece.
Of course the post isnt directly saying "buy my product" by showing it in your face, but then again I have not said that either.
Its super easy though to click on the logo in the top left corner of the post to realize the site, on which this blog is on is directly advertising itself as "Defining Platform Engineering"
heneforth, yes i conclude shoesalesman trying to get you to adopt their platform through a linkedin advertised blog post =)
I bang this particular drum quite loudly, for sure. However if you'd indulge me as to _why_ I hope I'd be able to change your mind.
To start with, I'd like to state that while I'm not a fan of useless things, I can definitely make peace with non-perfect terms, systems and things that serve no purpose.
However, when it comes to things that actually cause harm: I can't abide that.
Why would "DevOps" as a term cause harm?; you might ask.
It's harmful because it is:
A) Rewriting history
B) Has no universally agreed upon definition
C) Is a term that's proginator can't even define in his own terms
D) Has no chance of ever having a universally agreed upon definition.
The deal is thus: if you have a devops team, some people will say you have failed devops.
If you have a devops engineer; some people will say you failed devops.
If you "practice devops" and you have no people focused on operations: I would say you've failed devops.
So what do we call the people who actually ensure operational excellence and best practice? Aha. DevOps.
I'm not trying to sell you anything, I just think this treadmill needs to stop.
I know it's unfashionable but hiring sysadmins into your project team and having them tied to your Agile methodologies is fine, sysadmins need to code though, as they always did. Heck, even Pat Dubois (originator of the term DevOps) originally envisioned the role as "Agile Systems Administrator".
You might think me pedantic or that I'm being needlessly overwrought on this topic; but precision in language is a very real engineering problem.
We must be precise with our terms, and if we can't even get our methodologies or job titles to align as an industry, can we really be trusted to get any other terminology right.
If you think I'm kidding about precision in language being an engineering problem I invite you to read: https://fileadmin.cs.lth.se/cs/Education/ETS170/IREB_CPRE_Gl...
Firstly, the idea that "DevOps" is rewriting history doesn't hold up well when you consider the natural progression of any industry. Concepts and terms evolve over time. It's not rewriting history, it's a sign of the industry maturing. Speech is in flux all the time, otherwise this post would also read like Shakespeare.
Second, the lack of a universal definition isn't necessarily a bad thing. In fact, it might be a strength. It allows different organizations to interpret and apply the principles of DevOps in a way that best fits their specific needs. This flexibility can actually drive innovation. Or differently set, the meaning is assigned and not a given. As a freelancer in the space I get a lot of different offers under the same umbrella, but the truth is that i can simply help in a lot of areas by definition.
Third, just because the originator of the term can't pin down a solid definition doesn't necessarily discredit it. It's common in tech for concepts to outgrow their original definitions as they are applied in new contexts and grow in scope.
Your point about DevOps teams or engineers signifying failure is also too black and white, in my opinion. The title is less important than the actual practices and methodologies being implemented. Whether we call them DevOps engineers or agyle sysadmins is secondary to their contribution to the project. That is what the job description is for.
However, calling this entire branch of IT "bullshit", will remain a marketing scheme for your underlying investment, no matter how much you now try to blame it upon technicality.
Yet words like "bridge", "ceiling" and "concrete" are unchanged for a century. Evolution is fine, but we're chocking on our own spit trying to be "next" and "hip" and whatever. So much silly useless effort, but this is worse anyway because we went from more formalised and easy to understand terms to a fuzzy one which nobody had clearly defined.
> Second, the lack of a universal definition isn't necessarily a bad thing. In fact, it might be a strength. It allows different organizations to interpret and apply the principles of DevOps in a way that best fits their specific needs. This flexibility can actually drive innovation. Or differently set, the meaning is assigned and not a given. As a freelancer in the space I get a lot of different offers under the same umbrella, but the truth is that i can simply help in a lot of areas by definition.
Sorry, no this is a super bad thing, you don't get to choose what "yes" means given a concrete context. This is not materially different. If you have a different job: it should be called as such.
> Third, just because the originator of the term can't pin down a solid definition doesn't necessarily discredit it. It's common in tech for concepts to outgrow their original definitions as they are applied in new contexts and grow in scope.
He never even supplied a definition, so we are stuck in this loop of defining it constantly. And our definitions almost never match. Lets just let it die. Almost as much time wasted as with timezones.
> Your point about DevOps teams or engineers signifying failure is also too black and white, in my opinion. The title is less important than the actual practices and methodologies being implemented. Whether we call them DevOps engineers or agyle sysadmins is secondary to their contribution to the project. That is what the job description is for.
I'm just pointing out that many people disagree on the definition, so many will just claim outright that you're doing it wrong.
I agree that actions are better than definitions, I'd rather have shitty definitions and functional actions; however that's a false dichotomy, we can have both.
> However, calling this entire branch of IT "bullshit", will remain a marketing scheme for your underlying investment, no matter how much you now try to blame it upon technicality.
I am invested in nothing, I have no financial or even sunk-cost gain from taking this particular fight, so you're concretely mistaken. FWIW, I am not the original author of the article; mine is here and was written a year earlier: https://blog.dijit.sh/devops-confusion-and-frustration
> "Sorry, no this is a super bad thing, you don't get to choose what "yes" means given a concrete context. This is not materially different. If you have a different job: it should be called as such."
Sure we do. It's a made up word for a concept. Especially when we are describing the widest range of possible work for a person working in IT.
I have a final point for you: If you really want to go down your own road, your job should have been called "Person who sets up the Kubernetes cluster with terraform, ansible and yaml". No platform is the same. No DevOps is the same. Henceforth we use umbrella terminology.
Distributed Systems Engineer
Cluster Administrator (the actual IAM Role, incidentally; weirdly not called "DevOps")
Container Platform Engineer
Deployment Automation Engineer
Or should I just call developers "IT"? ;)
> Distributed Systems Engineer
Probably best reserved to people that implement distributed systems. One can deploy and manage a K8s cluster with 0 distributed systems knowledge.
> Cluster Administrator
That's a good one. But many DevOps do more than managing clusters: I create and manage templates for our CI/CD pipelines, project templates and internal libs that abstract best practices, manage networks, databases, data lakes, identity providers... And develop internal applications to automate ops of all those resources.
> Container Platform Engineer
Same issues as the titles above.
> Deployment Automation Engineer
That's a good one, I like that. In the end of the day, everything a "DevOps" do is done to deploy things faster and more consistently to prod, as well as keeping deployments running correctly. I think we can even drop the Automation and use "Deployment Engineer"
> Distributed Systems Engineer
Not every distributed system is a kubernetes cluster
> Cluster Administrator
There are many services which should always run outside of kubernetes, loadbalancers are the easiest counter example. The IAM Role is called like that because its closely scoped to all things cluster in amazon.
> Container Platform Engineer
entails that you only work on the container platform, while a cluster admin for instance can setup everything on VMs or bare metal
> Deployment Automation Engineer
This sounds like you are only taking care of deployments, monitoring and logging come to mind as counter examples
> Or should I just call developers "IT"
I tend to call myself Freelance Sr. Engineer with Kubernetes / GitOps / DevOps / Golang Focus in order to bring german exactness to a widely defined field. So depending on my hat, I become what's necessary in that scope. But yeah, to most managers, I am simply someone from IT
Companies have a goal in mind when they go out an hire staff for their objectives. Overall I still think that DevOps is neither b*s* nor insufficiently defined to yield a job title. Titles are vague on purpose, because especially in the realm of DevOps your scopes will get extended to more things than just terraforming.
Sorry for being that guy, but it's "moot" not "mute". A moot point is a point with little relevance to the discussion at hand, while a mute point would be a point that is either unheard, or made using mechanisms other than speech.
Would people read this article less critically without knowing the author's position?
If so, then THAT'S the problem with devops (and every other software buzzword). If someone is reading online about ideas in their profession, why should they ever treat the ideas with anything less than maximum critical thinking and distrust? This is their livelihood, after all.
Ignoring all of the other changes it brought into the industry (the fact that CI/CD is table stakes now, for example), the single biggest thing that the DevOps culture instituted is the _idea_ that ops will need to code sometimes and engineers will need to _ops_ sometimes.
Sure, many engineering orgs translated this into creating a "DevOps" team that writes YAML and manages The Jenkins™. (At these places, this is, sometimes, the ONLY way that breaking silos can be done due to complex organizational rules. Also, many of these teams are composed of ex-sysadmins AND ex-SWEs, i.e. the literal embodiment of DevOps.)
But it also brought about operators managing high-complex software like Kubernetes in even the biggest of BigCos.
This simply would not have happened without the DevOps mindset taking hold.
For example, from the article:
> For every operations person without software development skills, there are FORTY engineers without cloud operations skills. If you are going to build an internal platform, you’ll need experts with overlapping experience in both fields working together.
Yes. That's literally what DevOps _is_.
DevOps as written is not what you're saying; this has been mentioned at least 20 times in this thread.
Sysadmins of yore were developers, they used perl and C and eventually ruby. Then there was a glut of sysadmins who couldn't code, then people acquired collective amnesia that sysadmins ever coded and invented a new term.
The issue was that the term that they decided to use for this new breed of developer operations folk was the same word that was used for a methodology of breaking down siloes (having developers and operations as distinct entities but working together)[0] and the same word that was being used to describe sysadmins adopting Agile[1].
The annoying thing about the passage of time is that you look back and see only what has changed and may conclude that those changes are a result of paradigm shifts: you could be right, but sysadmins were using proto versions of common paradigms now: version control integrations used to be bad and fragmented, but it was ops pushing VCS; infrastructure as code via cfengine; cluster management with clustered ssh clients (cssh).
I was there, I saw this.
It's hard to differentiate between the slow march of progress of operations tooling (rebranded as devops tooling) and the fact that it might not actually be linked to devops as a methodology.
[0]: https://www.youtube.com/watch?v=LdOe18KhtT4
[1]: https://www.knowledgehut.com/blog/devops/history-of-devops#
I really didn't. One of the definitive books on the DevOps ideology, _The Phoenix Project_, spends basically its entire plot driving home the point that DevOps is about dev and ops working together instead of ticket-passing.
To your point, doing this requires going back in time a bit where sysadmins used to do a little bit of software and devs had to be a little bit of a sysadmin.
(I blame Windows and the dot-com boom for the silos having been erected in the first place.)
That book (I've read it twice, I think it's the only book I read twice honestly) is about the Toyota method and Agile and now that I think about it: warns against having "super stars" who do everything; which seems to be some other folks definition of devops xD
Now enter the snakeoil salesmen and the agile consultants. Throw the work devops around. It can mean whatever the fuck you want. Devs? They cannot be bothered with operating the service. Look at Google! SRE!!! Ops? Are the addibg any business value? I mean c'mon they are "running" the service!
I know, i know: devops team! A team where we do devops stuff! And we'll replace the operations team (secret time: we will not. It's the same shit with a new name). Now we just wait for all that agile fairydust to kick in.
The well may be poisoned, but the original meaning is still perfectly well alive.
It was 100% career-bullshitters within the org wanting to make a name for themselves by "introducing" something that the clients wanted and was a "big gap" for the consultancy. They latched onto this buzzword, and at the same time it was combined with clients latching on to it as well from all the hype it was getting in conferences, talks, articles, and other places where bullshitter evangelists push their drug. Never mind that in every sales meeting with clients we "promoted" this buzzword - so no wonder they "asked" for it.
Had one of my projects burning for a year as a result, where we were told by leadership to rely on a "platform engineering" team that had a ready-made platform. Meanwhile it was just a bunch of arcane priests that knew some awful-looking JS/JSON inside a YAML templating language, running on some weird platform tool that manages K8s. Found out about the mess when I inquired after a month where this ready-made "platform" was, and they said they were still "bootstrapping a K8s cluster in AWS", then waited another 2 months for them to "bootstrap" the QA and production environments. Meanwhile the platform engineering team was led by a career-bullshitter and they had just a bunch of guys that basically did some bootcamp courses, certs and minor experience in "Their Chosen DevOps Tech".
Leading to such wonderful talks summarized as: "Oh you want to explore the Hashicorp stack and weigh pros and cons? Fuck you, it's a poor-man K8s, we know better, shut up and stay in your lane."
Never again.
From that point on, this function 100% stayed within the dev team's control. Architecture, platform-ing, SRE, devops, whatever you want to call it.
My senior full stack developer LinkedIN requests are pretty low volume for some months now, while he gets flooded with offers with compensation more than most senior devs I know.
His skill set is also (and I love the guy dearly, I'm trying to be objective and he also feels this way) maybe 1 / 5 of what a senior dev has. Basically he's "I’d wager most of what they are doing is using Terraform and YAML to do menial tasks for the engineering team."
EDIT: One thing I forgot is, that in order to be really good at DevOps, you need to know your stack _really_ well, which imo necessitates that you spent at least some time on the development side of things. This goes against every hiring managers instinct though, as people should only specialize in one thing and do it well (at least where I live).
Oh god I thought that was unlucky. I have huge respect for Terraform for popularising what they call “infra as code” (we used to just do this with the actual AWS API) but now there’s multiple modern IaCs that use actual code with real tooling.
I don't use OpenStack but there is a Terraform provider to configure it: https://registry.terraform.io/providers/terraform-provider-o...
It’s highly compensated because it’s very needed. A low low % of operations engineers are as good as they need to be. It’s not even always the organization’s or the engineers’ fault, either, as the author of this blog alludes to.
Picking a horrible position just because of a better salary is a great way to live an unhappy life.
Life is not only about work, I have to work to be able to pay my mortgage and to enjoy other activities.
Having happy or unhappy life, for people in the IT industry, is mostly about their mindset.
And we have HYPE all over the place: cloud services, frameworks, programming languages, tools, application architectures, etc.
And to my eyes DevOps is another hyped thing. Not only that, many jobs require you to be "certified" in a cloud service provider (a la Cisco like in the 90's) or on some particular tools (Kubernetes?).
Then we re-discover that the OLD way was really good, that actually there was no need for something that over-engineered, and we go back and start doing things the simple OLD way. After some time, the cycle re-starts.
What is DevOps anyways? I don't know, but today it seems that is a word associated to fancy cloud services and tools to deploy applications, way way different from the savvy system administrator from the 90's & 2000's that knew lots core concepts as: DNS, LDAP, SNMP, HTTP (Apache), SMTP, NFS, iptables, etc.
If I had to hire someone for today tasks I'd probably stick to the 90's knowledgeable system administrator.
It certainly helped with the social integration of software developers (which is definitely a good thing), but it also created this bizarre connotation that tech had to be cool.
It doesn't. Technology first and foremost has to work. Especially software.
I should add that it's now becoming even worse : tech now has to be politically correct and inclusive ( aka: boycotting a database technology because its author don't support some US legislation regarding transgender rights would be a totally ok decision)
Ai, ai, ai, generative ai!
Product teams need somebody who has the following skillsets: security, observability, deploy/release, containers/repeatable environments, networking/runtime/storage... if you don't have those skills on your team, you're going to have a hard time. Sometimes you find those skillsets in one engineer, sometimes among multiple people on the team. If they're in one engineer, maybe that guy's title is DevOps, maybe it's Platform, maybe it's SRE, maybe it's something else.
Stop sweating the title and start hiring for the skillsets you're missing. If you can't find anyone, find people on the team who are interested in learning, and give them time to train on it. Who cares what other companies are doing, worry about the missing holes on your own team first.
In my my opinion Microsoft came pretty close with Azure DevOps.
About half of the comments are people saying their there exact definition of "devops" is the one true one ( Handed down from Heaven by Patrick) and if people followed that then it would be unicorns and moonbeams
Another third of comments are Developers who dislike Ops and especially that Ops can't code [well] and would seek to eliminate it. They hate that Ops people keep re-appearing with a different name each time. They will lament that Sysadmins always used to be able to code "back in the day"
Some of the remainder will be people (like the author of this article) noting that Operations is a distinct skills and that most devs don't want to do it. SO rather than trying to eliminate it perhaps we could actually do it properly.
- ops is sufficiently complex - The incentive to deliver features means that ops corners tend to be cut
So the suggestion is to build a self service platform that encapsulates the complexity, while making it easier for developers to use.
Much like with the agile debacle, the best thing to do is continue working and try to avoid/minimise the harm they can do until they find the next idea to latch onto.
Incidentally, this video popped into my head as I always think of it when this comes up:
https://m.youtube.com/watch?v=0qFGs-SIWB4&pp=ygUeRGVlcGFrIGN...
A lot of Windows people disliked the concept of Sysadmins because they didn’t know how to program and tried to paint the concept as elitist.
A decade later someone announces the concept of DevOps, which is an infrastructure engineer that knows how to code to automate everything.
Whatever. it’s not worth getting worried about. At least automation is being considered important again, we have infrastructure as code now, the languages are better than Perl, sendmail is dead, and as long as you stop them installing kubernetes unnecessarily (which they will try) they’ll be fine.
… except that’s not what the concept was.
… on a metal server. In Bash.
But what do you think software runs on these days? Magical immaterial containers? Bare metal is still there, just a few (more or less terrible) abstractions away. And please check out https://github.com/kubernetes/kubernetes - Oh noes, k8s is 3% bash? The horror :))))
My point is that today sysadmin skills are not enough. You need Python (boto3), APIGW and Lambda, years of Cloudformation and stuff. Does not hurt tho :)
There is no DevOps role and never was. It was just management twisting it into marketing scam to get devs operate systems and to feel like it is their responsiblility.
Problem is that no company wants to have Sysadmin tied to one dev team because sysadmin won't have much to do most of the time - but that was idea of devops that you should have devs, admins, qa, business analysts working together reviewing tasks from the start and giving input on what is their specialization.
Well yeah no true Scotsman - but dev ops idea is not bullshit - way people do things is bullshit.
The idea of DevOps is to my knowledge that the role of Developer and Operator slowly merge together in one role. Which would also be in line with SCRUM and agile development, i.e everyone should at some point be able to do everything in the stack, if it's necessary.
What you're describing sounds more like the standard agile development process.
That is certainly not how devops was originally envisioned.
There are two major sources for DevOps:
1) the 10+ deploys a day talk from Flickr: https://www.youtube.com/watch?v=LdOe18KhtT4 which is where most of the actual methodology comes from.
2) The "DevOps Days" (first use of the portmanteau by Patrick Dubois) which was an agile conference aimed at Systems Administrators trying to get them to work more agile; the conference was called "DevOps" to make developers think it was also for them.
Neither of these primary sources said that we should erode or replace ops people; or that it was ever intended to be a role or individual responsibility. This is a lie that corpo's have made to reduce headcount.
I've been a sysadmin, then a DevOps, now a Dev. I loved system administration the most but then it got a bad connotation for management, recruiters, etc. Like sysadmins did not do automation or wrote code. So the jobs went away and then I got jobs writing YAML and using whatever alpha version automation tool was hyped at the time. So in the end I said fuck this, I'm tired of trying to put garbage code that others wrote into production using what they (who had no Ops experience) think makes Ops work. I want to write my own garbage code and let others use whatever born-yesterday YAML templating engine or DSL is trending nowadays.
As an Ops guy the only way to win is not to play.
> Need a database? File a ticket with DevOps. > Need an IAM role? File a ticket with DevOps. > Before long, your massive team of engineers has fully saturated your understaffed “DevOps” team’s backlog.
Our DevOps is under IT, not Engineering, so there's already an organizational gap there.
I've talked with DevOps multiple times over _years_ and laid out a rough architecture of how Engineering would like to work. To this day after CI/CD passes its still manual deploys with a ticket.
Render.com and other platforms do 90+%, if not 100%, of what I want as an engineering manager, but I'm hamstrung by policy and politics that were put in place 6 years ago when non-technology ppl decided we somehow had the resources to manage an OpenStack build instead of AWS. And that's before I realized everything on AWS is config config config and there's entire companies (like Render) that take away most of the work and just give you a dashboard to accomplish what you'd like to accomplish.
Know how to say no? Awesome, that’s what you need to tell the devs whenever they moan about not having fancy things like metrics or logs
(Sorry just ranting about my current gig)
Many of them are fresh out of code bootcamps. I'm uninterested in subsidising poor education an laziness. Ops is hard. Computers are hard.
The cost of those platforms (the borgs, the chubbys, the map reduces) is stratospheric compared to what a startup should be focusing on.
If you want to go somewhere new (as most startups do) you're not going to find a paved road to bring you there.
You can ignore architecture but it won't ignore you.
This part, while true, bothers me:
"Simple, just hire some more frontend and backend engineers to develop a great internal PaaS with all the golden paths your operations team is architecting while you are trying to build your actual product that runs on top of it."
I've been on that team, multiple times. It doesn't really work, because it's yet another silo, and it takes years to bear fruit. We shouldn't be developing platforms, we should be buying them.
It's like building and running a factory so you can have an assembly line to make widgets. Most people who want to make widgets today actually contract that out to a company in China who already runs a factory where they adjust their tooling to make your widgets. That's what economies of scale and specialization do: you pay somebody else to do the hard, expensive thing that isn't a core part of your business. You get to market faster, you waste less money, and carry less risk.
We're doing business wrong. You should not be building factories, you should be hiring a factory.
Do I mean "hire consultants?" No. What we need doesn't seem to exist. We need companies who do white glove integration of a software product with a bespoke PaaS. Like Heroku, but with more attention to architecture, monitoring, alerting, security, ci/cd, etc. Someone who will give you the tools, train your people, do integration alongside them, and provide support. And they could do this on a cloud account that the customer owns, with IAM guardrails to prevent them fucking up the platform.
There is no value added to your product by building a factory. You might reduce cost, eventually, if your business is huge enough. But there's a small number of such businesses out there, and yours likely isn't one of them. So let's start some platform companies and make those the default choice.
Every company has a very different culture and set of tools/skill sets.
It can be done for simple web apps — that’s sort of what heroku did and I loved heroku.
For a whole company integration that is a bit more complex ? I don’t think so .. maybe one day we will have a tool like sales force and platform engineers will be customizing it instead
Another one is OpenStack. Did you know there are web hosting providers who will give you your own OpenStack install and charge you based on only what you use? It's a cloud provider, but again, it's one set of tools that everyone can use, regardless of size.
All I'm saying is, form a company that specializes in these technologies, add a shit ton of documentation and web UIs, and also actually write some code, tell the customer how to change their app to work better, etc.
There are already well known DevOps consulting companies that do the latter, but they are not advertising themselves as a full service platform. I think they should, and I think we should use their platform, rather than build our own per company. Like the tool ecosystems mentioned above, there is room for more standardization, while simultaneously room for people to tell your software devs how to use it right. There's no shame in changing your company's culture (or setting it from the start) to align with DevOps principles. Might as well be the culture and principles of people who know what they're doing :)
I think one of the reasons this doesn't exist is that you are describing every single component of the software development process here.
I have a personal theory that software developers have developed severe tunnel vision about code in the last 10-15 years. Before then, most developers had some knowledge about deployment practices, shell scripts, apache, that sort of thing. They roughly understood how something like a web server worked; writing a toy one was in many CS curricula.
Whereas now software developers' number 1 skill set seems to be "I know how to solve complicated algorithmic problems". Which is fine and dandy, but it takes so much time to master that it drives out all other deployment basics that many of us acquired while tinkering with software. I am talking bare minimum here, by the way -- I have worked with developers a year or two out of college who did not know what a shell script was because they had spent their entire degree inside either an IDE or Leetcode.
This, combined with the fact that complexity has increased exponentially in the operations tooling space means that you sort of need knowledge about operations knowledge in your software engineering team along with complicated algorithmic problems. You could do this by hiring some very expensive consultants like you describe. I think you could also do that by altering your hiring pipeline to de-emphasize algorithmic problem solving, and increase the importance of ops knowledge and testing.
You're describing a period (well, not far off - maybe a little earlier) when the norm was to chuck a .war file over the wall to a sysadmin team who'd plug it into tomcat, and who absolutely definitely wouldn't give you access to the server and you had to ask very nicely to even see logs when it broke. The tunnel vision has always been there, as long as the hiring environment has thought that specialism is a good thing.
What's needed is for people to do more time in tiny orgs where there's no room not to learn how everything works, and where there's no space for complicated solutions that don't fit in one person's head.
It's interesting how differently we discussed DevOps back then.
EDIT: also ThePrimagen reacted to this article yesterday: https://www.youtube.com/watch?v=qVEEpUvl0Kw
Application developers ought to have a decent working knowledge of ops such that they can work productively with the ops specialists and vice-versa. Just as application developers ought to have a decent working knowledge of product and UX design in order to work productively with specialists in those areas.
And Amazon has a vested interest in making The Cloud as complicated as possible so that you wind up getting suckered into vendor-lock-in and can never migrate away.
Combined with all the overly complicated crap surrounding concepts like CI/CD, K8s and the whole CNCF landscape with a couple thousand different vendors trying like hell to get their tentacles into your organization.
If platform engineering can fix it, fine, but I'm not holding my breath.
Platform engineering = building black boxes (abstractions) to reduce Dev team mental load at the expense of nobody understanding how to fix issues
Couldn't this be caught in CI relatively easy? Either CI is not running in Docker or the logic wasn't tested
what? Why?
> so the team leans towards asdf and a README full of stuff to copy, paste, and pray.
It's really not that hard to run stuff in containers. Use a Makefile or similar so you don't need to type long Docker commands
> During a sprint, an engineer adds convert (ImageMagick) to the mix to support manipulating images and forgets to update the Dockerfile, and then production goes down.
And what... you can't run integration tests in a containerized environment in CI?
Don't get me wrong, I'm all for efficiency and the like, but if I were a mere "programmer" or "sysadmin" or "webmaster" I don't think I would earn as much as I do now that I'm a "platform/software engineer". Words matter. If you're called an "engineer" (regardless of whether you have a degree in engineering or not), then business people will think something more of you than if you're just a "programmer". Same for "platform" and "devops". Same for "data scientist" and "statistician".
You think knowing K8s justifies the $150K/year paycheck in your startup that has only 20 "engineers" and merely 5 services on production? C'mon. But here we are, and I'm more than happy to get those $150K/year for applying K8s and Terraform when it's not actually needed at all.
Posts like this one only reminds us that we live in a time where people is making money out of hype (scrum masters anyone?)... but as long as I'm making money as well, I'm happy to join the conversation and play the "let's fix this!" pantomime.
Edit: This allows the devs to focus on the business problems and solve them within the stack instead of spending time fighting the technology, it works great.
My experience has been that the places that treated development roles as cross-functional have been the most efficient, with the least amount of politics and dysfunction. It removes a whole cadre of people who are barely able to describe the features they request, while minimising the space for power plays.
1. Any task is an API away and everything except getting a Cassandra keyspace is self-serving. And getting Cassandra access is a matter of a 2-minute conversation.
2. There is always a UI for casual tasks, such as setting up a service cluster or CI/CD pipeline. This is crucial as engineers do not have to learn any non-transferrable shit. The UI will guide them to complete their tasks.
3. Powerful support of observability. Clients, servers, OS, networking... All kinds of metrics are just a finger snap away. In addition, adding metrics is dead easy. If we need something that is below the service level, such as SNMP, we just dropped an email to the four-person observability team, and the metrics will show up in a matter of day before I even had time to finish my own coding. Heck, we had full support of OLAP query on millions of time series with full support of boolean search by any tag (even metric name is a tag) back in 2011, when most people thought Graphite was the best thing since toilet.
God! I miss so much the Netflix days when everything was a non-event. Deploying a predictive autoscaling service to any application in production? Non-event. Getting Druid to production readiness in less than a week? Non-event. Having user-controlled fault injection via APIs? Non-event. Setting up Mesos and Zookeeper and opening up Zookeeper for the entire company to use in production within a week? Non-event. Creating a company-wide crypto service and using it to swap AWS keys automatically for all the Netflix services in production and in test? Non-event. Generating an event with a new schema without talking to anyone and see the event show up in a brand new Hive table under 15 minutes. Non-event, and it was back in 2010. Setting up eBPF traces and access all the pretty and dynamically updated charts for any server via an API? Non-event.
I don't know what more I could possibly ask for.
My opinion is that the platform must be:
Not built upon k8s. K8s is too complex by itself, too many moving parts, too many failure points. And everything built upon it makes it more complex. If you have team of devops, that's fine, they got work to be paid for. If you develop alone, that's a problem.
Leverage containers and docker. Docker is a thing that everyone knows, anyway.
It should be a ISO which runs in memory. Talos is a perfect example of this approach.
It should solve distributed storage automatically. You run this ISO on few computers, their local disks become a distributed storage.
It should provide overlay network for containers automatically.
It should have lots of low-level features built-in and without choice. I hate to choose between calico+istio vs cilium. It should not happen.
It should provide all the necessary metrics, centralized log storage, easily installable CI/CD pipeline, it should be resilient by default against crazy containers eating all memory or CPU or network or I/O.
I guess, it could be built upon kubernetes, but it must be a very well chosen set of components without a way to replace those components and it must be extremely resilient and well tested set of components without any nonsense. With some reasonable UI dashboard, so nobody would need to install those grafana monsters, for example. With built-in OIDC, so no need to spend months installing keycloak and hooking everything to it.
I saw some "platforms" built upon kubernetes, but those were duck-taped monstrosities which work as long as you pay their authors for 24/7 support and overvision.
For example, via a few docker containers & an nginx configuration pointing to those containers as upstream servers.
https://www.redhat.com/en/topics/devops/what-is-blue-green-d...
If you don't want to do it wrong also, prepare to be persecuted and insulted.
I'm not in a big corp, we are a rather small team, but we do all of our operations ourselves and it doesn't appear to particularly suck. Sure, the more the complexity of a system grows, the more you might have to split up responsibilities so that not everyone does everything... but not every business is giga-scale and needs kubernetes, at least not initially.
Then again, I also very rarely have deadlines and don't work at nights...
DevOps is fine, bad DevOps is pointless. Just like microservices are fine, but bad microservices are pointless. Same way bad metrics towards business owners causes you to seem to add no value (which goes both ways).
That said, you can still find plenty of job postings for 'DevOps Engineer', which is part of the problem. But I suppose culture changes are hard...
It feels like the need to define operations is a bit like if a dev role needed to be defined by IDE or language How about visualGolangDev or categorize devs after what kind of code is being written microserviceDev vs backendDev vs middleWareDev In all honesty, can we just go back to ops ? It was a fine name. Said it all and running k8s , a unix server or a ci/cd pipeline is essentially just operations.
What is wrong with ops ?
It will happen when the optimization of the autoscaler will improve the revenue, by itself and in a quantifiable way.
Otherwise, a fast autoscaler is just a small step in the right direction, as elegant code, clean financial accounts and free strong coffee are: all of them support the business, but it's the sellable functionalities that create the business in the first place.
And KISS & YAGNI should apply to methodologies just as much as software systems. Don't turn simple things into rocket science just to fit some management fad.
Have "process meetings" with everyone, collect feedback and discuss areas to improve. Be an ear first, mouth second.
In my current current company, the operations group’s name changed from “techops” to “platform”, … but there is no “platform” or roadmap towards that, and it is reflected in how people still call the group “techops”.
That would be a bit like expecting your house builder to also architect your house. That's not going to happen, but they need to be able to communicate and understand each other.
That's why tools like Ansible or Terraform, and the whole Cloud Computing emerged: so that infrastructure can be abstracted away, and you, as a developer who wants to deploy their code to production are only required to create a couple of simple yaml files, trigger the deployment pipeline, and be done with it. Ops roles are outsourced to AWS or other provider.
I know, that's just a theory, real life is nowhere near that ideal state.
this dream of self service devs is becoming a myth. devs do so much incredibly thoughtless stuff to infrastructure all the time. you need an infrastructure person in the process to make it resilient and allow them flexibility to do the stupid shit they want to do - with guardrails.
If you deploy something, it breaks, you need to be on call to handle it. If you're on call, handling incidents, you need to be also building stuff.
DevOps is not a role is a practice. That article is really bullshit.
The thinking and upside of this was that quality wasn’t an externality. If your code sucked, then you had a bad experience with ops, and you were motivated to fix it. Where that could fail is when managers didn’t have skin in the game- it’s important that managers have to feel some of the ops pain personally to prevent psychopaths from destroying a team just so they could look good schedule-wise.
Note also that “tech debt” either became an ops problem, which self-corrected as described above, or it became an agility problem, where it took forever to make changes to code. DevOps as I saw it didn’t have a good answer to the latter.
We had lots of different teams for different projects; having one team know everything about everything wasn’t a thing. Perhaps the problem is that DevOps is not well suited to smaller or less differentiated organizations.
This approach scales because it doesn't create a queue in front of ops.
When recruiters come by and say "I looking for a DevOps engineer" that's a big red flag: DevOps is not a job, it is a philosophy to be applied to all related jobs within the software supply chain! As idiotic as if people come by to present themselves as "Agile Engineers" or "Scrum Engineers"...
However: I do not think the DevOps philosophy is bullshit, far from it. Not perfect, like Agile methologies, meant to be bended and adapted. But once a month I hear stories from legacy ITIL-based organisations, which are point-to-point mentioned pain-points from "The Phoenix Project"...
The companies I've worked for which grokked DevOps didn't have dedicated DevOps teams. They had developers who understood automation embedded in the project team and the DevOps component was just one small piece they were working on. In companies which struggled, they setup separate dedicated DevOps teams and hired explicit "DevOps engineers" which you'd submit tickets to in order to get your pipelines built and changed. One of these enabled the teams to accelerate development and releases significantly, and the other just felt like you had a traditional ops team that was more of a barrier to getting work done.
I view building software delivery teams similar to (in my imagination) building a "special forces" team. You cannot afford for your team to be taking on a building of terrorists only to come up against a locked door so your team shrugs and says "guess we'll have to open a ticket with the door breaching team". You've got to have the "specialists" you need embedded in the team or your "mission" will grind to a halt waiting to borrow resources from teams with different priorities.
If you know what you are doing you can actually minimize your cost and complexity a lot and get away with not having to actually do a lot. I am a startup CTO these days. That means I'm not bored. Devops is not a person, it's me in what little time I can budget for this among many other more interesting, lucrative or tedious things that I also need to take care off. I need my servers to run and not die on me. I need decent uptime and build automation. I need backups, alerts, and all the rest. I've seen it done well and I've seen it done poorly. I don't like it when it's done poorly and I know what that looks like.
If I had the budget, I'd probably get a person to delegate to and invest more in this. But I don't. So, I go for minimalism and reduced complexity.
So, good bullshit free devops is what I do. No mico-services in my company. We have a monolith. That's not likely to change any time soon. We can't afford the operational complexity and I frankly don't see the added value of having to have that complexity. I don't use kubernetes for the same reason. Using kubernetes to run 1 docker container is nonsense. We run nice simple docker containers. Most common cloud providers have easy ways to scale and run those. I used google cloud run for a while (cheap and easy), but our use of websockets and threads made using traditional vms and a loadbalancer a bit more flexible. That's what we've had for the last two years.
I don't need a teraform script to create that for me and I just clicked that together in a few minutes. My infrastructure is fine. I do use build automation. We have a few lines of gcloud commands in there that tells the instance group to apply a new template with our latest docker container. It runs when we merge to our production branch. We use Github actions for this. Our master branch goes to a similar staging environment on every pull request merge. That's simple devops.
We also have a few managed services, alerts, monitoring, logging (via elastic cloud), and few other bits and bobs. I'm not a caveman. But you get the idea, this is not something that we need to spend a lot of time on. Set it up right and you can do more fun things. Our uptime is fine. I can scale this endlessly.
I can also setup a new environment without too much hassle. If that actually becomes a regular thing, I'll automate it further. But that just isn't a thing right now and not worth a minute of anyone's attention in this company. Automating things you only do once every few years is a waste of resources.
sysadmins: 2010
devops: 2017
platform engineers: 2024
Wonder what the next term will be for developers who understand operating systems and distributed systems holistically.
Afaik the term was coined at Netflix https://netflixtechblog.com/full-cycle-developers-at-netflix...
(Full Stack, a misnomer, is already taken)