Operations is not Developer IT
matduggan.com
matduggan.com
1. Application developers are your users. If we application developers took offense every time a user tells us that things are not working, we'd be pretty pissed off all the time. Educating and empathizing with your users is part of your job.
2. Talking about how it was better before: QA teams sure do buffer a lot of crap. They also cost a bunch and slow down time to release. Yes agile is causing problems. The bureaucracy and stiffness of organizations before agile was no nirvana either.
3. By your own affirmation you treat applications as black boxes that should be deployed using a runbook that should just work. This is ridiculous. Application's ownership is shared between everyone who works on it.
4. And yes, as developers, networking or physical drive space are things that we tend to abstract away. Maybe if the infrastructure people were involved in development discussion earlier, they'd be able to raise their hands and say: wait a minute, you're going to blow up our logs.
This all feels like someone who used to not do anything suddenly being asked to take part in what's happening...
EDIT: apologies for the strong language and sounding like an asshole myself, but I certainly feel irritated when someone takes the time to write a 5000-word article complaining about whiny developers who thought they could own ops but actually don't know anything and scream for help when they themselves are the cause of all evil.
I think both sides can do more to reach into the domain of the other. I get it- we don't want to deal with blinking lights and they don't want to deal a missing semicolon breaking everything.
Honestly I think "that's not my problem" is one of the worst attitudes you can have as part of an organization with common goals.
Different job roles exist for a reason. In the case of developers and ops, developers develop, ops manages ... well, usually literally everything else, and there should be _some_ overlap in the middle.
If you can't say "no, I can't do your job for you", then that area of shared responsibility just gets wider and wider and, if one side _is_ doing that (as developers usually won't step into the ops end too deep), it shifts further and further to one side. In the case of a typical organization where you have maybe an ops person per dozen or two dozen developers, that person very quickly becomes a bottleneck. That person gets burned out. You need to hire a bunch of expensive ops people to do work cheaper developers could be doing.
Literally watched the lack of role definition and a bunch of ops people that, by virtue of almost everything being their job, won't say "not my job" do this at a company I'm leaving. A couple months back I was literally writing code in one of our apps because the team that owned the project "didn't know elasticsearch and didn't have time to figure it out".
Hey, that's a pretty shitty way to start off a comment, don't you think? With a personal attack?
1. Yes. But it's not operations problem if you are whining that your PS5 game isn't running on the XBox. There is personal responsibility in this, too, and it's not operations job to hold your hand and explain how to do your job. If you aren't reaching out to operations to make requests, they aren't going to know what to do. Your entire comment shows that you think they are subservient to you, rather than you actually being an honest user. Tell them what you want, and work with them to get it.
2. QA teams do not slow down time to better quality releases. They do slow down time to half-baked or buggy releases. Regardless, the number of app developers to operations people is generally a very bad imbalance. I promise you, the good ones are working with the people that reach out to them.
3. Maybe if you invited the operations people earlier, they'd have some ownership in the product. But usually they release it without operations even knowing, and suddenly there is something in production that is half-working. They had no hand in it. They literally did not work on the project, so they can't know.
4. You can't abstract away things if you don't know how they work or account for them. Again, inviting operations people to earlier discussion is incredibly easy. You know what projects you are working on, they tend to not because there are far fewer of them than there are application developers. So, it's on you to reach out to them to get input. Yes, they have to make themselves available, but you have to invite. And guess what? When you do that, you get a wealth of information and makes the product better.
Your comment feels like someone who is used to expecting perfection from others while accepting their own mediocrity.
Wow... ending a comment with an insult is rather shitty, too. Why did you decide to go the route of writing a comment that starts of shitty and ends up that way?
Personally, I did it to hold up a mirror to you.
"Maybe if you invited the operations people earlier, they'd have some ownership in the product."
Awww... You like each other but none of you dare to make the first move :D
My experience has been that they can be very accomodating and supportive if you do talk to tham.
There is always a project for them to offer help with, since the business will not suffer devs to be idle.
Anyways, Ops have their own responsibilities and none of the core ones are centered on devs.
No, that's not smart.
IMO this is main topic of the thread and of the article.
There are groups of people who instead of spending time to figure out how to work together and understand what other side has to say, they just throw shit over the fence.
Maybe some could start by reading points at least couple of times and try to understand instead of trying to write personal experiences as fast as they can in reply to other comment that hurts their ego.
But keep blaming the war on peons.
Maybe it works if you assume all employees want to do bare minimum and don't get blame for what was not delivered.
What I see most of the time is that people want to deliver, people want to be valued by their work.
Of course I am cynical as the next guy from me in terms of "getting on high horse" but there is a lot of people who want to do their job and want to do it well.
Playing divide and conquer, playing non-existing scare deadlines is going to work once or twice and any smart employee will leave after that kind of crap. Other option is you are going to get smart employees who cannot afford to leave, but because of that crap they will just stop giving any fuck.
I sign up in reality into "self fulfilling prophecy employee", when you treat your employees or other people with expectation that they are thieves - in the end they will steal from you.
If you treat your employees as if they suck - they will suck.
Of course there are bad apples but if one goes the road that everyone want's to rob him, he will get robbed.
Only if you ignore reality. If you hold a meeting and don't tell me about it, how can I attend? Both sides want the ops people involved. Maybe the one having the meeting should invite them.
1. It's not operations problem for sure, but I certainly don't bash people for not knowing things I am the expert of.
2. Fine
3. The OP's saying he doesn't want to know!
4. Well, writing applications is sitting atop a stack of technologies more and more abstract. A developer not knowing what happens in an IP packet is the same as an infrastructure guy not knowing what happens in an NP junction.
I don't care if devs understand IP packets, TCP congestion control algorithms, or anything similarly low-level. If they do, that's awesome, but it's not expected. I do expect them to have a basic understanding of expected latencies for intra-DC vs. internet, why running Flask in production isn't a good idea, and if they're really sharp, an inkling of how Kubernetes networking works.
the flask build in webserver is not production grade software in my opinion.
> WARNING: This is a development server. Do not use it in a production deployment. > Use a production WSGI server instead.
I don't expect junior devs to have a sense for what is production-grade and what is not, but if they try to ship software that explicitly warns against being used in production, you've got a real liability on your hands.
Why do we call these people "software engineers" if they're not engineers. There is endless documentation, there is all the resources needed. Allowing devs not to know what they're doing is a complete failure of anything approaching professionalism.
There is a legitimate point buried there, but I just kept seeing red reading it.
So anything said before that happens, and it'll happen for you partly especially at first, is only going to enrage.
There is no attack in the text. There is a complaint that issues presented to operations often lack the basic level of detail and due diligence that they should have. You are free to disagree with the author's expected level of due diligence on issues; I think you'd be wrong to, but you can. However, it isn't an attack.
You perceive a non-attack as an attack, and respond with an explicit attack and name-calling. That actually makes you the aggressor.
Hmm, who is the asshole here?
Uhh... can we chill with the personal insults a bit?
"Ooops... deployment failed. While deploying your artifact we found the following:
- Nothing is listening on the nominated port
- Your deployment is utilizing 100% CPU while idling
- We detected an abnormal volume of write operations to the mount
Please fix these issues and re-trigger the pipeline at your earliest convenience.
Regards, Ops."
Now that just shouldn't happen... ie, we(ops) aren't going to deploy something that doesn't come with healthcheck(s). The healthcheck never passing(port isn't listening) is going to stop the deployment from ever completing. Ops job is to push back on developers if they try to hand us something like this to build a pipeline for. In my company, to hand Ops the name of a repo and say "build a pipeline"...there are a lot of requirements, and the biggest one is a list of SLAs. That list of SLAs is how we build monitoring for your application, and one of those should always be a list of port(s) and protocol(s) that are exposed; we build monitors against those.
Being a human kubernetes seems to be the crux of it.
In my company, to hand 'ops' (kubernetes) the name of a repo and say 'build a pipeline', it's basically a matter of committing a gitlab-ci.yaml file.
A better way to put it... most of my colleagues in an Ops roles already did development for 10-15 years, and moved on to developing the tools to deploy other people's products. Additionally, kubernetes isn't everywhere - they also build pipelines that produce AMI's, GCP images, and write the terraform/cloudformation/HEAT/etc to deploy those things. If you wonder "who automated blue/green deployments?", that's your ops team.
Also, in your example of "human kubernetes" - ops builds those clusters, and monitors those node pools. If you half less than a dozen clusters, or less than 100 nodes among the pools - you might not even have an ops team.
What seem like efficiencies can rapidly become barriers.
It’s like any case where two systems share a requirement - you can factor it out into a shared library or you can duplicate the code; in the case where the common requirement is only ‘coincidental’ not ‘instrumental’, you are better off duplicating the code so that the two systems can evolve independently and not take a coupling to a shared dependency.
The same applies to infrastructure. Sure, you’ve got a dozen clusters, and it seems efficient to have one team set up and operate all of them - but are you sure the efficiency of one big team is better than twelve much smaller teams, closer to their dev orgs, who each run one cluster more tightly suited to that org’s needs?
How far down can you push that decentralization?
With smaller and smaller units of cloud compute and storage being available as services, the answer is increasingly ‘all the way to each individual application’.
All of those things were written by developers.
I think for point 1, he's trying to say that application developers aren't doing their role as both dev and QA. I've witnessed the same issue where an DBA had trouble installing Maxscale on two identical servers. He was convinced that there must be something different between the two servers despite them being created from the same template, and only differing in IP/hostname. He had done no research, opened no tickets with the vendor, but instead wasted 30 minutes of my time arguing that it's not his fault. And this is common with many of the developers I've worked with in the last decade.
For #3, I don't own the application you develop. We provide you with a platform that YOUR application runs on, based on requirements you provide. If you don't do an adequate job of providing accurate requirements, that's on you, no my team.
And #4, developers don't abstract all those things away, they often fundamentally don't understand how they work at all, so they ignore them. This ignorance has damning consequences when they make blind assumptions about how things work.
If there is a problem with the deploy let's meet, fix the issue and most importantly learn from the problem, and document the incident for future reference.
And them move on without fingerpointing.
I've come to dread cute management phrases like "everyone should pull on the rope". I agree with the sentiment, but software development is not as simple as pulling on a rope. There are lots of moving parts and lots of things to specialize in. And I say this as a generalist dev, not as an ops engineer.
I agree with TFA completely. I was interviewing for a job recently, and one of the questions I would ask when the interviewer signaled it was time for me to ask questions was "how do you handle QA?" On some occasions, this got me weird looks, because "QA" seems to be an antiquated concept.
In a similar vein, my stint at Amazon taught me that one of the questions to ask my interviewers is to tell me about their on-call rotation. Is there any? How often are you on call and for how long? Who gets paged first?
Yeah, we're all on the same side, but there needs to be some structure and order. Otherwise, you end up with something like this:
"Twenty-seven people were got out of bed in quick succession and they got another fifty-three out of bed, because if there is one thing a man wants to know when he's woken up in a panic at 4:00 A.M., it's that he's not alone."
-- from "Good Omens", by Sir Terry Pratchett and Neil Gaiman
Not being on the same side should not be the problem, it's the solution if you divide the work as each side has to treat things differently while working together.
Operations takes care to restore the operational side. The faster the better.
Developers need or normally want to tackle the issue differently and most likely after operations reports updates they can decide whether the workaround is acceptable or development wants to maintain further on.
this normally works very well this way as operations does not have the time to maintain the application but the developers have.
I'm willing to help troubleshoot and provide guidance based on my experience, assuming the application developer has performed their due diligence. I have no insight into what their application is expected to do, or its failure modes. I have no input into the coding methods, the test harnesses, the deployment process. But when that shit breaks because the dev doesn't understand the difference between `rm -rf ./*` and `rm -rf /` that's his problem.
Now of course this is an org problem, not a team problem. As in parenting, setting boundaries and responsibilities is the key to success. Too many leaders in IT simply think that "DevOps" will be cheaper and faster and leave it at that.
My previous company had a HUGE problem with Devs cowboying off and doing whatever and dumping it on the Ops team at the last minute.
One of the biggest (but for damn sure not the last) issues was a dev who designed and built an entire new product around a MongoDB database, which wasn't something we had in production, and something he didn't mention during the months of development and demos to stakeholders. Week before the launch date he hits up our Ops folks to get production set up.
Ops was calm and collected about the whole thing. "We don't have MongoDB in production. Are you volunteering to learn how to correctly install it, write monitors for alerting, be paged with issues, figure out backups and how to ensure our data stays safe, secure, and available? You're not? Then get the [redacted] out and rewrite your app. Yes it will affect the ship date, and yes it's your fault."
I'd love to say we used that opportunity to shore up our processes involving kicking off new applications and including Ops folks in from day one, but that took years more.
Love the shoot-down!
A developer was tasked with adding a major new feature to one of our older monoliths. He added MongoDB as a dependency. The application already had a well managed Oracle database. Nothing about the feature required MongoDB.
When it came time to go to production, the DBA and ops teams responded similarly to how you did. I wish I could say sanity prevailed, but the business mumbled something about contractually obligated release dates and forced it through to production. Pretty sure it is still there rotting away.
I've worked mostly on the app side of things and this sort of thing just makes me shake my head.
If they listened to your DBA/ops guys no value would be gettig shipped ;)
Counterpoint: the dev is doing this to remain employable, so that they can ensure higher success in the future for themselves.
Their goals simply don't align those of ops and are at best parallel with those of the company as a whole - of course it's to be expected that they'll attempt to prioritize their own when there's a lack of governance and oversight within the company.
It's something that i've noticed more and more, yet is something that noone really talks about - people wanting to use bleeding edge technologies just because they're at the top of their hype curve: wanting to implement microservices when they're just maintaining monoliths and there's no need for them.
Personally, i'm an advocate of both microservices (or at the very least modular monoliths), containers and many of the new technologies, with the exception that i've initially tried all of those out in personal projects in the evenings and weekends. Yet what is the person who doesn't code outside of work supposed to do to remain employable? Would you expect a doctor to practice new types of surgery in their own time? Actually, why don't companies fund a week every few months for their developers to upskill themselves? Just a bit of time that's treated like a vacation, but during which they're expected to hack together prototypes etc.? Clearly most companies out there don't do greenfield or pilot projects, so something like this could help.
I don't think i have any good answers for this, but it definitely deserves more consideration!
Seems a bit of a waste to rewrite the app instead.
Not that I would recommend Mongo anywhere, production or dev, but it would apply for any other technology for which this happened.
> So, you could have delayed the app by the same amount
> but now have a mongo environment for production as well?
No, we couldn't have. Not just because we didn't want MongoDB, which at the time was notorious for data loss, but because our ops team didn't have the capacity at that point in their schedule or team size to handle it. Maybe had we discussed at the beginning of the project plans could have been made or altered, but we didn't and so they couldn't. > Seems a bit of a waste to rewrite the app instead.
The responsible dev took the time necessary to rewrite the data layer to better reflect the needs of the application.Is what I wish had happened. Instead the developer jammed the huge JSON blobs into a column on an MSSQL table and changed a few lines. lolsob.
Sounds like quickest way to deliver value to the customer. As described, was far too late in the process to worry about deploying with a clean, extensible architecture.
A reasonable amount of technical debt in order to ship in the timeframe available.
What if your mongodb database drops its data and now you have production impact? Are those losses calculated while making these decisions during development.
You're misusing a tool because you didn't do the correct application design in the first place.
NoSQL has it's place, mostly in the trash. Lazy key/value stores (which is all that NoSQL is) throw away all the benefits of relational logic for a glorified combination of a file system and grep.
That's not "delivering value to a customer", that's delivering crap.
Standard "Agile" response. It was only "far too late in the process" due to a complete lack of process, oversight, product quality ownership and capabilities.
If nothing else, that developer should be "counselled" as should the PO, the Scrum Master and anyone else involved that allowed the situation to occur.
And the ongoing capex and opex for the additional unbudgeted support should be pushed back on the PO as a requirement to fix.
Lol, I get your point, but that was also true for the dev organisation. Hence what you ended up with.
I doubt the needs of the application included a rewrite in MSSQL.
If it turns out it was product owner pressure, the product owner gets a call too. Possibly first.
I've had an Ops team that had a similar attitude, and they did a lot to help me become a good developer. Part of that was requiring that I come to them with identified problems. "Hey I'm getting this error, can you take a look at a stack trace in a language you've never used and tell me what's wrong?" would have gotten me booed/laughed out of the office, and for good reason.
It's not at all unreasonable to expect the developer to come around instead with "hey my application can't write to this NFS mount like I expected. It's running as $user, the permissions look right but I'm still getting permission denied. Any thoughts?" (A real situation I ran into, turns out SELinux had further permissions I was unaware of, and my Ops lead Chip was happy to show me what was what.)
Yeah, we're all on the same team, and that cuts both ways -- Ops should ensure Dev has what it needs, and Dev should make some actual effort to understand the landscape their production applications run in. Which seemed to me to be the entire point of TFA.
This a thousand times over ... If you can train your users to do this any customer relationship will be better off!
"a) I do exactly this, b) expected this outcome, c) but got this instead"
Short and to the point, it's remarkable how much easier it makes things for everyone. I think I got it off usenet at some time.
* What did you see?
* What did you expect to see?
* What sequence of events led you to see what you saw?
* any other info; (OS, time of day, location, anything that causes you to see something different doing the same thing)
On the other hand, everything else he says is wrong. Devops isn't "ops gets out of the way and lets developers do crap and nobody owns it when it breaks" its... dev and ops WORKING TOGETHER. If you don't know what their app is doing YOU ARE NOT DOING YOUR JOB. (On the same note, if they don't know anything about how the system is deployed or the like, they're also not doing their job). The rest of the article is literally about complaints about that. The point that you can't monitor an app you don't understand, and developers don't want to push out crappy code, is the whole reasoning behind devops. The people no longer being in their own silo, but instead working together as a more holistic team to own the thing from end to end. If you stay in your silo but add automation, you are 100% going to have pain. There is another pattern they could try, thats the platform model. In that case ops can stay in their little silo and present a platform with apis that developers can build on. Its what you're talking about here, and that can ALSO work. But its a different model. The old style of ops they would do as they were told, while trying to restrict anyone from changing anything. As a platform team, now they're delivering a product. They should be talking to users, and judging throughput, and iterating quickly, and being customer focused... basically, acting exactly as the developers are supposed to be. I am a strong beliver that the platform team should have product owners and customer metrics the same as the developers- heck, if they like QA, they can come up with a QA process. But yeah, I've seen a lot of low effort finger pointing in orgs that pretend to do devops from inside their functional silos. That point, the one in the title, is a great one.
Developers are not just "users", they're fellow software professionals who can reasonably be expected to work harder on troubleshooting than reporting "it works on my machine but not in the test environment :(" without even reading the error message or including it in the report.
As a general rule, when you have most of the control or knowledge of a technical process and you want someone else to help you with it, you need to give that other person as much transparency and info as possible. Because they don't control the process and will have to slowly, laboriously ask you questions, or ask you to do things, rather than just probing the system themselves.
They're taking time out of their day to work in a relatively inefficient and frustrating mode just to help you out, so jeez, have some respect and try to make their jobs a little easier.
If you don't and prefer to wear this entitled attitude, fine, but you're just as much an asshole as he is.
I agree with point 1, nearly completely. A lot of developers could take a lot more responsibility for understanding the environments their applications operate in, but I get it.
Point 3, at least in my shop, you're just wrong. I don't know anything about what you're writing. I probably don't even know what problem it is supposed to solve. You are mistaking the highway road crew for mechanics.
Point 4, in my shop, we provide a lot of documentation and guidelines for this sort of thing. Developers are responsible for knowing if their stuff is going to fall outside of those, and come to us to work something out. Again with the road metaphor, if you drive a semi into a single car garage, you're the idiot, not the person who built the garage.
On some of this, I'm taking a hard line. I do, in fact, end up doing a lot of troubleshooting with developers. But most of my team does not write code. If you want more senior ops folks who also have a coding background, come on over! There aren't that many of us who are any good, and I would love to hire more.
I don’t follow this. Developers are responsible for learning what kind of environment their application runs in, but ops is not responsible for having some clue about what they’re running? That cuts both ways, and it’ll help everyone out.
> Developers are responsible for knowing if their stuff is going to fall outside of those, and come to us to work something out.
I find this attitude fairly common amongst ops people. They just build something that is totally inappropriate for actual usage, and then dump the responsibility for figuring that out on the developers.
I don't think it's as cut-and-dried as your question frames it, but I do think there are fundamental differences between the two positions that justify some of the tension there.
The problem is the difference between domain knowledge and general systems knowledge. The former varies wildly from org to org, team to team or even within individual teams. The latter is more consistent across wider applications and over longer timeframes.
Developers usually need a lot of domain knwoledge to do their job, which can leave less space for systems stuff. But the systems stuff they do learn tends to be more widely applicable.
Ops folk often service many teams where the domain knowledge differs between them. The best of them might be able to internalise all of those differences but it's a big ask. And there's rarely any crossover.
This difference is also why developers tend to have a slower ramp-up time than ops engineers do on joining a new team. It's just the nature of the work.
I say all this as someone from the developer side of the fence. I'm fortunate to have some years in the bank now that the systems stuff comes more easily. The domain stuff remains really hard.
With the road metaphor, one issue I've seen is ops will create a rope bridge and get mad when devs need to drive a car over it. "You shouldn't do that! You idiot! Just walk over the bridge like we expect!"
Example: We have about 500 different applications in our company and the ops team maintains a single rabbit cluster for all apps (and everyone is supposed to use that one cluster). If an app gets too chatty on that cluster "Oh you idiot, why are you so chatty! You just sunk the organization!" Which, in turn, discourages the usage of rabbit (maybe that's the intention?)
> But most of my team does not write code.
I actually prefer this ( :D ), our ops team was a bunch of converted devs that decided the best way to do things was making a giant ops framework for all devs to follow. That ended up costing WAY more money than if they'd just used tools that were available. They fetishized trying to make everything "just one line!" which ended up breaking anytime you had a slightly different need (trying to take control right up to managing how version bumps happen).
Overly trying to force a single method of implementation has a lot of negative consequences. I prefer instead to have guidebooks and examples with the freedom to be an idiot and walk off the beaten path when needed.
Well, the main problem with the "bridge mismatch" is usually that resources required for an environment are not free. Its usually the opposite, most infrastructure is rather expensive, and running multiple systems side by side because multiple developers require slightly different versions of the same thing tends to explode cost.
How? Honest Question. If you know nothing of the application how are you able to offer any input into the infrastructure it runs on.
You try to respond to what people need and add things when there is enough demand. But I can't know what your business goals are, what your uptime metrics are, or who your users are.
At some point, your app becomes a black box that takes in requests, accesses DB/storage, and emits logs/metrics. I just don't have the brain space to be intimately familiar with each service.
Do you... work for the same company? Draw your paycheck from the same revenue stream?
If you're just a black box provider of undifferentiated compute/storage why the heck are we paying you? We can buy that from a dozen cloud PaaS providers.
I do my best, but I'm never going to be intimately familiar with your product on a technical or business level in the way that you are when you spend 20-40 hours a week on it. If you want that level of service, you're gonna need another couple million a year in ops staffing budget.
You're paying ops because someone needs to know how to fit all the lego AWS gives you together, understand what's inside the lego pieces to debug issues when things go wrong, be accountable for ensuring best practices are implemented as far as security, backups, etc, optimize spend, and figure out how to architect all this stuff to make sense on AWS.
We could get some of that with some of the more managed services like Heroku, but at our scale the premium we'd pay is waaaaaay more than two ops salaries.
The highway road crew know what a car is, though, right? They know that the road needs to be clear and flat and drained of water, and the markings need to be clear, so that cars can drive on it.
When the devs come to you complaining about flat tires, you can't turn round and say 'this is a mechanic issue, I don't know how tires are meant to work. They go on the bottom, right?' - you're meant to help check for rusty nails or bits of metal in the road that are causing all these flats.
'Oh, I didn't realize that was something that could cause trouble for cars'
Well then you're a pretty crappy highway maintenance guy.
The more they need me, the bigger my paycheck gets. My TC has increase 4x last five years and I do not consider myself some infra guru by any means.
Why is it Ops job to guess at your application requirements? You have the best understanding of what setting LOG_LEVEL=DEBUG is going to do to disk requirements.
And logs are not tracing or metrics or canaries or any of the other required operational monitoring capabilities that an application requires.
If your QA team is slowing down releases then that is the developer's fault not the QA team. Frankly, this move fast, don't do proper QA is irresponsible and a danger to users.
Then again, the entire point is to release after all the bugs are fixed, not to get all the bugs into production as quickly as possible :)
I wish more companies valued QA teams, then maybe I wouldn't get so many notices of security breaches and need to keep checks on my credit.
On our teams, we are all T shaped, so whoever is low on tickets might temporarily jump into the QA team as well.
They're a ratchet pattern, adding more is easy but once they exist it's very difficult to find someone with the authority to authorise removing them and the willingness to stick their neck out and declare that they aren't required and the willingness to spend time on low-importance maintenance. As a consequence logs build up until something gives and they become high importance urgent failure. The middle bit where they "aren't important" but they still waste storage space and networking bandwidth and processing power (and money) and when there is something to debug they waste people's time because the important details are needle-in-haystack among tons of low-value filler, all gets ignored.
At the limit, it isn't sustainable to print the complete internal state of a system at every clock cycle. It "should" be possible to do a lot better troubleshooting_power-to-log_weight ratio than "print every state change which feels important at the time in whatever semi-English message format is convenient", shouldn't it?
You really need to have that pointed out to you?No. Equating internal teams with paying customers is the very attitude that is causing these problems. Encouraging teams to think about their "internal customers" leads those customers to become entitled. We work together in the same company, our relationship is not the same as with actual external paying customers. I can't tell a paying customer that they're being unreasonable or lazy or unrealistic. We absolutely should be able to have that conversation with other internal teams when appropriate.
The post is describing the situation that has evolved as a result of QA being phased out. Telling Ops to suck up that extra work because "Dev are your users" is exactly why the post was written.
As an ops person I've had to explain the devs own architecture to them; they didn't know how it sent mail -- nothing to do with SMTP; they just hadn't shared the knowledge among themselves of the db/java app interaction.
I once had a developer tell me ridiculous things like "my java app can't write to java.tmpdir". They couldn't even tell me what file they were trying to write. I had to dive into apache docs and send it to them. I turned out to be a bug in an apache project code, nothing to do with tmpdir writeability.
The lack of basic responsibility and ownership was appalling.
Ad hominem is where you aim to refute the argument someone is making by impugning the character of the person making it. Obviously, it doesn't follow that because someone is a bad person, that what they say is wrong.
"Don't believe this guy, he's an asshole" would be an ad hominem argument.
But this is not an ad hominem argument being made here. They are, instead, making the logical claim that, based on the attitudes described, the person comes across as an asshole.
They aren't then saying that that invalidates their arguments - they are taking the author's arguments at their face, and inferring the person's character from them.
Which seems to be a mode of argument you're comfortable with, since you've just done the same to the person you replied to.
Seen from the other side, this is an insane approach.
If you cannot even estimate how much resources over time your code is going to use are you really a professional?
Really, I'm not even talking about a precise estimate, just a ballpark estimate.
2. QA is still needed sometimes. The tedious tasks need to be automated, but that never gets prioritized.
3. Traditional Ops only cover non-functional and operational requirements. What is agreed and documented with Ops?
4. Ops are rarely invited in early, if ever.
The problem is the siloing, aka not doing DevOps.
Maybe you're imagining the wrong scale of organization, here. An ops department isn't usually "embedded" into a development team (especially when there's more than one development team!); it's essentially a company-internal Platform-as-a-Service provider that the development team deploys their app to. From that PaaS's perspective, the app is an opaque workload. It's not ops' job to fix your app, any more than it's Heroku's job to fix your app.
> Maybe if the infrastructure people were involved in development discussion earlier, they'd be able to raise their hands and say: wait a minute, you're going to blow up our logs.
"Infrastructure engineer" and "operations" are entirely distinct roles. (Maybe not at a startup with five people, but once you get to even 30-or-so, there's a clear delineation.)
Infrastructure engineers are fundamentally software engineers, who happen to know a lot about infrastructure, distributed systems, networking, etc. They know about the operational constraints of software. And as such, they usually get put in charge of release management for the software—i.e. get put in the critical path for changes—because they have an eye for what changes to the software might break the deployment.
Ops people, meanwhile, aren't anything like "in the loop" of your software engineering process. Their day-to-day is spent managing servers and various well-known software systems running thereupon (e.g. Kubernetes, Nginx, RabbitMQ, etc.) They get handed opaque components (those well-known systems, and also your app), glue them together, automate "around" those components using runbooks, document how to get things back into working order when they crash, etc.
In small companies doing "DevOps", there are no real "ops people." There are only infrastructure engineers doing ops.
When a company becomes large enough, there is a transition point where managing the servers and all the standardized stuff running on them gets too distracting to your infra engineers, and they find it hard to help with the app, because they're too busy fighting fires and doing maintenance to the operational infrastructure. At that point, you hire actual "pure" ops people, and build an actual "pure" ops department, to take that load off the infra engineers' plate, so they can get back to their true comparative advantage, of guiding the app in an infrastructure-conformant direction.
But that separation necessarily means that you now have people managing your servers who aren't engineers. They're technicians.
-----
A labored metaphor, for your enjoyment:
Your ops department is like the service center for a motorpool. The people working there are automotive technicians. They are not automotive engineers. They can't make you a car, or change the components of your badly-designed car so that they're better-designed, or tell you what your weird prototype car means with its weird nonstandard error messages.
They can do standardized probes, get industry-standard error messages out, and do things about them. They can swap out broken components for newer releases of the same components. They can replace consumables. And they can notice if something is weird in a statistical sense (i.e. if some of the weird proprietary metrics the car keeps are not within historical reference range), and point that fact out to your automotive engineer.
But you've got to have those automotive engineers, on staff, in the development team, to deal with that information.
This has been the bane of my work happiness for a while now. I keep having to tell junior devs to actually _read_ the fine error message, just in case it actually _contains information about the error_, you know. Not that it seems to help much, it’s like they can’t get the concept into their heads.
This is 100% a problem with younger, bootcamp-”educated” devs, in my experience. I know the common wisdom on social media is ”no one reads text anymore”, but if that includes aspiring developers, it might be tough to replace the current workforce when that day comes…
well except for using Rust, and a bunch of dependencies from the web, version determined when downloaded at compile time.
"lack of newness" is a characteristic many will expend untold hours to extinguish. to my perspective, the "rewrite it in Rust crowd" is the peak; all non-Rust code is soiled, and worthy of replacement.
(it is very possible that the "rewrite in Rust" movement is just a guerrilla marketing project)
This lead to a lot of them learning common failure modes of the software they ran.
Ironically the people who were script kiddies in their teens have been some of the best troubleshooters I know.
I was doubtful that this was a universal truth then, and I think it's the same now: there are a lot of people who do mediocre work, and they are and were supported by a smaller group of people who do really good work. And the world keeps turning.
One of the joys of the Internet and of open source is the increased ease of sharing ideas and solutions.
I think the only solution for this is when "support" includes education. In the simplest form, we can give support by helping that colleague to find the issue himself, rather than giving the solution directly. In a more advanced form, you're making structural changes to your company. Like in how you share knowledge with the team.
We've all heard the stories where someone bright joins a new workplace, is assigned drudgework and after a bit automates it until they only work two hours a week? That's someone who is willing to learn, and everyone else in the office was willing to experience boredom in order to avoid learning.
For all I know, some of those people would have perked right up if the subject matter were milling, millinery or masonry moving millier-weights of stone. Not everyone is interested in the same things, and even when their job is all about it, sometimes people aren't invested in it.
If it works, everyone involved wins. Sometimes it doesn't work.
You're doing everyone a disfavour giving out full answers or not demanding some homework first, and often, it's really an ego-issue.
I think there's something about certain kinds of tools that are cryptic, unpredictable, and frustrating that can teach you to be helpless - to just Google and hope. It's fixable though.
It really is worth it to go the extra mile and write comprehensive docs, even going as far as writing them in a conversational tone as if it's a blog post or a book. I'm really happy I found a company who treats documentation and workflows as first class resources.
For a small team where only 1 person is working on this it helps eliminate the bus factor and it also makes it easier to have non-hardcore ops folks do code reviews on your IaC. Having them be able to get the gist of it with a little bit of background knowledge is so much better than nothing. All of this results in higher reliability of the services your company offers.
Maybe they have seen so many red herrings that they don't even trust that the error message could contain something useful and relevant? Or maybe they just learned to skim through everything, and don't actually read stuff.
I also don't understand why, when they ask for help, they can never be bothered to say what they're trying to do, what error message they got, etc. It feels like they're doing me a favor when I try to help them fix something.
Really helped me gain the mindset that not only was it my mistake that resulted in code not running, but that it was fixable. Like a game of ping pong. You hit the ball, sometimes the compiler hits it back.
Computers are scary things that fail in counterintuitive ways. When the handle of your tea cup breaks, the issue is intuitive and most people will be able to understand why it is happening and how to work around it(handle it carefully from the top end end enjoy your tea?).
But when it comes to computers, often you need deep understanding of its inner workings to make sense of your observations of problems. Why Xcode would say that it failed to compile my project because usefulExtensions.swift already exists? What it is supposed to mean, I see only one file with that name? That information gives intuitive idea about the issue only if you know how the compiling process works.
Why would I know why the package couldn't be found? Unless of course I know how that package manager works. Then I can check if the package manager is configured to look at the correct places.
Most error messages are like that. Instantly makes intuitive sense if you know how everything is glued together and makes no sense and needs study if it's outside of you domain of expertise. No one reads error messages unless they can recognise the pattern instantly and there's a data(like the name of the variable) guiding you to the fix.
Not being scared of the tool and believing one's inherent supremacy over it must be the most basic criterion for practicing this craft, but these days this fear is nursed, at times encouraged, at times even exalted (corollary of the failure fetish) especially by those who publicly place themselves as ambassadors.
Any introduction to computers must start with the statement that they are all heaps of plastic and sand and the only things they are able to do are because some mortal sat down and spent time figuring it out.
People starting out now are at a disadvantage because their first encounters happen mostly through extremely polished looking apps and it is hard to see at the outset how one could go from weird incantations in a text editor to that.
Personally, I love debugging things. I have a very good "theory of mind" for dealing with computer failures, and figuring out why the computer isn't doing what I might naively expect it to do is a lot of fun. However, it's only fun because I've been able to stay on top of the curve as the systems I work with have become more complex. Starting from zero today sounds a lot more daunting.
Nowadays, research skills are more important, but I see a lot of devs who just don't have them. Can't find the answer on the first page of your first (poorly formed) search? Run get the senior dev. To me it reads like incuriousity and laziness, or lack of training.
I don't mind doing some coaching, but if you're a dev, and you can't even be bothered to read the error message, what does that say about your effectiveness?
/rant
This scales to everything IMHO, everything is simple once you understand it. Levels of abstractions is what makes it scary and complex. I.e. electricity or fire is also not scary once you know how to handle it.
They're supposed to have that knowledge, or at least not be afraid to dive in and get that knowledge.
There's only one way to build an intuition of what kind of problem probably causes some error (most famously, if the error is completely incomprehensible, you missed a closing thingy on the previous line), and that's by doing the work a lot.
I don't think so. We can do so many amazing things with the computers precisely because we don't have to know how things work. Computers are so many levels of abstractions over printed metal on melted sand.
People who know what they are doing will understand the errors of their own creations and will learn the workings of the tools they use to some degrees and will be able to understand the failing modes of these tools with experience over time. No one starts with complete knowledge before start building things.
> or at least not be afraid to dive in and get that knowledge.
Of course they should have the drive but people's first instinct would be to make the error go away so that they can do their actual work. People have limited time and energy, you can't expect a JS developer, for example, to study inner workings of a Linux box to understand all errors. It's cool when they do and gives them superpowers but it also makes them less productive as JS developers. Sometimes you simply need to implement that button to render on the server without studying the server.
When even line numbers are missing, simple syntax errors can generate new errors and mind numbing troubleshooting.
How would they obtain those abilities though if not while spending time on the issues brought up and learning how to learn.
I think sometimes people are just bored and can't be bothered to find the cause and solution to their issues, and over a long period of time that mentality sticks and becomes second nature resulting in phrases like "this software sucks, I need to read the docs to use it".
The problem is, learning is taxing and many times you encounter these errors when you have more important things to do.
When you want to develop your game and the IDE is complaining about something about locating some files, do you think that it is good idea to learn how that IDE organises dependencies?
Sometimes you suck it up and learn it and you know next time. However, your first instinct would be to look for ways to make the error go away so that you can immediately start working on the task that you are supposed to work on. That's why we have abstractions and when things work fine we don't know how things work.
It shouldn't be expected of you having complete knowledge of all computer systems, tools and frameworks before you can make a ball image bounce on the screen.
If I know typical causes of errors (forgot to connect to the VPN, etc.), I'll include them in the log message as well as things to check.
Often you need to know a lot of context before you're even able to determine what the error message is! One error message can lead to a cascade of other error messages, or it's something breaking down as a result of multiple layers of indirection, requiring the developer to careful track the trail of what went wrong and led to another thing failing, which broke down the next thing and ultimately, decided to stop the program and mention only the very last thing falling apart to the user. There might be a directly sensible connection with the original error, but often it's quite unrelated. An experienced developer often immediately recognizes: this is not the actual error message, that other thing is! But for a junior it's all equally incomprehensible.
It is detective work with many false leads, and being very new at something it can be so overwhelming you don't know where to begin and immediately assume you will not succeed finding out 'whodunnit', asking your senior co-worker for help.
At least capture them so somebody who knows that area can make sense of them.
So yeah, I think I largely agree with your assessment, and would only go on to state that the path forward is slowing down to learn vocabulary and think critically. You really speed up after that.
1. If ops staff have limited expertise/authority, it's less likely they can resolve problems. They might acknowledge (so maintaining some aspect of client SLA), or have a limited set of pre-defined remedial actions (reset button). Anything beyond that, though, and it needs the dev team. So it's arguable whether the ops staff provide much value in the equation.
2. As a dev, there's nothing quite like the prospect of being paged at 2am on a Sunday to incentive more robust code.
End to end dev accountability isn't a panacea either - but the problem is more nuanced than just pay rates.
You broke it, you fix it.
And by "application" I mean the combined software, release, documentation, runbooks, etc. Not just the latest git tag pushed.
It’s usually not hard to figure out what’s wrong from the messages, but man do they look scary and hard to understand when they appear. Yes I’ve been writing C++17 lately using some very template heavy libraries.
When I did C++ we sometimes made little competitions for the smallest change that can produce the craziest error messages. On the other hand I always found it extremely satisfying to make one little change that removed thousands of errors and warnings.
Easily diagnosed if you're working incrementally, one small change at a time, and making checkpoints with version control: `git diff`, carefully review the diff of what you changed since the last checkpoint where things were more or less working. I must have not been disciplined enough to work like that at the time.
Troubleshooting systems integration failures is also character building for getting better at diagnosis from errors. Sure, it's failing, but let's try to figure out the immediate layer of failure from the logs, error messages, symptoms: name resolution? tcp? tls? http proxy? authentication? authorisation? api spec misalignment? error in our application code or the system we're directly talking to? unexpected data? error in some other system that we depend upon transitively? each time you hit a new novel failure mode, or fail at one level deeper, you're making progress!
In my experience it’s also common with older and college-educated ones; contractors trying to avoid extra hours; senior architects; and especially anyone who thinks ops is someone else’s job. It’s definitely not specific to age or training mode.
There are a few contributing factors I see: tunnel-vision focused on the particular detail they think they’re working on, causing them to ignore anything they “know” isn’t related; shoddy tools like much of the Java ecosystem where poor culture around logging trains every user that it’s normal to have huge amounts of log spew; etc. but the biggest problem I have seen is ego — either unwillingness to believe that the product of their staggering intellect could be less than perfect or that the mundane task of getting their grand vision to actually work is for the little people.
I’m thinking of a “senior architect” who was quite surprised to learn that networks are neither perfect nor instantaneous, and that his app might have some issues due to needing thousands of XHR calls to load the UI. It was so much easier to ignore the error messages and say the problem was Chrome. He had a CS degree – the problem was the wrong mindset and having been enabled to avoid good troubleshooting skills.
It's only after the technological illusion of Maya breaks that you realize floppies have read heads, hard drives have moving parts, CPUs have conductive traces and all of these are vulnerable to breakdown, entropy exists in the system and cannot be expelled, that the previous "ideal" state of your system was temporary, an illusion, that nothing always works the way it is supposed to and that your options boil down to "burn it to the ground and start over" or "leap into Hades both feet first to rescue the soul of what you love".
Most people go the first route. Buy a new one. Replace what is broken with something else. That way the illusions are never broken. The technology didn't fail, only its current & easily replaceable avatar.
As the Son of God once did, after its death it will rise again, immortally replaceable.
However, it is only after you have faced that 2nd trial by fire and returned with your elixir that you as a changed being can peer through the veil. The meme about "CPUs being rocks we filled with lightning & tricked into thinking" rings differently to you now.
You're touched the bones of the God and found that they crumble. There is no God here, only a beautiful shambling nightmare that has eaten the minds and souls of millions, built by mad scientists and engineers in a vain attempt to create the God whose physical absence they find themselves longing for the same way a neglected child longs for the embrace of their mother.
I wonder if it's also to do with the environment in which they learn. When I was learning to program, like probably others here, I didn't have anyone around me who knew anything about computers so was generally on my own until my first job and had to dig through stack traces and read error messages and had to try and figure out what was wrong. Kind of a blessing and a curse as I imagine my rate would have been a quicker and I wouldn't have hit so many brick walls but I learned to debug independently.
teach people to read error messages and simultaneously improve the readability (and utility) or error messages
i don't know why we put up with such bad error messages anymore. i imagine it's a function of stockholm syndrome and the difficulty in getting messages changed
I have an example:
I built a logistics and invoicing tool a few years back with error messages that where human readable with clear proper messages that told the user exactly what they did wrong and it even proposed how they might fix the problem.
I don’t know how many times I had to go to the users workstations read the text out loud for them like they where a 5 year old and ask them what they thought it meant. They always knew what it meant but I had to read it for them it was embarrassing.
And these where university educated accountants that where using the software.
After a lifetime of garbage error messages like “error code 4513” people just zone out.
I get codes back in the day when storage for a whole book was costly, but that isn't the case anymore. Just tell us the error, show us the pointers, and then tell us what typical fixes are instead of expecting us to go to the internet for a solution.
As developers, we also have to be used to a lot of completely unhelpful errors. Yeah, couldn't connect to the DB, sure... oh, but actually because my code ate all of memory, why didn't you say that in the first place?
Software just kept the tradition of error codes, since that meant you could also sell that juicy documentation (localized into whatever language you wanted) to the user as well. I suspect it also a localization issue because OracleDB would never return the table/column in the error message, so as to be easier to translate.
The modal dialog breaks UX spectacularly.
Even logs should have messages that actually are intuitive to follow.
Or cause the user had filled in part two of a task but not part one and then tried to continue with parts of the task that where dependent on filling in part one.
Or the user tried to synchronize orders from the erp system but the erp system would not return any orders.
I resent this. Not because I'm a bootcamp-"educated" dev. I'm not. But it suggests somehow that devs with CS degrees are somehow better in this aspect. If anything, they're arguably worse (obligatory, not everyone disclaimer).
I think the people most likely to fall into the “it doesn’t work” category are people who don’t have much experience troubleshooting difficult problems.
In the end it’s about compassion and understanding of each other. Unfortunately in a lot of companies the only direction people are getting is “get it done on time”. It’s rare that management asks people to have empathy for each other.
They're not used to reading error messages because they've been brought up seeing nothing but completely useless error messages.
You'll get a lot of pushback here but it's definitely true.
That doesn't mean it doesn't happen with CS grads as well, but it's quite rampant among bootcamp devs. I think the reason for that is that, since the bootcamps are so short, they "stay on rails" and mostly work on simple projects (that will give out something they can push to a github repo and use as a portfolio).
It's the same with git. Every bootcamp will use git and claim to teach it to their grads, but then watch them do anything on a repo with multiple users. A lot of them just rote memorized commands to pull and push to main and that's it. Branching? Rebase? Using the commit history? Never heard of.
For new hires from serious Engineering or CS Degree, they should have had at least a few classes dedicated to projects where they built something non-trivial. On top of theoretical classes teaching the fundamentals.
We built a couple relatively simple applications for an enterprise client. It took their ops teams months to get both applications running in K8s, even though our deliverable was a fully functioning container. They were largely incompetent as far as we could tell.
But, I don't think it's worth being unkind or judging them. Every time they asked us a question we made an effort to point them in the right direction. There were other times it was a problem we couldn't help with, we kindly let them know that.
I think the reality is that the demand for competent IT and developers outpaced supply a long time ago and it's not getting better. Those of us who know and care about the difference should make competent co-workers and executives part of the job evaluation. Or, accept incompetence around you as a reality, help and avoid as wisdom dictates.
But, complaining that it exists and framing it as competent ops vs incompetent developers is both untrue and unhelpful IMO.
The latter part of the article that talks about the pace of features, complexity, and the lack of time is spot on though IMO. I think the article would have been better focusing here and avoiding the IT vs devs angle.
Competent and incompetent people exist in all areas. Some of those incompetent ones can get better with time and support, and some can't/don't.
If it truly was a competence issue, how did leadership solve it?
Let me share an experience. In 2010, I worked on a project for a large business in the US(Fortune 100). The process was set so rigidly that it worked well, but I was among the group of people who were mad at it saying”why is this so rigid? Trust us and let us do things faster!!”. Context : There were change management rules in place. The software was to be released only on a regular cadence of about 6 months, only after thorough integration tests, and approval from the change mgmt board. Should anything go wrong in “move to prod” there will be representation from dev, QA, Ops, change mgmt, and Mgmt orgs to immediately decide on actions until the release to prod is successful. There will be thorough documentation of what to do (run books) on what changes occurred, what their impact could be and how to rollback if something unexpected occurs. It was always a party after a successful release :-)
Trust me there were a lot of bugs, but they were mostly found and fixed during the laborious QA and integration tests by people whose job it was.
Fast forward to now, I am a “Cloud Engineer” in a small team that does everything from app development to building CI pipelines to running services on AWS to being on-call to keep them running.
I must say, I wish for the old days back. Sure, it was slow and laborious, but it resulted in better outcomes and manageability. IMHO, it also resulted in better reliability of software due to the diligence done by several layers.
It is easy to say do the same just faster in your small team. But, in practicality it just doesn’t happen. I work on setting up Observability one week, then onto designing infra for a new service, then onto some development and so on. I feel like my scope would have been limited, and I would have had an easier time becoming an expert at something than becoming so broad skilled like I am today.
Sometimes, old, slow, and mature is not so bad. Not everyone needs to follow the FAANG SV companies to be successful.
I like to distinguish between “product developers” (i.e. building products for consumers with guaranteed scale, so do it right the first time) and “project developers” (get it done ASAP and cut the corners you need to do so).
In the “project developer” world, 50-75% of your requirements gathering happens before a line of code is written. There is usually a “right way” to implement a process of which technology is only one component and figuring that out as you go will actually slow down the project due to the maker / manager schedule conflict. True “agile” in this environment just leads to scope creep as there usually aren’t dedicated product owners to say no to every little request.
I’ve stopped pushing agile as hard because the corporates simply can’t afford the kind of engineers to make it work correctly, and they don’t have the roles required to gather and feed requirements to a dev team in an agile format. Sprints are a good way to time-box feature development, but most business projects work better with a more waterfall approach. Your customers and project plan operate under waterfall so there’s less downside to begin with.
Waterfall model has its downsides in extracting the requirements out properly whereas the Agile approach(the little I have seen of it) seems to lose the layered stability of a waterfall based approach.
Those were also the days where it took many years to go from Java 6 to Java 8. Or perhaps to try out Kotlin.
They were the days where legacy code was the norm, and we kept supporting it because nobody dared to change anything for the better. In practice, that's just not something you can maintain in a competitive market, because your competitors _will_ use new technologies and faster/better development processes.
"it just works" might be good enough for maintaining your application, but will it be good enough to find people willing to work in that code base or that environment?
I work for a large business where both the old and new practices are in place (mostly the new ones, though). Focusing on "going fast" is definitely not a good idea, but I believe there's a sweet spot in between.
I'm on a project now that has not released to prod. It has a lot of new legacy code.
Responding to "Someone will always have to own that gap and nobody wants to, because there is no incentive to. Who wants to own the outage, the fuck up or the slow down?" with "Not me." is not sufficient, it's a very valid question for which any organization definitely needs an answer pointing at some specific people - if it's not going to be pure ops people, it's IMHO not going to be the feature-developing devs as well, that would likely need separate 'site reliability engineer' teams as some major companies do.
Seems like OP works for a shitty company.
Throw in some red tape where I can't have access to logs myself? Then I don't care to fix it at all - chasing another team, that has diverging priorities, is complete a waste of my time.
If your Developers are tossing shit over a wall, I'd bet top dollar you work in organization B. In which case they are behaving accordingly. Don't empower me to identify and fix issues? Then I won't (and I won't lose sleep over it either).
But of course they required two-factor authentication in O365, so it was SOX2 compliant.
I think this is the original thesis for DevOps.
It wasn't supposed to mean no more division of labor. Division of labor is a key innovation in human society that enables civilizations to exist. It was supposed to mean the teams in different categories of labor interact throughout and consider each other's needs, and not throw shit at each other over a wall and only ever interact through a ticketing system.
This waffly "breaking down the silos" stuff is a later redefinition, i believe reacting to the fact that the original meaning was extremely unpalatable to existing organisations, with existing employees and hierarchies who would be severely disrupted by it.
From Patrick Debois in the interview "Later, I saw a talk by Jean-Paul Sergent about developer (Dev) teams and operations (Ops) teams working together."
So no, breaking down the silos was baked in from the beginning.
The team that writes the code should also deploy the code and get paged in if there are problems in production.
That creates a tight feedback loop that requires developers to learn and manage the whole stack, code defensively, and test enough to be confident to deploy to production.
Didn't test your code enough? You will be paged in the middle of the night to fix it. It creates a strong incentive to make good decisions because you will be living in the mess you create.
I have heard of companies deciding to do 'devops' and it turns into a free for all of dev teams having to handle/build things end to end. Everyone loses in that scenario.
This is how it is at my startup. All of the engineers are involved in managing the infrastructure for everything we build. I find it gives me much better insight into my app, and the feedback loop is much tighter since I am in control of everything.
To be fair, Developers are getting slammed with their responsibilities too. At one time it used to be that they could just know one programming language really well, like java, compile their code and hand it off to QA.
Now they have to know a dozen languages, frameworks, do their own testing, deploy the service, monitor it and trouble shoot everything in production in some 'cloud'.
Or they are just being lazy and this guy is sick of it. That is when you do your best to train people up and get them to put in the leg work. Ask pointed questions about if they Googled the error and help them work through the problem. Then add some things to the docs to help others out in the future.
Oh no, responsibilities!
Meh, the days of throwing a tarball to QA and log off at 5pm are gone, thankfully.
Some organizations empower their ops team to close support tickets by just saying "not enough information to diagnose a problem".
I've seen with my own eyes team metrics improvement (SLAs etc) by just counting the time the ticket was "in progress" to the ops team instead of waiting for the developer (or customer, whether internal or external) to reply.
I am always saddened when I hear "our organisation has a DevOps team" - immediately this demonstrates the fundamenetal lack of understanding the very premise of what DevOps set out to solve: Bringing Development and Operations together.
Even the very name "DevOps" was constructed such to symbolise the combining of the two domains into one. But no. Now we just have a new cool title to throw on people who will be ringfenced just as they were before.
"DevOps" today is just codified Shadow IT.
Yep, because developer might not know what ops needs in terms of traceability, logs and so on, to be able to run their code in production without having to wake them up at 2AM. Similarly Ops knows a lot about what can be done with existing infrastructure, or off the shelf components, which can save a huge amount of work, while providing a more stable system.
I do mostly operations now, and I'm lucky enough to work with really talents developers, who care to listen to input, before writing 5000 lines of code. I also work with customers, who have their own developers, with their own weird ideas about the world.
The biggest problem I see right now, except for occasions cowboy pretending to be a professional developer, is developer picking technologies without understanding it. We work with customer who picked technologies because they're interesting, not because it's what they need. When performance is terrible it becomes and operations issue and being told "Kafka is not actually a database and should be used as one" often isn't the answer they want. Or try telling a developer that the code he worked on for three months can be done by the existing load balancer in a few hours or that the ORM is actually writing terrible queries.
DevOps team, as in: "We use the shared knowledge of both parties" is fantastic, but operations is frequently an afterthought and not involved in the design fase.
If we're to take "DevOps" as developers doing operation, I'd prefer that we do the opposite and let operations do development. I think we'd get better results.
If I’m asking a question I explain what I’m trying to figure out, what I’ve tried, what I expect, what I’ve researched. Basically helping the answerer not waste as much time covering the same ground.
If I’m answering questions and don’t get this info, I ask it. And establish the expectation that this info helps me answer their question.
About 70% of the time, the asker adds in more info. 25% of the time I don’t hear back. 5% of the time I get a complaint that they are too busy or can’t answer the questions.
”Often they have not even bothered to do basic troubleshooting, things like read the documentation on what the error message is attempting to tell you.”
This happens, but this just means that your Development Team needs some coaching or to improve their quality.
This tells more of a quality of the development team you have been working with. You have to pass along this feedback and ensure that Development team also works with professionalism as everybody else.
DevOps would tell be that "Dev & Ops" would look up issues together (Yes, he will be blocked as well WORKING with you), if you find that it was developer's fault. Tell them: "Hey, this is on your side. You saw how we troubleshooted together. Now each of us has new tricks to use in the future".
If you don't do that, you are the shortest path to get THEIR problem solved. And it is too easy to go that path.
Good management will ensure that if a problem like this occurs, they don't "coach" but they "counsel" the appropriate dev to do their job and not waste everyone else's time.
I have been in this industry for 15+ years, and as a developer, I have a surprising amount of experience dealing with customers. Of course, when a customer complains about some feature not working, I would not just take their word for it. Customers mess up too.
What I would not do is brush their complains off. "This is a systemic issue". "They are causing problems". "They don't know better". "They don't have the correct incentives". Try telling that to a customer, or to your boss.
The obvious disconnect from his own team is the problem.
People from dev regard this as "dev bubble" = your last 15 years of experience. Not a joke.
This is the crux of the problem. Coding in isolation. Replies of 'It's java, it should work anywhere' etc.
The other gear grinding commom theme is not even doing basic troubleshooting. To the point of not even googling the error message or the symptoms, and being 'blocked' because they are waiting on a ticket they opened with the 'other' team.
There’s some very clever ppl that know all about how networks/vm stuff work, and I’ve learnt enough from them that I can fix most of my own infra related things - or at least give them a run down of what I’ve done first to save them some time.
It got me back into hardware and networky stuff, so now I’ve got a MikroTik at home, some proxmox machines, Tailscale network etc - more fun than just spin up a box on DO and be done with it.
A lot of ppl just aren’t interested though, they just want to code (and maybe learn a new language) but because a lot of stuff is now PaaS and it’s super easy, there is no need to learn it (in their eyes)
I had to chuckle - everyone (not just developers) seems to blame the network first! (including blame the firewall rules)
And now people who call themselves programmers are like this themselves.
Now can I have some other kind of future please?
Especially when the dev wanted to migrate from Java 8 to Java 11 and didn’t even attempt to lookup our documentation on how to change JVM parameters.
Your code threw an error? Did you read it? Did you search Google?
A monitored metric dramatically changed with the last deploy? Have you investigated why that might be?
I am more than happy to help troubleshoot a tricky problem if you tell me what you've already tried. If you truly don't know where to begin, I'm also happy to teach you. What I am not going to do is fix your problem for you, with you retaining nothing.
Programmers have to push out an endless stream of features; DBAs have to deal with ever greater amounts of data; network people have to deal with an enormous amount of endpoints (and now the network extends itself beyond the firm, so security concerns have grown exponentially).
The real challenge is to make your IT departments realise that they are not each other's obstacles.
As with most organizational dysfunction, middle management fiefdoms are to blame.
It always helps when the executive can see through this bullshit and ask the right questions, but often by the time this happens millions of dollars have been wasted.
Developer: Host XYZ is very busy.
Sysadmin: Yes, Yes it is. The top 10 processes are your Java App.
Developer: Fix it.
Sysadmin: ???? You can request a larger virtual machine, you can try these options to the JVM, or you can fix your code.
Developer: Can you do it?
Or: Can I get a bigger server... Yes, but you have 32 cores and 256GB of RAM, and your applications isn't that complex.
There needs to be a rethink of how infrastructure, development and deployment is handled.. maybe the solution is to slow things down and insert a little carefully thought out bureaucracy between the layers (can't believe I'm advocating for more bureaucracy!)
I feel like devs at small companies are doing everything -
coding, testing, supporting customer, deployments, troubleshooting and of course straight over ssh+winscp cuz vps/bare metal are cheaper
large organisations usually have a much bigger responsibility and is held to higher standards, frequently audited and controlled to stay within rules and legal compliance
If you've got 2 developers, they're both doing everything and on call 24/7 and all have read/write access to everything on demand.
If you've got 200 developers, you're going to start wanting a team of shift workers keeping an eye on the systems, and maybe you won't want every developer to have read/write access to production data.
If you've got 20,000 developers your working practices and infrastructure are almost completely cemented in place, and anyone who doesn't like them has already left because it's easier to change jobs than to get 20k people to change their behaviour.
> It is baffling on many levels to me. First, I am not an application developer and never have been. I enjoy writing code, mostly scripting in Python, as a way to reliably solve problems in my own field. I have very little context on what your application may even do, as I deal with many application demands every week. I'm not in your retros or part of your sprint planning. I likely don't even know what "working" means in the context of your app.
The point about not being in retros or part of sprint planning... I take up arms against that. I've worked for companies that have gone from waterfall to hybrid agile because we cannot get buy in from Ops to actually... you know... come to our retros, sprint planning and scrums.
Some things in this article is just pointing out the obvious... mediocre developers who push their problems and/or lacks on other teams. However, that quote the Author needs to look in the mirror. They exist only because of the products offered by the Company need resources. They have a responsibility to be business partners in that. If they aren't the company needs to re-align some priorities and it could start with Ops. Ops doesn't get a pass in an agile organization. The whole point of agile is to destroy them ivory towers. And if they were in those planning sessions, the developer might have already gone over the type of destructive testing that would have emerged from that collaboration and their DevOps relationship would be even richer.
Another symptom of this is that when the QA/Staging function went away, load testing became perfunctory. Many of the performance problems we see should have been caught in QA. Devs are anxious to ship and get on to the next sprint, leaving app support and operations on the hook.
I think it goes a step farther back to product. PMs and analysts put constant pressure on developer teams to complete work quickly and that time pressure shows up on the next guy's plate, etc
"Trickle down software engineering"
I've had previous coworkers approach me about API "bugs" because they didn't bother to troubleshoot their app code and just immediately assumed it was a server-side issue.
Then I spend 10 minutes debugging the issue only to point them the error in their own code. I don't know if it's laziness or inability to troubleshoot, or both.
As for me, I'm currently watching a good devops team go down the drain because of a bad manager, so I'm seriously considering trying to move to management so I can help my employer do better.
I guess these developers will end up writing code directly on github online editors...
Now I'm in a dev ops team (as a dev) and we spend a very large swath of our time---troubleshooting infra issues. It's all AWS and our problem now.
Running code at scale turned into a very challenging comp sci program, and uptime vs code slickness is getting prioritized by clients.
The career support and innovation in that corner of the world (ops eng jobs) reflects it. Sort of gets after what software architects do, but the requirements to know that come way earlier in the career for Ops. Ops Engs with cloud knowledge, Python, and IaC tend to go far.
Similar in nature to "Running arbitrary containers" but without the human trust-to-do-no-evil policy in place.
Devs build a reporting framework. They test it. All is well.
Devs deploy reporting framework to production. All reports suddenly time out.
Devs complain to the network team that the network is broken.
Network team spend weeks checking every single interface and link, in all involved network paths (8with multiple daily "we are not seeing any network issues, is the application still having problems?", so definitely not done in complete silence).
Network team eventually pulls out the packet sniffers. Network team then asks if they can see the relevant code snippet.
Root cause is that in the test environment, the RTT between app server and (test) DB was on the order of 0.1 ms. But, in production, the RTT between the app server and the DB server was ~20 ms. And the code generating the report did NOT pipeline any of the data gathering, effectively making the report generation taking ~200 times longer.
Unfortunately, the network team was unable to deploy a higher light speed limit for that network data path and instead suggested that the application fetch multiple data items in parallel.
The problem is the Ops teams get ZERO credit for enabling the shoddy work done by devs, the devs meanwhile get patted on the back, and frankly continue to be romanticized.
"move fast and break shit (and let Ops fix it silently)"
This is a huge problem. Working on reliability and security is hard, shipping broken features is easy.
In that regard, those roles are slowing things down and costing money
Sometimes, when old IT generalists work with new IT specialists, these sort of misunderstandings occur.
Of course, this is just my personal opinion based on what I have experienced over the last 40 years.
I think the actual boundary lies more in big vs small companies: small companies do not have the resources to hire specialists for every little subproblem, while big companies typically have enough employees that specialisation becomes a possibility.
No one should ever get to play "not my job" while simultaneously throwing complexity grenades over to another team.
QA as a discipline has evolved but from the sounds of it, it's not been widespread enough.
Function-organized teams encourage knowledge siloes, and its-some-othrler-teams-problem-ism.
It seems more and more places want it to be. DevOps is all the rage.
In the end, overall, it sounds like an incompetent ops, who is unable to keep up with the pace of this industry, complains about incompetents devs and that industry does not wait for people to catch up.
These are all wrong.
Complexity comes with valid reasons. High demands thin out the pool of competent devs/ops. We, the lucky bunch who earns 2x-3x salaries of other professions, should be the most humblest bunch, especially if there is frustration within ourselves about the fast growing reality.
In shared hosting times. Ops maintained Puppet definitions, create new "deploy environments" (using Puppet), gave/revoked server access the employees, monitored servers. Devs maintained the source repos and deployed to the environments provided by ops.
Now we live in virtual machine times (docker). Ops does the cloud infra (terraform), monitors the services, gives/revokes access to cloud services. Devs maintaines the source repos and deploys to the cloud clusters provided by ops.
Don't get me wrong, I appreciate well operating systems and the people who make that happen. There is going to be a beancounter wondering whether the very expensive, trained engineer can operate faster/cheaper.
The author seems to want it both ways. They want the devs to fix their own problems, but at the same time give them zero control of the stack (we have to provide them with guide rails to prevent them from hurting themselves indeed).
One should not build silos where experts sit.
One should participate in a team of many different experts.
If you still have to call "whatever department" for fixing your slow SQL query, your disk space, your repo access, production deployments etc. Then in most cases I feel you should get out and find a place where silos are not being exercised.
This is a thing of the past.
Is it? In highly regulated environments, silos are practically a requirement. Security access controls are intentionally put in place to limit access to systems and their respective pieces. If you need repo access, or access to the CI pipeline, or access to the database, you have to go to the appropriate channel.
- no need for anything but minimal security, perhaps because the business is trying to surf the margin between their AWS bill and their Google ad revenue.
- there is only a need for security in the part of the business that deals with money
- some stuff that the users do, they would prefer to maintain integrity but they don't care a lot about confidentiality
- the users want reasonable confidentiality, too
- everything about the business is money or secrets
Where your business is on that scale determines how much you are regulated and how many internal gatekeepers are necessary.
In a sense, kubernetes is the new Linux / Bash of our time.
If it's painful, maybe it's just the abstraction not done right, but not the fault of abstraction itself.
But in the cruft of legacy systems it probably won't be for a long time, probably never.
I'm a freelancer. I "use" the project managers of my clients, I don't hire them myself.
Same goes for my applications. I use the cloud, managed services and such. The providers hire operations people, I don't.
Yes, developers should understand the operational environment a system runs in, and should be capable of advanced troubleshooting. But the rest of the post is simply tired screed about how "the old days" were better, despite the fact that they manifestly were not.
I have been at this for 25+ years and I remember the days of silos. I remember developers passing off code to QA and it coming back days later with bugs. I also remember a lot of 'not my problem' coming back from developers. Sometimes it wasn't their problem; often it was.
Either way, the person with intimate knowledge of how the code works should be the first person that looks at the problem. In SaaS, that is production and should be the developers (within reason).
As Ops, my goal is to make that as easy for possible for developers. That means automating everything I can so that the right tools are in place for deploying, monitoring and alerting are there. It means that I have to automate spinning up and destroying infrastructure as quick and easy as possible so that we can meet your needs and also keep costs down.
I have also seen companies fail at becoming 'devops' in the most terrible way. They took developers and made them own everything from code to deployment to VMs. The developers had so many pieces to understand that the only guarantee was failure. That was a terrible startup to work at.
Exactly. They're specializations, like heart surgeon vs. orthopedic surgeon.
> I have been at this for 25+ years and I remember the days of silos.
32 years for me. Silos evolved out of the wild west of the 90's and early 2000's. Which evolved out of the strict controls of early computers run by a cult of Operators where the devs couldn't even access the machine directly. It's a cycle, where management tries to remove people, only to have to put them back later. I've seen it over and over.
> I remember developers passing off code to QA and it coming back days later with bugs.
It is literally QA's job to find bugs that developers missed.
> In SaaS, that is production and should be the developers
SaaS or not doesn't have anything to do with it.
> As Ops, my goal is to make that as easy for possible for developers.
As a developer, my goal is to deliver high quality code that meets the requirements for performance, stability, monitoring, security, and functionality.
> They took developers and made them own everything from code to deployment to VMs. The developers had so many pieces to understand that the only guarantee was failure.
I've never seen DevOps done any other way. Hence my original comment.
It was called a scrollbar.
There is a big difference between the mindset of what makes a good application developer and what makes a good ops person.
Application developers, by and large, have a sort of "sandbox" within which features are developed. This sandbox results from working with abstractions, each with some kind of guarantee. For example, most application developers assume that what you write into memory will be what you get out. That is, hardware is abstracted. The idea that the memory chips themselves can have defects or can sometimes fail, even if one gets error-correcting memory chips, is a violation of guarantees. Memory that just works is taken for granted. Another example is assuming that the system clock is monotonic.
This extends to things like networks, storage, operating characteristics, and so forth. Very few application developers get into that nitty gritty, let along all the plumbing and interactions among different systems.
I've seen application developers get incredibly frustrated and angry when those underlying guarantees are violated in some ways. I've been like that when I put on my application developer "hat". The main reason is that the developer is holding as much of the state and logic as they can, and they do this by excluding things through abstractions. They want their tooling and platform to just work so they can focus on writing good software.
The thing is that, for a good ops person, all those nitty gritty and plumbing is the focus of their jobs. It takes a very different mindset to troubleshoot: you start looking at those "guarantees" and find out what they are actually doing.
I once interviewed at a place which has this amazing way of figuring out if someone has the mindset and tenacity to be a good ops person. It was not writing algorithims on a white board. It was a deceptively simple task of installing a piece of software. And even though there are documentation for installing that software, there are not documentation for installing that software for every single environment and requirements and its interaction with other systems. When adding in a time crunch, and the scrutiny of an observer, that simulates a pretty typical day pretty well. You have to have enough emotional intelligence to keep working through it until it works. Documentation is always sparse and can't be guaranteed to be correct. Runbook? Good idea, but there is no way even meticulously crafted runbook for one component is going to be able to describe how systems interact with each other. Someone, somewhere has to figure that out. (Well, they don't have to. We can just let the system fail).
And sometimes, as an ops person, you have to open that "black box" and read code. Just like sometimes, an application developer needs to pop open the abstraction layer and pull out netcat or sysdig.
In the end, I'm not lamenting that DevOps blurs who owns what. Maybe this is because I've mostly worked on small, early-stage startup teams. Complexity has to live somewhere. I like working on the teams where people talk to each other to figure things out.