DevOps Students Learn the Value of Uptime With 3 a.m. Calls
linux.com
linux.com
And on developers, so they realize that reliability is as much a requirement as any of the functional requirements they were given.
While it's normal and expected that the 'Ops' or 'DevOps' or the artist previously known as 'Sys Admin' is expected to be available off-hours for stuff they didn't build (read: emergencies-that-are-not-within-their-purview-to-prevent)... on the flipside, the tendency of many companies to reduce the burden of responsibility of developers and product managers who are often directly responsible for instability in the environment, is prevalent in a way that makes my blood boil.
Going back to this article, it's almost a celebration that DevOps folks have to shoulder that burden.
While it's normal to be on-call, being woken up all the time is a sign of a badly run infrastructure OR release/change-management practices that are rushed and/or feeble.
Everywhere I've worked I've stood up against this tendency, analyzing each issue that causes a page and seeing how it could be prevented. A quarter of the time it's technical: creating redundancy, deep diving into an ongoing issue, doing load tests and capacity management, etc. The rest of the time it's political: oh, the new code caused the memory to be sucked dry from the system, this started precisely after the last release (proof, here's a graph from my Check_MK setup); oh, the devs ran a crap query again on the Hadoop cluster even though we warned them not to do X. etc.
I think if they want to prepare future DevOps students, rather than using PagerDuty as per the article, maybe they should give them shots of liquor and strong beer, their livers could use the preparation. It's an incredibly political role.
Lots of words for "You can save money be shoving ops bullshit onto devs instead".
Oh right, it's all about collaboration and ownership and it's way more efficient this way. Just like open offices.
I am willing to acknowledge there are some pros to this system, but it's still screwing over devs by shoving extra (inconvenient) work onto their plates with no extra remuneration.
We don't run to cardiac arrests. Walk fast, with a purpose. I fail to see what production issue necessitates me running down the hall at 3am.
And DevOps students should be learning how to prevent (sorry, minimize) this, not conducting fire drills.
"That is why there is no upfront cost to join Holberton school. We only charge 17% of your internship earnings and 17% of your salary over 3 years once you find a job. If the company you join agrees to pay us a placement fee, this percentage will be reduced."
Its unsettling.
That's ridiculous. That's "Guess I can't take the dog outside this week" territory.
I realize the expectations are unrealistic, but as I'm sure most people know, your options are to leave or stick it out.
We're part of a very special group that employers are legally allowed to screw over regarding OT. Because reasons.
Actually I would really like to know the context of how that law got passed, if anyone has insights.
Look I've been you. I took the crappy job because I figured that was the only one available. There are better companies out there. Just ask around. Get a few offers.
There is nothing normal about being on call with a 1-3 minute response time. I'm pretty sure it's illegal as well.
Not at all.
What are you meant to achieve in 1-3 mins? Answer the alert or be at a workstation?
How are you meant to have anything resembling a normal life?
When it's built right, operations doesn't exist. On this theme - we're hiring. Central London. We need a platform/infrastructure engineer: solid unix, automation-centric, someone who will drive evolution of the infrastructure and release platform. https://clearmatics.workable.com/jobs/257440
And keeping students on-call 24/7 is abuse, nothing less.
This seems like a waste of a good night's sleep to me.
While there is always someone else's mess you have to support, that does not seem like a great teaching moment to me.
Because network outages never happen, disks never fail or fill up, memory is never an issue, programs always deal with only the data they were expected to, products never do more traffic than expected, and all infrastructure software ships completely bug-free.
If you aren't occasionally up at 3 AM fixing unexpected outages, then either you haven't deployed a project that requires uptime or you're paying someone else to do it for you.
I have never gotten up at 3am to fix an outage because outages do not impact systems that widely. If it is a network problem, let the network team do their job. All of your other problems are solved with multiple HA/ load balanced servers, monitoring, and proper testing.
several of these things are exactly the type of things wehre the whole idea of devops just falls apart. If a disk fails what use is a dev vs a good ops person?
We partnered with companies such a PagerDuty, Wavefront so that they have tools to help them keeping their uptime, we are guiding them to put everything in place so that their website/servers never go down. We never call them at 3AM and we actually never call them at all. However we do have challenges where we simulate hardware failure, traffic spike...
Just to make things clear: the goal is not to wake up students are 3 AM but to get them ready for production and this involve being on call and possibly getting paged at 3AM.