Incident Command for DevOps: Learning from the Fire Department [video]
youtube.com
youtube.com
The key to successful incident response is to design, agree and test the ICS processes, roles and Command and Control (C2) hierarchies before incidents occur, capturing the results in an Incident Response Plan. The IRP will typically involve standing up an incident response team when an incident reaches a pre-defined severity level. Incident response teams are often structured into bronze, silver and gold levels of command (gold typically including Cxx individuals such as CIO and CFO, not just DevOps roles) that temporarily replace Business as Usual (BAU) management with incident C2. In the UK at least, the bronze/silver/gold hierarchy was developed by the ‘blue-light’ services (police, fire and ambulance), and is itself a simplified version of military C2 chains of command.
If management don't buy into this type of approach, it doesn't node well for an organisation's ability to deal with a crisis.
If I understand/remember correctly, IS-100 and 200 are the required courses on ICS (along with 700, 800 for NIMS itself) for all frontline firefighters on any fire department wanting to receive federal money.
1. https://training.fema.gov/is/courseoverview.aspx?code=IS-100...
I tried to introduce the ICS concepts formally and up-front, but met a lot of pushback; "we don't need that", "we're not dealing with fires" or "it's too complicated".
By using ICS principles without calling them such, people usually see the value. I even got a sizable promotion and raise due to my "clear and concise handling of several serious incidents and putting procedures in place to handle similar in the future". All I did was direct people in to ICS functions and act as an IC.
I agree that ICS concepts should use used more widely and outside of just emergency services, but getting past the "stigma" of the title is the hard part.
Really useful training, have had to use it multiple times to coordinate responses ransomware, and other IT disasters, and does benefit from the buy in of senior leadership at the company I work for.
Biggest challenge we have is keeping ICs after deciding to centralise the IC organisation in 3 locations, all ICs were offered the choice to relocate or leave.
We have a lot of very new ICs.
ICS is crucial in any and all of our incidents and should be the model on how any disaster is handled
Can't be encouraging for Slack, couple of years. But, that's the deal now, people leave tech jobs every 1-5 years for something better or start their own company.
So far, a month in, I’m loving working for Slack; it’s a great company, and an excellent group of people. So there’s every possibility that I’ll be there longer than 2 years!