Facebook Workplace co-founder launches downtime fire alarm Kintaba
techcrunch.com
techcrunch.com
One of the nicest things was that we use a single deployment tool for almost all deployments, and it could insert the deployments into the timeline so we could see every deploy both during and before the incident.
But the biggest issue was voice. We always had a conference call running for incidents because some people would be in a place where they couldn't open chat (driving to the office or similar) and it was always a pain to integrate the voice data.
We got to a point where a recording of the incident call and a transcript could be added to the incident, but during the call the Call Leader usually had to quickly type the voice highlights into the chat.
I'd love to see a real time voice transcription as a feature. But it would have to be pretty good to not just get in the way.
So it was a few teams all reporting to the same Director. We even had the customer service tools reporting to the same director so we could get some real time info from them (and to them).
Each team built in what they were most comfortable. Most of the tools were built in Java, but my team built in Python (mainly because I write in Python and biased towards hiring people who also like to write in Python).
Sidenote: Other reason I asked is cause apparently both Hulu and Netflix are listed on the CherryPy docs as orgs using them.
I was going to link to some open source with CherryPy, but it looks like they've all be rewritten in Go or Gradle. :(
Netflix is definitely on the good end of companies with custom tooling in this space. Would love to chat more about how you do things if you don't mind me reaching out personally?
My investors are a bit tough, I get no equity and no seats on the board (there is no board). Vacation and pension’s good though.
Paul Graham wrote: never do things for money or ego, but techcrunch clickbaiters writing about facebook employees, sorry, founders, supposedly dont agree.
P.S. this whole title reads like from git manpage generator.
Anyway it's an interesting idea. I worked in a support center for decades and demands for updates and managing P1 type situations, giving updates to the dozen interested parties (each of whom were competent in some ways, less so in others), managing misinformation, and the varying politics was a constant hassel for folks who were technical.
There were times where "oh man my phone just went down" (I unplugged it).
That's no small thing to try to deal with, with software.
We really wanted the interface to feel comfortable for non-technical folks who needed to stay up to date on incidents so we focused on bringing as much of a human element as possible into the dashboard. Hopefully we can find the right balance of friendliness and informativeness, we certainly don't intend to become a social network (but we would love it if we were as easy to use as something like Slack or Discord)...
We experienced the same pain you're describing with keeping others updated and the politics (and subpolitices and subsubpolitics) that come out of major incidents over our careers. Luckily we also saw companies that did it right, usually with custom built tooling. Our biggest discovery was the more open you make the whole incident process the better everyone understands the work being done and also the less of those insidious little subpolitical conversations happen where facts are skewed and people are blamed for process challenges.
It's definitely a big hill to climb.
It wasn't the worst way but also not a "system" for sure.
Half the battle with political stuff was really was defining the problem in a way that everyone outside engineering understood, and keeping everyone updated / aware of what was factually (not rumors or panic) going on, what was known, not known, what was happening, and when they would get their next update.
And that's not even the technical stuff where folks here are asking about recording conference calls and etc ;)
Looks like StackOverflow to me but clean. I like the UI. I don't know if I will ever see that utility where I work since we wikify things like this post-chaos. The office I work in is a small enough team to where we have plenty of flexibility.
Like when I think of dealing with a situation I see that and think of some sales guy going off in his comment about
OMG THIS IS JUST LIKE THAT OTHER TIME WHY ISN'T THIS FIXED!?!!?
Now the whole comment thread is off track because sales guy brought his own FUD to the game, and god help us all if someone brings it up on a conference call as fact.
Now obviously that is not the fault of the UI if you've got a sales guy does that, but as things scale up organizationally that kinda thing ... happens ;)
Granted this system might be intended for engineering teams, but it's also something that inevitably leaks out when folks find out it is there.
Even if they don't have an answer some way of making sure FUD doesn't take off too quickly.
I'd say a high % of managing any situation and the politics that follow is accurately defining / revising the problem and making sure everyone is on the same page with it. Dodging / handing the inaccurate information leaks sometimes is a big job.
Granted that's all office politics stuff and not just technical... still gotta solve the problem ;)
No one wants to rewrite PagerDuty internally -- why are people all writing their bespoke incident management, response, review, and reporting toolchains internally too?
Because the good ones are tightly integrated with the rest of the internal tooling. To use a 3rd party incident management tool usually means you have to run your operations they way they expect to get maximum value. A lot of times its easier to build it yourself based on how you operate.
However, as the way people operate become more standardized, the third party tools will become more useful.
Here's hoping that the calculus changes either as these tools grow more robust or, like you say, as people begin to manage their software systems in less bespoke ways.
For example, during an incident we'd like to know "what changed between time [X] and [Y]" (deployments, configuration, experiments, other service outages) while we're trying to determine a root cause and fix the problem. And much later, after the incident, we'd like to auto-compute a metric like "what is the change success rate of [Y service / services owned by team Z]?".
This aren't complicated concepts -- it's not like we're trying to reason about causality with machine learning to reduce time-to-resolve for incidents! Still, our incident-management tool will really behave better if it's aware of our what-changed-at-time-X tooling and our incidents-caused-by-changes reporting. If this is an external tool, yikes, now our incident tool has sprung awareness of changes and reporting?!
One of the reasons we started Kintaba was that we're noticing a shift as more and more employees are leaving successful companies that have invested tons of time into their incident management process (most of the bigtech companies) and are now joining/starting their own companies where they want to spin up a mature process quickly.
I'm sure that stuff exists (I think Datadog sort of does this, etc) but I've yet to work anywhere that doesn't just create some #shtf-$date slack channel which eventually gets lost in the black hole due to cost prohibition or time required to get it going while a fire is going on.
One thing Kintaba is really good at right now is wiring your slack channel directly into the activity log so it's properly attached to the incident and ultimately the postmortem. This helps avoid that channel getting lost over time, but there's still lots to do for sure in organizing all that data to make it more useful after the fact (one thing we currently support is #tagging for quick search within the log).
Few things are as annoying as going back in repo time and trying to figure out if you're in quite the same place / right place.
A couple notes: - the verification email went to my spam folder on Gmail - acknowledged is misspelled on this image https://kintaba.com/images/collab_splash.png
Thanks for the spamboxing report. Seems gmail isn't a fan of us today... working on it now.
Just making you aware, as the site causes lots of problems and is not GDPR compliant. Top comment on the thread above: "Yahoo/Verizon is cancer and should die in a fire."
Yeah. <closes-tab/>