Post-incident review on the Atlassian April 2022 outage
atlassian.com
atlassian.com
But the people who work there are human, and I know what kind of a toll a protracted recovery effort can take. I know it from chronic pain two years after the burnout started. No one should experience that.
I know it’s probably not the thing Atlassian the business thinks they should communicate to customers/general public, and it’s not surprising it’s absent, but it’s a shame they did not address the well being of their staff. Especially those who undoubtedly put exhausting effort into recovering from this incident.
And sure, that would be unusual as a business communication. But should it be? Even before this severe burnout, I would have looked at this list of lessons learned and thought “a lot of people are going to be overworking the same as they’ve been through all of this, but now they’ll be doing it invisibly.” And that definitely doesn’t inspire confidence that there won’t be another catastrophic mistake in the near future—it doesn’t make me more confident in the business.
None of this is a strong criticism, just some observations from being on the other side of a marathon incident recovery.
Some Atlassian customers might have had much more severe (mental health and other) problems than the Atlassian staff, so I think it can be perceived as mildly solipsistic to be praising the staff and (being perceived to) forget about the customers' wellbeing (I know you didn't say anything about forgetting about the customers, but affected customers may perceive it that way).
> I know you didn't say anything about forgetting about the customers, but affected customers may perceive it that way
No, I didn’t mean to suggest or imply this, and hope it won’t be taken this way by anyone. It was perhaps a mistake taking it as read that obviously their customers were also harmed. But I do think it’s a reasonable point to make more explicit, and I do agree that it likely had a similar impact on customers to the experience I described. I’ll repeat that no one should experience that.
That said…
> As someone who's worked on similar situations, I don't expect any pat on the back, but see it as my responsibility to make sure it doesn't happen in the first place, and when it does, my responsibility to fix it without complaint.
This seems like you’re addressing something entirely outside the actual content of my comment, and perhaps projecting your own priorities onto it.
First of all, I was in no way suggesting anyone get special reward. And I was in no way referencing any Atlassian employee’s complaints nor airing my own. I mean it sincerely that I hope they are okay and that their mental health needs are being respected. It was entirely a statement of human compassion, and an observation that it’s one which can go unstated/understated in these discussions.
Secondly, I added my own experience as a personal reflection on the toll it can take. It’s not easy to say in a public forum that I’ve suffered years of chronic pain after addressing an incident. I added this for context because I think it is easy for people to dismiss the impact serious incidents have on the people responsible to them.
A side note before I get to thirdly: I think this is also true for people in many careers where incident response is a primary job responsibility. Sometimes it prompts explicit acknowledgement, often it does not. I just think the world would be better off if more people’s legitimate pain and challenges were acknowledged.
Third point: I took special care above to say “responsible to”, not “responsible for” or simply “responsible”. I’m certain, given Atlassian’s size and the impact of this outage, that many of the people involved in recovery efforts played no role in the mistakes leading to the outage.
In my own anecdote, I played no role in causing the incident. I have to walk a fine professional line here because I have no intention or desire to criticize anyone else involved. I think I can reasonably say this. I did what you described:
> see it as my responsibility to make sure it doesn't happen in the first place, and when it does, my responsibility to fix it without complaint.
Even so, I’m experiencing chronic pain years later. And now I will complain: it sucks being in pain every day for years. I wouldn’t go back and do anything differently as an IC, except perhaps to tell past me when to slow down and that I have more ability to affect short term prioritization than I once realized.
Lastly,
> praising the staff
I sincerely hope this hasn’t had the negative impact on you which you’re describing as hypothetical for Atlassian customers generally. If it has, I hope you’ll hear me when I say any sharp tone in this response is not from lack of solidarity.
But if it hasn’t and you’re replying only out of work ethic and blame-placing: kindly go re-read my comment and recognize that the only praise I expressed was for the humanity of human people whose experience I wish to include in the conversation.
You can praise staff but cannot praise the customer unfortunately as that would be inappropriate. You can work on relationship of course, by not limiting your reconciliation to contractual credits on SLA. These will mean nothing to the customer staff that took significant part of the hit in work stress. Not sure what would be a good way to regain the trust at that level.
Edit: fixed accountability vs responsibility in wrong order
Of these, losing access to OpsGenie for this long was a massive problem, dwarfing most other systems. OpsGenie is like PagerDuty in the Atlassian world.
I spoke to several engineers at impacted companies who could not believe their incident management system was “deleted” and had no ETA on when it would be back, or Atlassian could not prioritise restoring this critical system ASAP. JIRA and Confluence being down was trouble enough, but those systems being down for some time was things most teams worked around. However, suddenly flying blind, with no pager alerting for their own systems? That is not acceptable for any decent company.
Most I talked with moved rapidly to an alternative service, building up oncall rosters from memory and emails - as Confluence which stored these details was also down. Imagine being a billion dollar company suddenly without pager system: and no ETA on when that system would be back, your vendor not responding to your queries.
I talked to engineers at such a company and it was a long night to move rapidly over to PagerDuty. It would be another 7 days they could get through to a human at Atlassian. By that time, they were a lost customer for this product. Ironically, this company moved to OpsGenie a few years before from PagerDuty because OpsGenie was cheaper and they were on so many Atlassian services already.
The post-mortem has no actions on prioritising services like OpsGenie in reliability or restoration, which is a miss. I can’t tell if Atlassian staff are unaware of the critical nature of this system or if they treat all their products - including paging systems - as equals in terms of SLAs on principle.
Worth keeping in mind when choosing paging vendors - some might recognise these systems are more critical ones than others.
I wrote about this outage from the viewpoint of the customers as it entered its 10th day and it was discussed in HN, with comments from people impacted by the outage. [1]
This is not very far away from AWS hosting their own status pages.
I wasn't affected by this in the slightest and just found out opsgenie exists from the parent comment but even I can understand that this decision would almost certainly be driven by things like "we're already using atlassian for everything else and will benefit from the interop" and "we already trust them with everything else and they haven't let us down or we wouldn't be using them for any of that stuff either."
When we switched to OpsGenie it was also to consolidate billing with other Atlassian products, but we all knew PagerDuty worked better and that OpsGenie was still pretty new and rough around the edges. We certainly didn't gain any business advantage by switching, but we did have to do a lot of work to switch which took away from other things we needed to get done. But ultimately we had no say because somebody just wanted to save a little money. I also doubt that there was a thorough vendor assessment before we picked it up, since it was a vendor we already used.
When you're ready to prove that your services "don't go down" send me an email and I'll come work for you.
If your incident reporting is down, hopefully you completely stop changing anything about your infra.
14 days without pagers ...
From friends and people I know I heard nothing good about Atlassian products. And that was before that 2 weeks downtime.
It looks like the products are duct together with duct tape, spit and a little bit of dirt.
There’s an interpretation of that where life is great
1. Establish universal "soft deletes" across all systems.
2. Better DR for multi-site, multi-product incidents.
3. Fix their incident-management process for large-scale incidents.
4. Fix their incident communications.
Regarding #4 in particular:
"Rather than wait until we had a full picture, we should have been transparent about what we did know and what we didn't know. Providing general restoration estimates (even if directional) and being clear about when we expected to have a more complete picture would have allowed our customers to better plan around the incident....
[In the future], we will acknowledge incidents early, through multiple channels. We will release public communications on incidents within hours. To better reach impacted customers, we will improve the backup of key contacts and retrofit support tooling to enable customers... [to make emergency] contact with our technical support team."
However, there's one point that makes me skeptical: there are no organizational changes, or changes to leadership, or anything in that direction.
This sounds like "the tech guys screwed up, culture and management is fine here". Which it might be, or it might not.
I would have loved to see
5. We will stop pushing customers so hard towards using our cloud
for one, but that wouldn't be convenient for Atlassian.
So you can’t say it’s slow or unperforming.
> 3.3. Restrictions. Except as otherwise expressly permitted in these Terms, you will not:
> (i) publicly disseminate information regarding the performance of the Cloud Products; or (j) encourage or assist any third party to do any of the foregoing.
Used to be common to see this kind of comparisons between databased for example.
> Except as otherwise expressly permitted in these Terms, you will not:
> ... (i) publicly disseminate information regarding the performance of the Cloud Products;
But usually when we people talk about "corporate pr" they talk about weasel words, non-committal statements, and deflecting blame away from them. I think this PIR does a decent job of acknowledging the mistakes they make and how they can improve.
You missed your recovery time objective by ~2 weeks and did not properly communicate the issue from senior management until around a week into the outage. It is great to hear about how the company plans to do better, let's see how the next outage improves things.
Nobody cares that you lost very little customer-submitted data if customers rely on you (mission-critical) to continue to accept data, and the outage prevented that.
I’ve been there with much less significant incidents when a “routine” change turned into a potentially resume generating event. It’s not fun.
Ultimately the responsibility is with the organization that made it possible for an event of that scale to happen rather than the individual person who happened to trigger it but that doesn’t make it feel any better.
Recently, I was asked if I was going to fire an employee who made a mistake that cost the company $600,000. No, I replied, I just spent $600,000 training him. Why would I want somebody to hire his experience?
Mistakes almost certainly follow some kind of pareto distribution instead of it being evenly distributed.
That's why 2% of doctors are responsible for 39% of malpractice lawsuits. If your doctor cut off the wrong leg and you sued (and won) $2 million, would go back to the same doctor and just chalk it up as "well, now he's got $2 million of training"?
I see it as: Your employee just learned a valuable lesson, and you definitely don't want to hire another new employee to make that same mistake again.
Ultimately, it's the owner taking responsibility for the fuck up because they hired the employee, but still standing behind their decision to hire the employee.
If your employee is an idiot and makes up 39% of mistakes, of course you won't keep them around when they make a big one. But most employees are not fuck-ups, as you alluded to with 2% of doctors making a majority of malpractice suits.
Like they mentioned they had a delete script that worked for all types of unique IDs so that also can dilute the feeling of ”it’s all my fault” hopefully.
# find /tmp -exec rm {} /;
But instead she’d typed # find / tmp -exec rm {} /;
I learned a lot about Unix (and, as it happens, X windows) file systems that week. Note that this was long before tools like sudo existed, where cleaning up tmp was often a manual process.There is no way we could blame a person for a single keystroke error - it’s crazy that the command worked - but I now look at the find command with an enormous amount of fear.
rm -rf ${basedir}/*
without setting set -u
to exit if there are undefined variables. I suggested a few other sanity checks but was told I was adding friction.Just this week, I changed a spec in one of our proposed endpoints that did exactly that. We passed in ids of various types of objects to perform actions, and I changed the api so that it would be forced to pass in a struct that contained an object type and object id. Explicitness is so much safer in the long run, especially in enterprise apps.
And have them type out both the number and they thing they are deleting.
Crockford’s 32 encoding with a domain prefix works well enough.
1: https://docs.snowflake.com/en/user-guide/data-availability.h...
When new data is written, the entire block is copied and rewritten rather than changing the data in-place.
A write-once immutability design (whether in the DB itself, or an extension, or a userspace implementation) lets you reconsider and rework mistakes after they’re committed, not unlike how you might do with git rebase etc.
Datomic describes its capability here: https://docs.datomic.com/on-prem/reference/excision.html
If you are just interested in global time traveling, there are many solutions, such as replaying the oplog from snapshots in time or delayed replication.
Am I mistaken?
The only problem is that you need to keep a map of timestamps to txids so you can find the txid that was valid at a particular moment in time. This doesn't sound to me like a significantly difficult problem, but maybe I'm mistaken. That said, it's not like you need super high time resolution for the use case in question.
An intuitive way to think of bitemporality is it's like MVCC, but with 4 timestamps per row version. One pair describes a range of time in "outside" or "valid" time, ie whatever semantic domain the database is modeling, the other pair describes a range of "system" time, which is when this record was present in the database. This lets you capture and reason about the distinction between when a fact the database models was true in the real world, and when the database was updated to reflect this fact (some people call this "as of" vs "as at"... the terms here aren't fully settled but the basic distinction is). So you can revise history, do complex time travel queries, all sorts of stuff. It's a very useful model that directly aligns with what sort of questions businesses need to answer in the context of a court case or revising their ground source of truth due to past bug or error.
The downside is your database balloons with row versions, and many queries become far more complicated, perhaps needing addition joins, etc. Also from the perspective of database implementors there's a ton more complexity in the code. So that's why it's not widely supported despite the standard.
There's also a niche of databases built around this model from the ground up, usually based on Datalog instead of SQL. There's also overlapping work with RDF and Semantic Web thinking (as awry as all that went).
In practice how most organizations address this is operationally, by keeping generational and incremental backups that let them restore previous database states as needed. Though as the original post we're here for proves, that kind of operational solution can bite back hard when it goes wrong.
I'd argue that SQL has some pretty big guardrails. 'DELETE FROM CUSTOMER WHERE' will generally fail, because there is lots of data referring to the CUSTOMER table and the system will insist the data remain consistent.
https://www.dolthub.com/blog/2022-04-14-atlassian-outage-pre...
"To manage the restoration progress we created a new Jira project, SITE, and a workflow to track restorations on a site-by-site basis across multiple teams (engineering, program management, support, etc). This approach empowered all teams to easily identify and track issues related to any individual site restoration."
Great case study of a monumental fuck up!
[0]: https://www.twilio.com/docs/glossary/what-is-a-sid#common-si...
It's puzzling to me that we, as software developers, spend so much efforts trying to automate such one off deletion tasks, and the automation would inevitably go wrong and result in data loss.
That "any other compliance situation" includes things you probably want to have a soft delete for, like deleting/closing an account. Customers accidentally deleting their accounts, then wanting it restored happens more frequently than one would hope.
GDPR does give you a grace period , so you can soft delete immediately and then hard delete after some period shorter than the GDPR deadline. However, actually implenting such a system can be rather difficult and potentially expensive.
Alternatively move the deleted data to a temporary location and then delete the temporary location after a short period of time.
Or better combine both patterns where expired rows get moved to a temporary location before hard deleting a period of time after.
GDPR says you have to delete data when requested. As far as I know you have 30 days to acknowledge the request and up to 60 days to action the request. It’d be completely reasonable to do a soft delete for 7-14 days before doing a hard delete to prevent these kind of errors.
Yes, GDPR has a grace period of 30 days or so, it's never been a problem in practice.
* hindsight is 20/20 - it's easy to lecture after the fact about how mistakes could have been avoided
* modern software systems are very complex
* have you never made a mistake?
It is however noteworthy that they have done the wise thing from a publicity perspective, which is post this on a Friday, hoping that by the next tech cycle more interesting things will have happened for the press to report than a rehash of this outage. That's politics.
Yup that deletes something… anyway…
> Establish universal "soft deletes" across all systems.
It’s just easier that way to observe what might happen.
Would be super interested in a technical writeup on how they do this.
In previous companies I have worked for, we did instant soft-delete, then hard anonymisation after 15-30days and then hard delete after a year. That means the data was not recoverable for customer but could still be recovered for legal purpose.
A simple technical solution is to store all data with per user encryption keys, and then just delete the key. This obviously doesn't let you prove to anyone else that you've deleted all copies of the key, but you can use it as a way to have higher confidence you don't inadvertently leak it.
Of course, this means trusting Atlassian to actually delete the key on request, but there's not much reason for them not to.
Hopefully they also change their API, so these two very different things don't use the same API call.
So how would one have a clean "partial backup" strategy if something similar would happen in his company?
Depends a lot on the situation and technologies already in use. For example, if you lost ticket sales, you might monkey patch it by having two systems at the entrance: one with the main data, one with the restored backup. If the person isn't in main, you can check backup. That could be deployed more quickly than trying to consolidate the two states.
In another situation, you could isolate customers so that a delete and a restore simply affects everything at that customer and such a partial delete is not a problem. You could still have trouble if there is a partial delete within a customer system, but restoring part of 1 company is a lot less work than restoring parts of hundreds of companies.
Yet another case where using types would have prevented a massive problem.
How I imagine this to work would be that the team who wrote the system (function, script, whatever) for deleting things would encode it in the types that their deletion system would only accept specifically app IDs, rather than both app IDs and site IDs (which I think reflects the specifics of their post-mortem).
In order to get a value into the deletion system for processing, the ID value would need to be parsed into the specific type that the deletion system accepts.
This is of course similar to just saying there should have been validation, but I think conceptually, validation and parsing into a narrower type are two different things. Gary Bernhardt calls this "functional core, imperative shell". Michael Feathers calls this "edge-free programming". Probably the best literature on this difference is by Alexis King, here: https://lexi-lambda.github.io/blog/2019/11/05/parse-don-t-va...
But it had the risk of catastrophic failure the whole time. I wish we had better ways to measure and communicate risk.
Big red flag. If two teams own something, no body owns it.
It could be argued at the scale of a company like Atlassian that this level of redundancy is prohibitively expensive, that's a lot of databases and files to have sitting around doing nothing 99.99% of the time, and it's hard to argue for prevention of something that's never happened and would be a costly thing to tool up for. But you can definitely factor scaling your redundant capacity into your model, both pricing-wise and engineering-wise. It's not like Atlassian products are cheap to begin with, I'm sure they can sustain some velocity / bottom line hit for the sake of something as basic as fully replicated staging environments. I definitely don't think this is on the engineers, it's a strategic oversight and shows where ultimate priorities lie within the company.
Putting your trust in a cloud service to take care of things you'd otherwise have to worry about yourself is a major decision, and safety is one of the top priorities of basically every user, and seeing the lack of process and glib approach to staging is a major red flag.
Anyway that aside I do appreciate their detailed write up and it does feel like a bluntly honest and truthful disclosure. That goes a long way to restoring trust, but it does also expose some of how the sausage is made and it's clear some of the ingredients are questionable. It does bear the hallmarks of a small successful software startup hitting the big time and scaling with acquisitions faster than supporting processes can safely scale; they have a team of engineers and it's up to them where to engage them and it seems being able to do proper dry runs of destructive changes wasn't seen as more valuable than getting more services on the products page.
Hopefully they'll act on the recommendations of the report and implement the improvements they said they would and not just refocus their efforts elsewhere once the spotlight moves on. I'd like to see regular updates on this as a long-term Atlassian user as it would factor greatly into me recommending Atlassian products over other stacks in the future. They could easily set up a public Jira / Trello board so we can keep track of progress on these promises.
Obviously this is not unique, these mistakes have happened before, so it's not just the kind of stuff that seems obvious in hindsight. I am sure there were engineers highlighting these issues internally but scaling redundancy is never as sexy as onboarding a new product and adding its customers (and revenue) to your quarterly reports. Hopefully the reputation hit is a stark reminder to the c-suite that yes, they are running a technology company, and that means that technology and engineering should be just as important as growth and penetration.
Anyway, good on them for being open. Well done to the engineers who worked to untangle the mess, good on management for allowing this level of transparency and taking ownership, things could have been a lot worse by the sounds of things.
I'm sure there are many companies who would lose 24hr of tracking to get back online in a few hr vs 14days. In the real world with deadlines and commitments this is madness.