Update on the Atlassian outage affecting some customers
atlassian.com
atlassian.com
Their original and early Jira product was great and put them out in front, but I couldn’t name any other product of theirs except Bitbucket which is very much a second place competitor in their market. Jira itself has faded and been caught up with too.
Reading through the post and seeing the level of support some people on Twitter have been tweeting about, I would take this as a great time to look at other offerings out there.
https://techcrunch.com/2010/09/29/atlassian-buys-mercurial-p...
I haven’t used it in years, but that seems accurate. I could never find anything unless I knew exactly how it was worded in what I’m looking for.
That is, the pop-up results that appear as I typed where significantly better than the actual "search results" page.
I did see some examples of editing HTML through similar approach but never tried it myself.
You lost me there. In no way was FogBugz a better product. It was at best mediocre by 2010.
Jira is much better. We also use SharePoint and Confluence is so much better.
Honestly I get the criticism of both JIRA and Confluence but till you've used the other big names in that space you don't realise how bad it can get.
I'm so sorry. Fogbugz has been in maintenance mode for years now, stuck in the middle of a half-finished redesign.
Tell this to all the big corps still running a slow, bureaucratic JIRA system demanding one to update multiple tickets multiple times.
I have yet to find something as good, even with all of JIRAs pain points
Maybe I should check it out again
The other two I've honestly never heard of but that says nothing about them as products except they don't get talked about.
Funny edit, Atlassian actually made a plugin for Jira called greenhopper that was very much like trello. But they discontinued it only to realise a few years later that they liked trello and paid a massive amount of money for it. If that doesn't exemplify incompetence of a companies strategy / directorship, then I don't know what does.
My interpretation is... Atlassian Cloud is hosted in shards where a shard is allocated a single AWS RDS database with (say) 100 customers.
Their script was run improperly and dropped tables from 400 customers across 400 different shards (across different databases)
Their backup strategy was just to restore RDS backups (recover whole database with many customers data together)
But they can't simply restore from RDS backups due to the co-mingling of customers. As otherwise unaffected customers would now lose data.
I suspect they're manually restoring RDS backups to new DB instances and having to scrape out the data for a single customer - a process they'd not built any tooling for as it was a fault case they thought they didn't need. Possibly as running the Jira/confluence backups are very heavy and create a lot of load/disk that they can't easily accomodation in a multi tenant hosting architecture.
At least right now they are focussed on this issue and that should stop from outputting yet more unwanted bloat in my daily ops for a while.
Guess its time to pay Asana another visit to see where they are at. Didn’t feel it was worth the move 2 years ago, maybe things have changed.
Bad script! Bad! So faulty, running itself with the wrong list of IDs and running itself in the wrong mode.
Lesson learned is not have any semi-automated script dependent on human input/communication/decisions without also having a recovery plan for the worst-case scenario of said script being run with the wrong input/communication/decisions.
I doubt anyone can claim that for all such potentially dangerous scripts. But something that should be reconsidered for anything like a compliance delete script, even if the design intent is to specifically bypass the normal soft-delete process for something where it has been specifically decided that it should NOT be restoreable.
Fix in this case would be to make it a two-stage process: first an easily revertible soft-delete, evaluate the impact, and then a hard delete designed to only delete data that has already been soft-deleted. Never hard-delete active, live data directly.
It's also not clear from the post if someone had used the "permanently delete" flag accidentally before. It's very possible they had, but hadn't ever run it on a wrong id before.
While yes, someone ultimately did hit "enter" on a command with the wrong ids and the wrong mode, the existence of such a script enabled that mistake to have enormous consequences. Calling such a script "faulty" feels reasonable. In the same way having tooling which allows you to delete prod database disks without any additional safeguards would be called faulty.
It sounds like they have a "type system" problem, e.g. you shouldn't be able to give a script an ID for a "cloud site" when it was expecting an ID for an application. That is, this sort of outage should not have occurred because it should not have been representable to begin with.
In fact, recognizing that is precisely why you can’t even do that easily anymore. Setting aside that you probably need to sudo the function, rm prevents you from running the command unless you pass an additional long flag that basically say you really really want to do this.
Programs which expect perfect user input, and allow minor user mistakes to lead to major negative outcomes are faulty programs.
Nonetheless, even with rm you can override the check you mentioned with a flag. In principle, that seems to support the post you are referring to, not refute it (not in detail but in spirit).
When I do big deletions in an automated fashion, I always make it a two-step process.
The first step identifies exactly the items which will be deleted.
This gives you the user a chance to look at it and say, "Why are we deleting 100x as many entries as we expected?"
In many tools, this is as simple as running it twice, the first time with `--dry-run`.
In a microservices-centric architecture, it’s entirely possible that the script was used to orchestrate the change across a number of other systems, and an update on one of them led to to undefined behaviour and/or unforeseen consequences.
Hopefully this event has enough of an impact for them internally (and financially) that they seriously reevaluate things and maybe invest in some SRE capability, and better tooling.
I'm sure that they're going to make their customer at a time thing properly automatic now... Still a big failure, but a little more subtle than you make out.
(disclaimer - worked for Atlassian years ago)
I am actually more surprised that they are able restore the data at all. If you can recover from a Compliance Delete, is it really a Compliance Delete?
By removing data and only keeping a 30 day backup any compliance deletes still simply take ~a month to complete.
My team and I run emergency services communications infrastructure with 99.999% availability requirements and our company has to pay an abatement if we don't maintain those levels. That's actually mission critical. We don't run ANY script or SQL command in production that deletes data without having tested it in our dev/test environment, documented the steps of the change and the results of testing, having the change reviewed internally, and approved by the government.
We can get these changes approved and into production in around a week normally. "Move fast and break things" is great, until the wrong thing breaks.
Lesson learned is not not have any semi-automated script dependent on human input/communication/decisions without also having a recovery plan for the worst-case scenario of said script being run with the wrong inputs.
This reminds me of the time I was tasked with changing the gender of a user to female on the production database (since the feature was not available through an interface at the time) and so I opened SQL Explorer on our live database and ran a quick update. Except I forgot the UserID param and accidentally changed over a million registered users to female.
A colleague came up with the idea to use the person's title as a hint to repopulate the database, which worked for Mr, Mrs and Ms.
I apologise now to all the doctors that found they had undergone nonconsensual gender reassignment.
https://dev.mysql.com/doc/refman/8.0/en/mysql-tips.html#safe...
"Something went wrong. We're moving mountains to get it sorted."
Works as axpected.
Edit: After reviewing the snippets below looks like the script has gained sentience.
This phrasing is ambiguous on a major question.
If the reality is that no customers had more than 5 minutes' data loss from prior to the incident, then it'd be good to say that.
Not to mention all the data that was lost, because you can't create it, because the servers are down.
Does this event make you re-consider using their software?
Also, I'm not clear if the site was down for all customers for a few days and now only 400 (0.18%) have an issue or if it was only down for only 400 customers. After looking at many tweets, comments here and reddit it looks like it was down for everyone but I'm not sure.
Still odd that they didn't have any tools in place to help them with it. Only having a way to restore an entire backup is only useful for disaster recovery, it's nigh useless to correct mistakes.
Something went wrong. We're moving mountains to get it sorted. View our status page and subscribe for service updates.