Incident Report for Confluence
confluence.status.atlassian.com
confluence.status.atlassian.com
This does not suprise me at all. I came by acquisition years ago, and was wondering when something would like this happen. They've deleted internal Slack, internal wiki before. Nobody cares about stability or scalability at Atlassian, their incident process and monitoring is a joke. More than half of the incidents are customer detected.
Most of engineering practices at Atlassian focus on only the happy path, almost no one considers what can go wrong. Every system is so interconnected, and there are more SPOF than the employees. So there are million ways this could've happened anyways.
Even a minor feature requires 6 months of planning with 3 timezones, while architects and middle-management in Sydney overriding most of the decisions and there is constant reorganization. There is a general sentiment of "We've built JIRA, our stock increased >10x, so everything we say and we will do is true" and all of the other ideas are invalid to them. Also, all other products other than Jira is second-class and not important.
Actually using JIRA on the other hand…
Some customers are also told they should expect to be down for up to another two weeks, making the outage a total of 3 weeks or so [1]. 400 companies are impacted by this outage and at least one of them is a YC company, Bitrise, who are still down, Atlassian is not telling them anything specific, and data loss is a possibility for them [2].
The outage is strange to last this long as it contradicts Atlassian's own admission of how they plan for resillience and their statements of how quickly they can restore data for customers. They specifically say how they utilize multiple DCs, have zero data loss scenarios and their recovery time objective (RTO) is 6 hours for their Tier 1 products (like JIRA). They are failing this objective big time, which questions what is happening behind the scenes. [3]
The outage is especially ironic given that Atlassaian have discontinued their self-hosted Server product and are forcing all customers to move to the Atlassian Cloud - this very cloud that is down for many customers. They stopped selling Server licenses in Feb 2021, disallowed changing tiers in Server products this February, and are stopping support for Server products in Feb 2024.
This long of an outage for software which companies depend on is mind-boggling. The only similarly long outage I can recall is the 2011 Playstation Network Outage which lasted for 23 days.
Good luck with the restoration, which sounds like it's manual work. Hopefully we get a public postmortem from Atlasssian: they are not the company known to publish these, though.
[1] https://twitter.com/kjartanmuller/status/1513462616030683138...
[2] https://twitter.com/gabornadai/status/1513481270411636738?s=...
[3] https://www.atlassian.com/trust/security/data-management
I've seen supposedly very secure and extremely well run orgs lose some or even all of their data to - retrospectively - ridiculously small chances and yet, they happened.
You should only consider the current situation, and plan for the future.
We all know they are working on breaking encryption and anonymity that comes from it, so if you don’t think they would at least allow some major event to happen just to roll out the iPatriot Act, then I would like to talk to you about the once in a lifetime opportunity to buy a bridge in a highly desirable location.
Which would be folly under the surveillance schemes you warn of!
Never threaten with what some may legitimately consider a good time. You'll be astonished who comes forward.
I don't consider the surveillance a good time, for the record.
Do yourself a favor and plan the sunsetting of every piece of software and service that you use for your business. At the moment you start using it, plan how you will eventually move away from it. It always happens, and planning ahead leads to much fewer headaches.
This applies to life in general, too; have an exit strategy.
My Multics and MVS Tk4- runs pretty well, my DosBox and Nintendo Emulator too and hell have you checked FreeDOS?
>Services come and go
Yes
>Code dies
Don't have to.
I was working with a DevOps team a couple of years ago and they had written some projects to deploy internal web apps to AWS. It came time to remove one of the apps, but the removal date kept slipping. After much hand wringing and meetings, it was discovered that they had no mechanism to uninstall one of their apps. They simply hadn't considered that an app might not be used anymore.
Nor did they ever consider needing multiple instances of an app, so all their stuff was using hard coded names.
It wasn't a fun project.
Maybe the problem isn't that code eventually decays maybe the problem is the vision of what something is supposed to do is lost and then it has no place later on.
It's not that our values as engineers have shifted, it's that old software was built on static platforms that rarely changed. New software has to constantly adapt to changing conditions in terms of security, stability, and performance.
Cloud-based services inevitably die, sure, but this is not true of software or machines in general. There's an entire YouTube genre of people finding and restoring cars that were built 30-90 years ago and abandoned for decades and getting them to run. We just don't build things to last any more, but a lot of that is by design. Subscriptions and planned obsolescence lead to more revenue.
Atlassian could have built software that doesn't rot. Step 1 is don't get rid of the individual server licenses so people can self-host. Step 2 is probably don't use Java because eventually users will have to upgrade to a JVM that isn't compatible to avoid security problems. 64-bit ELF will probably work fine for another 60 years. Step 3 is open source it so users can build it themselves since you'll eventually hit binary incompatibility problems otherwise. Do all of that and there is no reason Jira couldn't be as durable as sed and grep.
[Citation needed]
I trust the managed language more in this regard. I don't know specifically the situation on the BSD side, but on Linux userland, once it uses anything that is not the Kernel or glibc, it's anyone's guess if the ELF will even run if not on the exact same packages on the same machine that compiled it.
“Successful” is doing an awful lot of work in the title.
If you blew up our entire AWS account save our backups we could be nominally back online within an hour or two and fully within 6.
I am genuinely puzzled what could be happening with their infrastructure to cause such an outage. Are they self hosted on bare metal and there was a fire or something?
Do you regularly exercise this, or is it just a plan?
Our Chinese installation for instance runs on servers controlled by the Chinese government. I believe they run an AWS compatibility layer from what I've heard, I haven't personally worked on this.
We've got a decent amount of experience bringing the whole thing up from scratch, and we build with that in mind. Most of the time in my estimate is for importing the larger detail tables, which we do regularly for our staging environment.
They may also be gated by limitations in-between customers and the backups. Like backups going through a transit VPC, for example. You don't see the bandwidth limitation when you're sending incrementals all day, but try and do full, parallel, multiple customer restores through it, and ... (and perhaps they are afraid of making too many changes on top of the current chaos to fix something like that)
What makes you say this?
Also matches with it being a cross-product outage.
But just a guess.
Though, going by their speed, this seems to be a weirdly manual process, not sure what is going on there.
I always wonder if most SaaS providers are fixated on building massive systems when simple shards would work for most customers. If you compare something like Gitea and GitHub you can probably see the reason. With Gitea your "shard" is siloed away from everyone else because there's no federation (yet?). GitHub is more like a huge social network where everyone can interact.
I think the vision for a lot of SaaS is to facilitate B2B collaboration by acting as a proprietary social network rather than via federation between relatively independent shards. The reason is simple IMO; vendor lock-in.
- GitHub has way more features. More features means more scattered data (e.g. different micro-services databases, file storage).
- You cannot use Gitea in an organization with hundreds or thousands of engineers. It provides nether the features nor the scalability you need. I like Gitea but I'm happy that my company uses GitHub.
- Individuals and open-source use GitHub because it's free and has all the features (including stuff like free CI/CD which is expensive). This is all being paid for by companies like mine who buy expensive enterprise licenses for thousands of engineers.
- It's not uncommon to shard by customer/account, even in the cloud and GitHub is actually trying to do this right now for their primary database. Doesn't solve the problem with different storages for different micro-services though.
I hope they get a bonus and some time off once this is resolved.
It's almost the same as saying life is the cause of death in a human.
Change is life, software that don't change is dead software.
"sites were disabled", making it sound like they flipped a bit in the database, when they have clearly destroyed a lot of data.
Jira is emotional abuse software.
To me that shows again that self hosting often beats cloud services. Yes, the vendor knows more about their own product, but delivering services at scale is so much more complex than a small on-prem instance that on prem usually shows higher resiliance. I had the same experience with self hosted Gitlab and Mattermost (as a Slack alternative).
That's also the basis for the cross-org reporting that makes the MBA types happy. Yes, you can debate the usefulness, but that's who is signing the checks.
As a long time user of both, I'm much happier with Gitlab, but manager-types still struggle a bit, because too much technical clutter in their UI (they just want issues/boards, not CI/CD).
I got used to less features, and kind of see it as a benefit. We tend to push a product to its limits, then throw it out the window when it becomes to complex.
It still has many positives as an alternative to GitHub that also happens to be open source, of course.
The downside is that DC is about twice as expensive.
If you're using any paid add ons, you'll probably need to pay more for them as well.
It turns it really expensive if you've only got 50 people. I can't see many businesses that size being very excited.
Whether that's a risk people are willing to take is another question.
Since Jira is entirely development tools, it sees no benefit to being hosted offsite.
(We're a 1-dev startup so YMMV)
https://fossil-scm.org/home/doc/trunk/www/index.wiki
Before the JIRA days there was https://www.redmine.org/ as well.
Also side note, I feel bad for the people working those 24/7 shifts. Burnout starts to kick in and more mistakes can happen.
That's not the title.
Quoting from the last update on the page:
no reported data loss
> we have rebuilt functionality for over 35% of the users who are impacted by the service outage, with no reported data loss
The page title is bereft of more information:
> Multiple sites showing down/under maintenance
I can’t really imagine how you would design a system in this day and age that could fail in such a fashion. But anything that discourages the use of Confluence is ok by me.
I mean the output of one three-week sprint (that’s their cadence) was “we are turning some stuff off.”
https://docs.microsoft.com/en-us/azure/devops/release-notes/...
Looks like most of their sprints are adding things or fixing things, not removing things.
https://docs.microsoft.com/en-us/azure/devops/release-notes/...
https://docs.microsoft.com/en-us/azure/devops/release-notes/...
https://docs.microsoft.com/en-us/azure/devops/release-notes/...
https://docs.microsoft.com/en-us/azure/devops/release-notes/...
To three years ago:
https://docs.microsoft.com/en-us/azure/devops/release-notes/...
https://docs.microsoft.com/en-us/azure/devops/release-notes/...
https://docs.microsoft.com/en-us/azure/devops/release-notes/...
https://docs.microsoft.com/en-us/azure/devops/release-notes/...
To be honest at this point I don’t see us ever switch away from BitBucket as our environment (non tech employees use SourceTree, bitbucket pipelines etc) is centered around it.
Meanwhile we still have some self-hosted services (like SVN) which has problems relatively frequently which require manual attention from engineer's time. I probably spent more than a week maintaining our self hosted services last year, compared to 0 minutes for the cloud services we use. Not even including that our self-hosted services are probably not very secure since we sometimes forget to do updates on our servers for a long time.
Bet they wished they spent those hours elsewhere now.
“Affected” is doing a lot heavy lifting in this headline. The linked page states that a small number of customers are affects. So it’s 70% of a small number. Where “small number” is defined by Atlassian so we don’t really know the impact here. As best as I can tell my own usage has not been impacted at all.
I also don’t presume everyone to have read every thread on here every day.
I think Atlassian are probably not helping themselves by calling the number of customers "small" rather than saying the actual number or specifying a percentage of total customers.
This is extremely forward of me, and I hope you'll forgive my professional proactive nature here, but I'm a member of the ClickUp sales team and I'd be happy to learn about your teams and see if we can better serve your needs.
If you're interested, or you think your organization could benefit from at least TRYING to see if there's another solution out there, feel free to schedule a meeting here:
https://a.clickup.com/c/coreyelder#/select-time
Humbly,
ClickUp Cowboy