Atlassian: We estimate the rebuilding effort to last for up to 2 more weeks
twitter.com
twitter.com
Worse, outside of "we have rebuilt functionality for over 35% of the users", I haven't seen any reports from the people who have ostensibly been recovered.
Next, their published RTO is 6 hours, so obviously they must have done something that completely demolished their ability to use their standard recovery methods: https://www.atlassian.com/trust/security/data-management
Finally, there have been some hints that this is related to the decommissioning of a plugin product Atlassian recently acquired (Insight asset management) which is only really useful to large organizations. I suspect that the "0.18% impacted" number is relative to ALL users of Atlassian, including free/limited accounts, and that the percentage of large/serious organizations who are impacted (and who would have a use for an asset management product), is much higher.
Those last two are examples are very unlikely but no company is going to say RTO = "probably 6 hours but it could be three weeks if we get ransomwared"
- Some employee has root access to AWS account and uses it operationally
- Given wildcard S3 permissions to an IAM user and allowing delete bucket
- Not enabled object versioning
- Cross Region replication not enabled
- no large bucket protection
- don't have basic security monitoring and setup of Cloudtail alerts
- have not invested in full fledged tools for IDS and so on.
If some vendor have any of these issues I don't think any customer would approve these software to be used, these are not normal or best practices .Large apps have detailed playbooks on how and what gets turned on in what order, and most do DR drills and time those runs periodically. These are well established workflows in any large org.
Yes in a real world downtime you can't have planned for every scenario, maybe you miss the target by 25 % like 2 hours more, or maybe in a very situation you double or even triple it say 12-18 hours. You don't go from 6 to 600+ .
The way RTO is calculated starts by looking at limits on cloud/ hardware / bandwidth/ machine sizes, if basic limits are not factored in like cross region concurrency there is no point in RTO being computed. Even if something like that was missed and you spend tens of millions of dollars on AWS then AWS will work with you and relax those limits .
100x missing the plan either means extremely poor planning or they screwed up something very very badly.
e.g. versioning may be enabled, but not cross region replication because it is cost prohibitive. Someone runs a job to clean up a bucket that includes deleting old versions. They point it at the wrong bucket or wrong path in the bucket. Or a malicious user does it on purpose. Monitors and alerts really tell you after the fact that you now have a major problem.
Also limits (like cross region concurrency) may not be known about until it is time to actually do a mass scale restore. DR tests might have been done but only in isolation of one app at a time. By the time you realize your mistake you're dealing with physics. Maybe AWS can bump it a bit to help you in that particular circumstance though.
No idea what happened at Atlassian. My only point is it is very hard to get it right without a huge amount of effort.
However going like 100x is not probably cause this is hard to get 100 % accurate it look more likely deleted data as being rumoured and more importantly not actually having functioning backups that were ever tested and manually reconstructing from logs and other sources.
More than just RTO, they are not going to be able to meet RPO objectives for affected customers , depending on how much loss that is going to pretty bad.
Like most airliner accidents, this is probably an unfortunate combination of both of those things happening at the same time. My guess would be they have fairly decent planning overall but there's one (or more) small-ish areas where their planning is extremely poor - which crossed over with a screwup in a very specific fashion that laser focussed on that particular piece of poor planning. The "this can never happen" immovable object and the "You can't do that" irresistible force.
SaaS has lot more tolerance for failure, so my money is it is something simpler but difficult to get implemented in large org.
---
In a ideal world this incident should impact their revenue, growth and stock price substantially.
It is unlikely to do so, because of stickiness of enterprise customers, no better alternatives, compared to say Google, Facebook, Amazon where a minute of downtime is immediate quantifiable revenue loss so FAANG really obsess so much over how many 9s of uptime.
The typical management of enteripse app companies like Atlassian have no incentive to do anything beyond cursory lip service and get away with under investing in tech.
---
[1]3 years back i would have stood by that, but after Boeing 737 max twin disaster and systematic problems leading to it , i am not so sure those lessons are not forgotten.
you fail over to warm replica;s that are already staged with data with in your RPO.
Gray•Duffy, LLP Settles Two Massive International Data Loss Claims Arising From Computer Server Failures
Arc Touch, Inc. vs. Atlassian PTY, Ltd
http://m.grayduffylaw.com/?url=https%3A%2F%2Fgrayduffylaw.co...
https://www.theregister.com/2012/05/09/atlassian_cloud_stora...
In my limited experience these diffs can be missing information. I recently had to reconstruct an issue description using these email diffs after two people where editing the description at the same time and it was not 100% accurate, several lines were missing. Going to the 'history' tab on the issue I was able to get the missing lines however, if all you have are emails though you might be out of luck.
When a node went down there was no hope of ever coming back up unless you shut the other nodes down as well. While this was going down of course none of our ingress data was being inserted so it built a queue. When we turned things back on the queue would overload vertica again and we had to repeat the whole thing.
Fortunately for us we only stored analytic type data on vertica where customers usually were only interested in the last few hours anyways. So we ended up deleting all historical data and just reprocessing it over months while occasionally prioritizing customers that complained.
3-2-1-0 Applies to all data, at all time, in all places
3 copies
2 different formats (ex. HDD, cloud)
1 offsite
0 lost data? :D
That's not what "format" means, it's more like DB-Dump and DB-VM-Dump, or pure Files and VM-Dump, or something like restic-repo and rsync(pure files).
The idea is to not have the backups stored on the same hardware or even same type of hardware. Same hardware is obvious but same type of hardware is listed because if a manufacturing defect or a known vulnerability is present it would make all of your backups at risk. So you want to have backups stored on 2 desperate types of storage media. HDD and Tape, or Cloud etc...
As far as running veeam on esxi, you would need to elaborate more on that
I had some backups that where not backwards compatible from versions that worked with ESX but NOT on ESXi. There's your "backwards compatible".
>As far as running veeam on esxi, you would need to elaborate more on that
Yeah you know exactly what i mean, no need to elaborate on that.
wow, ESX, I have not seen anyone talk about that for a long time, you must really hold a grudge...
I never used Veeam with ESX so I can not comment on that.
>Yeah you know exactly what i mean, no need to elaborate on that.
No I really do not, you do not run Veeam on either ESX or ESXI, veeam connect to vmware with API calls, so .....
I made efforts to buy HDDs from different sellers even, to avoid sequential failures from singular bad batches. That's something else I'd want to add to a "3-2-1", with regards to HDD as a form of backup or storage media.
Technically alot of Backup Planning has moved to 3-2-1-1-0
3 Copies
2 Different Media
1 - Offsite
1 - Offline / Immutable
0 - Errors from Verification Tests
My 3-2-1 comes from a personal non-professional standpoint, thus not having the extra 1-0. However I have been considering immutable offline backups, using burned DVDs or Blu-Ray discs. That's another project for another time though, for now I'm trusting paid cloud providers.
As for verification tests, hashsums are a simple solution in my opinion, but I've moved to ZFS and BTRFS to avoid having blips.
They're almost certainly rebuilding something from scratch.
If it were an AWS system limitation, almost all of those can be lifted if you ask nicely and are a big account.
There’s nothing that makes me happier than the fearsome squealing noises that enterprise sales drones make when you drop the sales equivalent of a Paveway IV on their pitch.
My favourite one was running some software supply chain compliance software on itself and explaining how it was constructed on top of a CVE riddled garbage dump.
"Bite the Bullet" is a phrase from an Englishman, Rudyard Kipling, in his first novel, "The Light That Failed". It's believed to have come from the other English idioms, "to bite the cartridge" and "chew a bullet", which date back to 1891 and at least 1796. [1]
Not an idiom I'm familiar with, but it's grammatically correct and clear so I'd count it as proper.
"Bite the bullet" is commonly used.
And it is the exact same software as Server with some extras enabled like support for multiple nodes, so upgrading to it is as simple as pasting in a new product key.
You’ve never actually had to update an on-premise JIRA instance yourself, I presume?
Atlassian completely ignored the large number of smaller customers who are legally forced to use an on-premise solution. If the software industry was so hell-bent on SaaS there would be a great business oppotunity in creating an on-premise Jira competitor.
One is that we've had many bad reports from partners that Jira Cloud is incredibly slow, even when compared to the already underperforming Jira Server and I wonder what their performance guarantees are. The other one is that it's so, so pricey.
It helped, but not in the way I thought it would! There's been no reply, but also no more spam emails now!
Thankfully we haven't been impacted by this outage.
> or maybe it's my aging Macbook
on an m1 pro: its slow as molassessometimes i have to switch to the "old ui"[0] to get any use out of it (not sure what the cause is, but sometimes its literally unusable)
[0] https://community.atlassian.com/t5/Jira-questions/Re-How-do-...
I haven’t used Jira Cloud in any great depth, but I did play around with it as part of trialling the free plan and was amused that you could quite easily come across warnings to backup your installation and consult your system administrator before proceeding…how exactly do I do that for a cloud service? Doesn’t exactly inspire confidence.
At my last job we used Bitbucket Cloud and that was awful. Dog slow, ridiculously low threshold for being unable to render diffs, and constant incidents. We used to joke that they could make the “Bitbucket is experiencing an incident” banner a permanent fixture on the page and it would be right more often than it was wrong.
We’re still using on-prem Jira at my current job, but we just migrated away from on-prem Bitbucket to GitHub, as Bitbucket was becoming infeasible and the cloud offering is a bad joke.
I have. It's painfully slow.
But slowness isn't really the problem, the problem is that it's unpredictable.
I wait for the interface to be fully loaded, so I click on a text box and i start typing. Then FU--ING something takes the focus to some other element in the web page and now i'm typing random shortcuts (like reassigning tickets, changing status or whatever).
It's painfully slow but the real problem is that it's unpredictable in its behaviour.
I think software companies need to have a serious “Come to Jesus” talk with their users about who needs to control what.
This happens to me often and it's absolutely infuriating. I'd prefer a blocking spinning wheel of death over that nonsense. It's all but ensured that I'll be looking elsewhere when choosing project tracking software in the future.
I have to say it got better after they switched to AWS but Jira not working/being slow is still an inside joke in the office
In other words, this outage is not -because- it's cloud software. It's because someone, somewhere, broke something fundamental. That can (and does) happen in on prem at a much higher rate.
* It's been deleted for a week already, they estimate they might need two more weeks. Three in total.
* They claim to have "extensive backups", and hundreds of engineers working on it.
What? How? This simply doesn't go together. Why would restoring from backup take three weeks?
Either their backups aren't complete, or they need new software written for the restore, or something else doesn't add up.
I haven't administered their software yet, but what I've learned from the sidelines, at least Jira doesn't seem to be rocket science. A database, an application server (maybe a few instances for larger sites), a bit of config, some caches. This really shouldn't take three weeks to restore.
Four drive libraries are pretty common, too.
Isn't this how an incompetent, insincere and desperate company being subjected to a ransom attack would communicate publicly?
I remember a situation where we had a near miss with data loss (replica failed and master had a bad disk). We didn't want to put the production database under extra load by taking a live backup while it was handling all production traffic, so we restored a backup. But it was "bad". Tried the one before it, and the one before that. Apparently they were busted for over a month due to a config change. We restored a month-old backup and started applying binlogs (which thankfully we had been backing up). But that meant replaying a month of transactions into the restored database. I can't remember the details but I think we ended up replacing the bad disk, resilvering the array and live-cloning the primary before the binlogs got fully applied to the one we restored from the old backup.
You can’t just do a full recovery as that would mess those customers who were not affected (it likely takes time to notice the mistake - others have continued to use the system). You might need to write some tools to migrate the data from backups. Also you really need to test everything very carefully - otherwise you might be in even deeper trouble (looking at corrupted instead of lost data).
In large organization this kind of ”manual” recovery might require people from multiple teams as no single person knows all the areas. This adds overhead. Throwing too many people in does not help either. When you start thinking about it, few weeks is not that long.
And JIRA is definitely not simple. It’s complicated beast and likely the SaaS features combined with all the legacy makes it even more complicated.
Something that can make restore-from-backups harder, and that I've seen happen, is when the backup/restore systems themselves get destroyed by the same black swan event. Then you have to first recover those by doing fresh installs, and you have to have all the people on hand who know what the configurations would have been to be able to then use the backup library. Then you have to begin restoring a few target systems to check that everything is OK with the restore process, then you have to restore everything though you'll be limited by the restore system's bandwidth.
How could this happen? Well, a disgruntled employee could make it happen. It happened at Paine Webber in 2002 [0]. In that case the attacker left a time bomb in the boot process on all systems they could reach, and that included the backup/restore servers. Worse, the time bomb was in the backups themselves, so restored systems ate themselves as soon as they were booted, which slowed down the recovery process.
[0] https://www.independent.co.uk/news/business/news/disgruntled-worker-tried-to-cripple-ubs-in-protest-over-32-000-bonus-481515.html
https://www.justice.gov/archive/criminal/cybercrime/press-releases/2002/duronioIndict.htmThe only way out is to figure out the bugs and continue migrating forward, fixing issues as they appear one by one.
https://twitter.com/Atlassian/status/1511870509973090304
Most likely they wiped the data
If that is too complicated to retrofit then have any mass cleanup script move the records to a CSV file or temporary table.
Never ever ever be in a situation where a rogue script or bad SQL WHERE clause means restoring from backups.
If you organizationally cannot prioritize quality then nothing can help you.
We had a few windows laptops where something caused them to time travel to 8000 years in the future. Then, they'd slowly spend a few hours deleting every local profile, as nobody had logged in to them for 8000 years. Then, they'd do something to their time zone database and travel back 8000 years.
When they started the process, it was unstoppable. Trying to modify the system clock to something sane just caused them to depart to the future again, even if disconnected from the network. None of our users was very amused by this behavior, even if everything important was backed up.
I agree that just having a deleted_at timestamp and old entries are never pruned would not be a good faith interpretation of the law.
> The data subject shall have the right to obtain from the controller the erasure of personal data concerning him or her without undue delay
> “Undue delay” is considered to be about a month
Tip: Begin an SQL session with BEGIN TRANSACTION; at the end you can either COMMIT or ROLLBACK.
Always use a copy of prod on a staging server and run your queries there for testing.
They’re lucky they have a sound backup strategy in place, and that the amount of data lost is appearing to be minimal.
But I would like to warn people about certain implementations of database "soft deletes" that I'm not a fan of. To be clear, I'm talking about the idea of having a "deleted" and/or a "date_deleted" column and using those columns in the WHERE clause to filter out rows that shouldn't be visible.
That pattern complicates the table structure, queries, and indexes. It increases table and index size, thus more data has to be sifted through (either table data or index data) to ensure only non-deleted entries are returned. More data to go through means slower queries. It's also really easy for people to write SQL that accidentally leaves the "deleted" column out of the WHERE clause. Then old, irrelevant data is being returned.
Accidentally deleting data that needs to be undeleted is usually rare so I don't think people should optimize for it. We should optimize for things that happen frequently.
I have dealt with the rare "Oops! I deleted important data!" by restoring from backups and it has worked fine. I think it may be too strong to say you should never be in a position to restore data from a backup. In fact, I think it's important to streamline the restore process.
For cases where we know ahead of time that we want to query deleted data I'll move deleted data to another database table that exists solely for maintaining history. For example, an ORDER table will have a DELETED_ORDER table, or an ORDER_HISTORY table. The HISTORY tables can also record data overwritten from updates.
These tables take up disk space, but never affect the structure or size of the original table and its indexes. Queries to the original table don't need to be modified to account for soft deletes.
To guarantee that things go to the delete/history tables, I'll usually put a trigger on the original table to move data over to the history tables. This way no application-specific code is needed.
That's very use case dependent.
We've made it easy for people to undelete data they've accidentally deleted simply because they used to do it so often and the only people who could get it back were our tech team. We're a devops org so part of our job is of course to support the systems we build, but our time is better spent on building solutions to business problems than to repeatedly providing support for issues that come up all the time. Part of building those systems is of course engineering in solutions that make it hard to screw up, and easy to unscrew when things inevitably do go wrong. No mean feat given our platform dates back over 15 years and still includes a lot of legacy from the time when tech was just a couple of people.
I suppose the object lesson here is that edge cases in one system or company can be part of core business in another so it's best not to make too many assumptions.
Then weekly, a task went in and then purged rows with those non-content placeholders to completely purge that user, if a user-purge was requested.
Usually this goes along with "Oh and the other team did some important work at the same time" so you can't just restore a backup. You either tell them to deal with it or start writing custom scripts to copy out only the data you want to restore.
A more sane solution would be soft delete for x days and after that it becomes a real delete.
Temporal tables should really be used more often.
As a second step, restore from the backup at a set frequency. This would force orgs to automate and optimize not just the backup flow but also the restore flow. Tear-down and restore entire systems from backups. Of course, doing so enormously adds to the cost, but when there's an outage, it will pay itself over.
That is very interesting. This implies they are backing off, at least somewhat, from their very aggressive microservice strategy. Perhaps they feel like they have gone too far in decomposing their products.
1) This outage will get their organization to prioritize work such that it never happens again.
2) This outage is representative of a dysfunctional organization that can't prioritize work correctly.
If you've been using Atlassian software for a while and are used to how they prioritize tickets then one of those options seems far more likely than the other.
It has already happened in the past.
10 years ago to the month.
So as long as everyone plans a vacation for this time in 2032 it'll be OK.
The only options now are the $$$$$ "datacenter" license, migrating to the dangerously unstable cloud, or not doing anything and running unsupported EOL software.
But the lack of transparency is the worst. Another post speculated that Atlassian has lost data, doesn't even have backups, and is re-creating it by munging their emails and diffing them to re-create history. I can't really imagine that's tue - but what if it is, and Atlassian is concealing things?
Data Center is pricier than Server, but isn't it still cheaper than Cloud? And you control your own back-ups as with Server, so Atlassian cannot lose your data.
Absolutely no one even knew this was happening and doesn’t give a shit now because it’s a project death march.
JIRA as a whole has been a fucking shit show of a product over the last decade even on-prem.
We self-hostd JIRA from 2008ish to 2014ish. (memory is fading on exact dates) By the time we decided to stop using JIRA, we fracking hated JIRA and would never return.
Since then, GitHub Issues, Trello, Clubhouse (neé Shortcut) all provide less friction in day to day use. As an Enterprise, I do believe Shortcut is your best bet.
> Clubhouse (neé Shortcut)
It's the opposite: Shortcut (née Clubhouse), as "née" means "born", so it's the name it had at birth, the old name
The more you know!
I was kinda like....."not really the point of bringing it up."
Its worth noting we have had them just delete things within our account before. In fact one of our Senior VP's had their account just....disappear one day. We couldn't @ them in chats, tickets etc. Atlassian just shrugged and "restored" the account and said it was some issue with a stored proc on their backend or something.
I have always felt uneasy about how flippant they are in their processes. But it seems that is not shared.
Now it has endless competitors and I'm led to believe that it has accumulated lots of features that businesses can't live without but which the average end user never touches.
I don't envy engineers there right now, some dpts. Stay strong, don't burn out!
https://architectureau.com/articles/worlds-tallest-hybrid-ti...
Those sort of growth rates are incredible for a company this size.
They can easily afford their timber treehouse.
Sadly it seems society has assimilated vendors behaving poorly. Data breaches and incompetence does not seem to phase purchasing choices these days
Obviously not universal, but if upper management has decided to spend a lot of time focusing on a marque building they're not focusing on the business itself.
And there's still very little movement: https://community.atlassian.com/t5/Backup-Restore-articles/E...
But don't worry! It's in the Cloud! It's all fine!
Could Atlassian be liable for damages?
This would allow them access to more investors.
I'll probably should buy their stock.
Some companies cannot operate effectively without atlassian products, so a fuckup of that scale might just have legal consequences depending on whom it hits.
Kind of their problem to be frank.
> so a fuckup of that scale might just have legal consequences depending on whom it hits.
Any contract those companies signed would have a cap on the retributions by Atlassian for trashing SLA targets
In the external services I use, downtime of one service or other is to be expected at least a few times a year, and the “sunsets” happen occasionally.
Thing is, public services are solving a much more difficult problem (keeping things running safely for millions).
All that to say, I don't think it's a weird phenomenon, it's just you're realizing that you're paying someone else for something that's not delivered on.
so selfhosting may still have certain upsides even with such outage
In this case, it seems like the company took a risk and it did not go well. The possibility of being able to restore from backups might have been factored into this risk, but the latency of doing so might not have been.
If you're self-hosted, you dedicate as many people as possible/necessary to restoring your service, and it becomes their top priority.
You also have a lot more insight into the detailed inner workings of the restore, making it easier to plan against, instead of just vague "we're working on it" messages for days a time with no clear end in sight.
I’ve never – not at any point in the past 10 years – gone over 24h of downtime.
JIRA will now have 3 weeks downtime.
Distributed systems have complexity that grows superlinear, which leads to more and longer incidents.
I don't want any of my company trapped on it, but if they were I'm sure as well not going to self host that spawn of hell.
Mostly downtime is just upgrades. I can remember a few times we've had to add (JVM) memory as our usage increased. Not sure what we're going to do with the discontinuation of server product line. We self-host to keep source code, etc. more than one configuration mistake (or zero-day) away from exposing it to the world.
But then, there is no way to keep it in sync. I have to blow that project away in jira cloud, and migrate it again.
So I Have to hard-cut over projects, on a system that has dozens and dozens of projects, and somehow have people figure out which ones are where. or one really, really ugly night to cut it all over, and hope it goes well.
I'm looking for alternatives, but our team is so invested in some very, very customized workflows, its going to be a pain.
I think one reason Atlassian was successful is that they always invested a lot of effort in building tools to migrate to their products from any of their competitors (obviously not the other way around).
Maybe. But you're counting on your sysadmin(s), who are also managing dozens of other things, to keep up to speed on Jira and its quirks, and apply patches and new versions as they become available without missing any steps or screwing something up.
On average, you're still probably better off having a company that knows the product also host it for you, but obviously they can make mistakes too, and the downside is that when they do it might affect all clients, not just one.
This is also potentially an upside. For example when us-east-1 went down recently, customers were somewhat understanding because it was "amazon's fault" and everyone was down - it was in the news, etc. If we ran our own data center and that went down, our customers would've just said "why did you morons roll your own data center instead of just using aws?"
#3,198,191 Don't automate deletion scripts w/o sufficient recovery options.
What I will say is it's important for the customer to HAVE THEIR OWN BACKUPS. Don't rely on the vendor - that's the lesson here. If you have all your stuff in AWS back that data up someplace that's not AWS, etc.
Some examples: https://www.ownbackup.com/ https://www.backupify.com/ https://www.cloudally.com/
Edit: to clarify I would say if the data is important to you, then “the ability to back up the data” should be a requirement when selecting saas. See my other comment in this thread on ms planner.
Even for Email or storage or any other open system, UX changes and feature differences can take a lot of time to train properly, you don't migrate from one vendor to another vendor just like that.
That would be too good to hope for.
Since Trello is part of Atlassian aswell - what are good, reliable and above all lightweight alternatives for managing projects without the “pseudo-agile” rabbit holes of functionality?
My current employer shut off Trello and forced us over to Jira and is threatening to disable Planner, so I'm "not allowed" to rely on Planner enough day-to-day so it's possible it is either better or worse than I remember it being in that department. But this Jira outage has me reevaluating, and they haven't turned off Planner yet.
Realistically, though, you're more likely to be able to convince other project members to use the issue tracker of whatever forge they're comfortable with, for instance Gitlab, Gitea, Pagure or Sourcehut.
I've been pushing for an exit from Jira for a little while now, but this doesn't really add much ammo to that argument for me. It's like pointing at a plane crash and trying to justify the company no longer fly people places.
Their competition is equally bad and offers far less features for larger programs.