Inside the longest Atlassian outage
newsletter.pragmaticengineer.com
newsletter.pragmaticengineer.com
1. They can confirm that they have backups of our data (about a thousand stories, substantial confluence, opsgenie history, and three service desks).
2. Will our integrations, configuration, and customizations also be recovered, or will we need to rebuild those once our data is recovered?
I have received no response, and no human is even willing to acknowledge those questions. The service desk staff ignore them as if I never asked. Repeatedly.
Also, I've been asking around, and haven't been able to find a single story from somebody that can confirm that they were down, who has had their data recovered.
1. They do, every 25 hours, via snapshot. I have spoked to their team since the incident and that same thing is in this article. 2. Yes, they recover all of it. Some things have had issues, external mailboxes attached to service management projects, some attachment rendering slowness. Filters needing to be overlayed into our instance again, but otherwise it is running fine again.
Not sure what to tell you other than they are fixing life saving companies first, then the rest. That is what they have told us.
If the script was used in "permanently delete" mode, which is intended for compliance... how do you restore?
Is it the only explanation... if the deletion is non-compliant?
> Second, the script we used provided both the "mark for deletion" capability used in normal day-to-day operations (where recoverability is desirable), and the "permanently delete" capability that is required to permanently remove data when required for compliance reasons.
[1] https://www.atlassian.com/engineering/april-2022-outage-upda...
https://www.itgovernance.eu/blog/en/the-gdpr-how-the-right-t....
I once worked a company that had a data loss issue. There was nothing else we could do, we had exhausted every option we had over almost 40 hours. At the end of the second day, it was decided to restore from backup.
We had done this before, as a test. It took about 12 hours to restore the data and another 12 hours to import the data and get back up and running.
One small thing was different this time, and it had huge consequences. As a cost-saving measure, an engineer had changed the location of our backups to the cold-storage tier offered by our cloud provider. All backups, not just 'old' ones.
This added 2 additional days to our recovery time, for a total of five days. Interestingly enough, even though we offered a full month's refund to all of our customers, not even half of them took us up on it.
Business-wise would be to stay in their good graces and keep those customers by offering the refund, but you don't lose any money to those who either don't care or won't move to a competitor.
Focussing on communicating open and honestly allows them to explain the crap they’re going through because of your mistakes to their bosses, so in fact you can help them save their asses, and they’ll save your ass in return. This is much more important and valuable than a refund.
So you should ALWAYS communicate open and honestly, and offer the refund as an option for clients who do not have a boss to account to.
2 hours later I walked back to see what they found. I figured it would be several hundred dollars for a new clutch, and I'd have to borrow money or something to get it done. I talked to the owner who told be it was an adjustment on the cable. Just needed to be scootched up a bit and it was probably good for another 30k miles.
When I asked him how much I owed, he laughed at me and said, "For that? Not worth writing it up. No charge. You want me to show you how to do it yourself next time?"
The shop could very easily have charged me 1 hour of labor at their standard rate, maybe $75 or so. Plus a diagnostic or test drive fee. Whatever. He could have told me, "$123.98" and I would have paid it. I wouldn't even have been mad. But I sure as hell wouldn't have remembered the experience so clearly. Nor would I have told a dozen people over the years to take their cars there. And I definitely would not have driven 20 miles out of my way to return to that shop in the future years.
Being cynical about this stuff will hurt your brand. It's not obvious. It doesn't show up on the earnings report as a line item. This is service segmentation that seems like a no-brainer to a clueless MBA, but actually matters in the long run. How people view your brand is immensely important.
Not forcing customers you already screwed over to then spend more time chasing a refund is not only the right thing to do, it's also good business.
If you were charged $123.98 and you said, "hey, I told you where the problem was, why am I being charged a diagnostics and driving fee?" and they corrected it by telling you the whole thing is on the house, is that not good business sense?
Even by your own admission, you would have gladly paid that $123.98 with no issues and you wouldn't have been mad about it. So from a business perspective, if they can provide a service, get paid for it, and the customer has no qualms or issues with the transaction whatsoever, in what way is that hurting the brand or being cynical? I think that's a much more business-wise action to take than to give away your services.
No. I'll be happy that I saved on the money, but I won't trust them in the future. They're now "the place that tries to get away with things" in my mental Rolodex. Better to stick with the fee and know their value. (I didn't tell them where the problem was. All I knew was that the clutch wasn't grabbing anymore. I assumed it needed a whole new clutch.)
> Even by your own admission, you would have gladly paid that $123.98 with no issues and you wouldn't have been mad about it. So from a business perspective, if they can provide a service, get paid for it, and the customer has no qualms or issues with the transaction whatsoever, in what way is that hurting the brand or being cynical?
It would have been a fine decision, sure. But in that case that would likely have been the only business I did with them. Not out of spite or anger, but because I'd have no reason to pick them for future business. I would instead ask friends for recommendations, or pick some place closer to my future residences.
But what actually happened was that I was the one steering people to them. I also went out of my way to return to them for brake jobs, simple oil changes, etc. I was a loyal customer, and probably spent or caused others to spend over $5,000 there.
He had absolutely know way of knowing that would result. But if you just treat people right, the way you'd want them to treat you, you build a reputation. It pays back.
I know this story comes off a bit pollyanna. I get it. For a cynical and non-altruistic explanation: when it takes a technician literally 5 minutes to twist an adjustment nut and verify that was all there was to it, stop and think about the bigger opportunity before you robotically mark '1.00' in the "LBR HRS" field on an invoice. Especially if you're operating in a field that's notorious for rip offs.
> I think that's a much more business-wise action to take than to give away your services.
I'm not saying businesses should give away major services. But they should avoid the temptation to nickel-and-dime as well. That's on the other end of the optimization curve. Not good business.
I think he absolutely knew that building trust is key to solid, long-term, repeat business - not only from the direct customer whose trust he has earned but also the zero-effort initial positive trust-balance he will have with his future/potential customers, even before he has done anything for them, just via word-of-mouth referrals. Such a simple concept but it just doesn't compute for some people.
> But if you just treat people right, the way you'd want them to treat you, you build a reputation. It pays back.
Couldn't agree more.
And if they can't sustain that, then it's even more imperative that those customers migrate away.
“Corporations are people too”
When the entire service is hosed, that's a totally different set of circumstances, and you have to look at what the RTO/RPO are for basically restoring the entire service for all customers. And since the have more than a thousand customers, it totally makes sense that it would take orders of magnitude longer to restore the entire service.
I'll also ask (since nobody else has answered, I may as well ask you as well):
1. Are the customers actually being restored from backups (and additionally, by a standard process)?
2. Will the recovery also include our integrations, API keys, configuration and customization?
It is explained here that Atlassian runs regular DR planning meetings with the engineers spending time planing out potential scenarios, as well as quarterly tests of backups and tracking findings from them.
So, with those two things happening, I the imagine recovery time objectives of <6 hours was taking a typical "we deleted data from a bad script run affecting a lot of customers" scenario into account with the metrics from the quarterly backup tests.
That doesn't even come close to the recovery time we are currently seeing now however. We're coming up on 2 orders of magnitude more than that.
The above doc seems pretty far our of line with what is currently happening.
[1] https://cloud.google.com/storage/docs/storage-classes#coldli...
This hasn't been written up at The Register yet, so I don't have a single URL I can share with you.
I think many would have worse uptime even with more headcount
For example, for our own service: If you have a hundred or two hundred licenses, you can drop our system on a linux box and usually you have to throw a yum update and one or two service restarts at it every few months and it just works. I honestly wouldn't be surprised if many of our small on-prem solutions have better uptime than the SaaS clusters, or be capped in uptime by some externality, rendering the system downtime irrelevant. If their VMWare cluster is down, our system is down, but no one cares.
This also mirrors a lot of our internal systems. At a small scale, you can just dump chef, jenkins, sonar, nexus, whatever on a linux box and forget about it.
However, this changes with high license counts. We have singular customers in our SaaS offering that are more than 50 - 100x bigger than the small on prem systems. At that point, our SaaS offering is better than anything the customer could to on-prem. I'm confident to say this about all of our customers, except maybe 2.
If anything, a smaller company with smaller footprint and fewer total requirements is going to be more likely to manage a vertical slice of some SAAS product.
The reason things like github go down so often is because they are public/shared resources.
Very much this. Managing shared resources at scale is pretty hard. We have a bunch of internal sites made by interns as part of their internships, and, funny enough, those sites have much greater uptime and appear more stable than our own multi-tenant SaaS solution made by seasoned devs.
Experienced people hosting and tuning Atlassian products has a greater success rate than someone doing it alone for a large company. Almost every time I’ve migrated an old Atlassian installation under our wing it’s given me shock how users have been made to suffer the loading times and perfs that come from underprovisioning (db or actual machine) and messy configuration. I’m not blaming the former admins but it just happens. Usually end users are happy after we clean the mess up and everything feels snappy.
Disclosure: I’ve worked in this kind of expert role.
A lot of these SaaS are just glorified Rails apps with a patina of professional "security" and "reliability", and loads of extra junk that your co will never use.
Maybe i'm wrong, but the impression I had of Jira is that just like using sharepoint for file storage, the C-level people want it because they were told that's what big enterprise are using. And if it doesn't fit the need of the company and everyone hates it, they just blame the employees or lack of training.
If it's some multi-tenant solution it's no better.
Many medium-large corporations have their own cloud environments that their IT Ops control. Solution providers can host Atlassian stacks on their own cloud environment where they are not affected by data privacy concerns (it's in their already green-lit cloud providers data center) so they can host it behind a firewall with only VPN access allowed. They can also do all the magic you can usually do with web software like put a frontend proxy in front of it, or use more flexible/legacy authentication methods. Not to mention that for example you could have a Jira Cloud that you would need to integrate with a SCM program. Jira data could be "OK" to live in the cloud but code would be a big no-no. These problems can be solved by having them all live behind the firewall.
A competent managed solution provider also has consultants that can train or instruct on usage. It costs but it is simpler and faster than having to go through the forums or send a support ticket for every small issue to Atlassian itself.
Cloud just outsources that problem to another business. Sure, they have better reasons to actually cover those positions and make sure they have on-calls and backup and a disaster plan, but just because you pay extra money for it doesn't actually make it work better if the company underlying it sucks.
Atlassian is in the process of killing the on-premise small/medium business option, already announced an EOL date.
Move to the cloud, buy a 500+ user solution for a much higher price or migrate away are my choices. Of course I use the local database and have local services JIRA/Confluence talk to so it's not really an option to move to the cloud.
I assume lack of competent on-site staff 24/7, having someone else to blame as well as lower costs are why people choose the cloud over on-premise though.
Mattermost is so much worse that the slowness and general issues are not worth it. And in the end it is more down than Slack ever was, because it has performance issues.
I am not sure if it is Mattermost fault or our fault; but my friend from other corporation has similar experience with it. But maybe in general just don't know how to host MM, I donno
I swear if IRC just implemented emojis.
Lately we use it more than mattermost :)
My understanding of survivor bias is that you're getting a skewed picture because some of the data was excluded completely.
Whenever you see a talk like this, always assume that it's BS. It might not be used by any real customers, or might still be in development. There might be a bunch of fires happening all the time due to things the talk doesn't mention. And it might be shuttered the next month if it's too expensive, complicated, obscure, or hard to support. These talks should only be considered aspirational sources of ideas, but never taken as a gold-standard battle-tested model, until they tell you how it fails. Only after you know how a system fails and how to respond to it can it be said to be reliable.
It is only through understanding what can fail that you can figure out causation.
And since Atlassian failed here, the talk might expose some of the failure's causes, or at least cast doubt over the usefulness of the practices presented.
I much rather running Gittea on a raspberry pi that I CONTROL than having to have the impotence of doing nothing for more than a week. + having work at cloud companies and having been requested to "collect customer data" to hand it over to the government I would NEVER move critical pieces to anyone else's infa...
(Note: I am not supporting crime, but I rather to have privacy and criminals than living on an authoritarian regime where a dictator who knows everything abot everyone keeps "peace".... Yes I am looking at you China!)
If mistakes will be made, at least I wont pay others to do them for me....
For most SMBs, it's cloud or nothing (or a different vendor, of course).
We use Jira, but it's self-hosted for my team. Maybe other teams that have transitioned to the cloud version are aware that there's a problem, but I haven't heard about it.
Granted that's how all lies start / what sometimes people assume and they're wrong but ... maybe this is that time?
Maybe it is in fact so bad that honesty would be a push or worse?
In my opinion, such a scenario does not exist. Transparency always in all things.
What's the right way to structure your data here that would make restoring more straightforward here? Is this backup/restore scenario niche or they should have designed for it?
a) overwhelmed by creeping featuritis, each customer's data has relationships to global tables, and
b) they backup their entire database cluster in one snapshot
and there maybe other gotchas for restoration, like relying on denormalized views and caches that have to be rebuilt. they may also have erroneously assumed that data protection's main value driver is whole-of-system disaster recovery, which can lead to pathologies such as "we don't have a single-customer restoration tool".
this is not a niche scenario
What are the downsides to this?
* disaggregates data that the SaaS might be interested in querying/updating as an aggregate
* not all ORM frameworks handle this case well, if at all
* dumps are more than a single trivial command
basically all your data operations gain an additional dimension of complexity, and you may not perceive the benefits until much later
typically this is probably for internal reporting/metrics. But yeah, a custom script with direct SQL is in order. Personally my opinion is avoid ORM at all costs. Never seen a benefit that wasn't trivially done in SQL, and the downsides are incredibly painful.
The big downside of sharding out, per customer, is that's a lot of databases to migrate on upgrades. Or rollback if shit hits the fan.
The upside? You can have customers on different versions of your app if you really wanted to do such a thing.
In any case, proper tooling goes a long way to making it the difference between wonderfully manageable and torturous nightmare. Think idempotent backup scripts that are capable of failing at any time and resuming where they died, etc.
At least explain why there was such a total communication blackout company wide. Even support staff weren't allowed to discus it. Why?
A standard RFP question for SaaS should be:
- Can you restore data for a single customer, and if so, what is the RTO for that operation?
A smaller SaaS could be excused for only thinking about full database restores. When you're a scrappy upstart, thinking about hypotheticals is less important than survival.
But for any decent size multi-tenanted SaaS, it's imperative that you have the ability to selectively restore individual customers.
The usual approach is to do a full database restore into a separate instance, then run your pre-prepared "restore customer" scripts to extract a single customer's data from there and pump it across your prod instance. In Oracle for example you might use database links to give your restore code access to prod and also the restore instance at the same time.
Atlassian - MUST DO BETTER.
If you contact them and say "please restore our data to as it was last week" those I know do not offer this.
It was initially a cover-my-own-ass design, but it turned out to be an extremely popular feature that was never even used for disaster recovery. Instead, it was used for audit support, trial scenarios, projections, and all kinds of other stuff.
We restore deleted accounts on request sometimes. There was a client, for example, who forgot to renew the subscription and did nothing for 30 days, so their account was automatically deleted. We restored it from backups. It helps that every tenant has their own isolated database, so it's mostly a matter of restoring that one single DB. Some microservices store data without DB-level sharding, so we have a script which is able to make a partial dump for a specific account.
There's also a popular option to restore deleted data - nothing is ever hard-deleted (it's marked deleted but stays in the DB) and we have a script which can restore individual records (and related records). There's maybe 5 such requests per month.
We don't offer rolling everything back to a specific point in time, though. Technically it's possible by undoing the event queue but it's untested.
We also have a script to migrate customers from cloud to on-premises and back.
I really wonder what these Altassian restore tools look like it if takes "hundreds of engineers across the company" to restore 400 accounts. Are backups siloed across many teams?
I’d rather use request tracker or bugzilla over Atlassian these days
Is it standard to (in addition or instead) to have something more general/forward-looking like: how do you watch other providers' postmortems and apply the lessons to your own system?
> - Can you restore data for a single customer, and if so, what is the RTO for that operation?
If I were to aim something at this specifically, it'd be: can you restore data for N customers or N% of customers, and if so, what is the RTO for that operation?
I mentioned in another comment that Gmail had a similar outage in which they had to restore from tape. https://news.ycombinator.com/item?id=31017160 They had a tool for restoring a single account but not for restoring N accounts in bulk, which would be significantly more efficient than doing the one-account process N times. (E.g., in the case of tape backups, imagine the difference between pulling data from the tape library sequentially for each user vs all N at once, particularly when one tape may hold data for many of these customers.)
Plus coming up with an answer to the vague question on "describe your project methodology" (I build what you want, it works - nope, they expect half a page). Or the 3 questions on project management systems and communication software choices that to my reading should have the same answer.
That said, I've never had an issue copy pasting the answers to similar questions. As long as you answer the question with the answer!
Regarding bulk restore, a big customer doesn't care if you can restore all of your customers' data, they care if you can restore _their_ data, and fast, hence the question of "can you restore data for a single customer?".
This outage should convince them to care. The problem isn't that Atlassian can't restore a single customer—there are people reporting that they've been restored. [1] It's that Atlassian can't restore 400 customers efficiently. So unless the RfP also has a question "will I be first on the list?" and the answer is yes, single customer restore is the wrong scenario.
[1] https://news.ycombinator.com/item?id=31023163 says "I was down, my instance is fully restored right now. ... Not sure what to tell you other than they are fixing life saving companies first, then the rest. That is what they have told us."
it’s not like people will stop using jira and confluence, lol
they basically have a monopoly there
Same with SAP. Once your ERP has same feature parity with SAP, it will become SAP.
JIRA is SAP of engineering... yeah why I haven't thought of that before.
If Jira was a product used by individuals I'd get it. Maybe a database is overkill for a sole developer. But pretty much all users of Jira are companies with tens or hundreds of users on average. I don't see how separating on a db level is overkill in that situation.
Really annoying things that slash your velocity. You can't easily run pan-customer queries, can't aggregate data.
Ironically too using a single database makes full backup and restore much easier.
I understand the sentiment, but This is a pretty simplistic take that I very much doubt will hold true for meaningful traffic. Many databases have licensing considerations that arent amenable. Beyond that you get in to density and resource problems as simple as IO, processes, threads etc. But most of all theres the time and effort burden in supporting migrations, schema updates, etc.
Yes layered logical separation is a really good idea. Its also really expensive once you start dealing with organic growth and a meaningful number of discrete customers.
Disclaimer: Principal at AWS who was helped build and run services with both multi tenant and single tenant architectures.
And for migrations and schema updates I'd see this as a huge advantage. Migrating customers one by one is much easier than everyone at once. You also never have the issue that operations at one customer could cause a global lock affecting other customers.
Of course resource sharing isn't easy in this scenario, but you'd never want to connect data between customers anyway so I don't see the issue with that.
But maybe it works harder in a cloud environment where more is abstracted away.
Very fair call out on having more granular, discrete, instances for things like DML/schema updates and expensive queries. I love fault isolation and have had many sad days oncall when we exceeded the capabilities of The Database.
I wouldnt say it's harder because it's more abstract. I think the general motivation is to desperately avoid anything that scales cost/effort with the number of users. Even if it's sublinear a team can really drown under the cost of scaling up a service. And that's a serious consideration when a baseline expectation is to go from 0 to 10,000 or 50,000 active customers in just a few years. The care and feeding of (for example) 10 multi tenant partitions is just simpler than having to monitor & operate 10,000 independent databases with wildly divergent usage profiles. I will grant this hyper growth is not a common scenario for the industry, or if it is then its "one of them good problems."
I'd also say I have worked on a project that did have independent data tables for each customer instance. And we spent a meaningful amount of time abstracting away table creation/migration/etc, a common DAL that abstracted away the multitude of tables, common monitoring, etc. It has made some things around data migration & management easier but I honestly don't know if it's more efficient than multi tenant clusters in the long term. But the only way the economics and operational effort has worked is by going "all in" on using "serverless" technologies that efficiently scale to zero and have no carrying cost when idle
* Managing schema migrations across every DB
* You cant query across the DB, want to know some cross tenant thing for ops? That's now a lot harder
* Connection pooling and resource usage can be harder to manage
Most systems I've worked on use a single DB with a `tenant_id` col on every relevant table, it's easy to have your query builder slap in the auth'd tenant I'd. This approach does come with issues like saving and restoring an individual tenants data
Like a lot of things in life, it's a trade off
They did last night: https://www.atlassian.com/engineering/april-2022-outage-upda...
Ouch. I hope no one person got the blame. This is a systemic failure. Regardless, my regards to the engineers involved.
I think what we’ve lost in the post-XP world is that just because you build something incrementally doesn’t mean it’s designed incrementally (read: myopically).
My idiot coworkers are “fixing” redundancy issues by adding caching, which recreates the same problem they’re (un?)knowingly trying to avoid, which is having to iterate over things twice to accomplish anything. They’ve just moved the conditional branches to the cache and added more.
Most of the time, and especially on a concurrent system, you are better off building a plan of action first and then executing it second. You can dedupe while assembling the plan (dynamic programming) and you don’t have to worry about weird eviction issues dropping you into a logic problem like an infinite loop.
More importantly, you can build the plan and then explain the plan. You can explain the plan without running it. You can abort the plan in the middle when you realize you’ve clicked the wrong button. And you can clean up on abort because the plan is not twelve levels deep in a recursive call, where trying to clean up will have bugs you don’t see in a Dev sandbox.
Deleting 500 users…
Versus Permanently deleting 500 users…
Maybe with a nice 10 second pause (what’s an extra ten seconds for a task that takes five minutes?)Here's the way such a script should be done. You have a dry-run flag. Or, better yet, make the script dry-run only. What this script does is it checks the database, gathers actions, and then sends those actions to stdout. You dump this to a file. These commands are executable. They can be SQL, or additional shell scripts (e.g. "delete-recoverable <customer-id>" vs. "delete-permanent <customer-id>").
The idea is you now have something to verify. You can scan it for errors. You can even put it up on Github for review by stakeholders. You double/triple check the output and then you execute it.
Tooling that enhances visibility by breaking down changes into verifiable commands is incredibly powerful. Making these tools idempotent is also an art form, and important.
Because the names were hard coded, I had to get changes approved in GitHub. Then the script would run on Jenkins.
That script was also only for that purpose and nothing else. It made a mess because I needed a ton of functionality around creation and querying, too. I just copied the script to folders and modified them as needed but a better solution would’ve been to make a python module. I just liked the code itself being highly specific to what the script was doing to help reduce mistakes. If I’m running a script to delete repos, I need to go to the delete-repos directory.
So what they are saying is that they are not testing scripts at some staging server before running them in production. It's wild that they've managed to scale their products so much before something like this happened.
I hope they've learnt their lesson and they set up some QA process for that stuff.
I'm not looking to help their specific problems, but this is more from a general question I've thought of doing but never have done just because I'm sure I'd get laughed at for blazing my own trail
WHENEVER a human is involved in the chain, UUIDs can be suspicious because there's no easy way to verify what it is, whereas a human has a good chance of realizing that $1,342.34 is probably not a valid date.
This might be customer id and customer name, article number and article description, invoice id and invoice number etc.
Then it is usually very clear to the recipient what they've been handed.
Also, for internal autoinc-type id's, we mostly use sequence generators with non-overlapping "series". That is, we'll start first one at 1 million, second at 2 million or similar. Not perfect but can be useful.
If there is no feasible way of replicating their production environment somewhere else, then there should be some sanity checks in place. Something like "if an abnormally high amount of customer sites go down during the script's execution, kill the script". This is a 20/20 hindsight approach though and if Atlassian engineers can't solve I doubt a random HN user like me can.
Perhaps my ad blocker is causing that stupid highlight to tweet js they are using to break.
They did, this post by Atlassian from yesterday is referenced in the article.
https://www.atlassian.com/engineering/april-2022-outage-upda...
Still doesn't excuse them for the time taken to come clean.
1. It's snappy. Moving between issues and views is a breeze.
2. The UI is extremely functional and consistent.
3. Great keyboard shortcuts!
Using Jira felt like having brain fog, and Linear is such a relief after that experience.
When we chose Jira, one of the points that was made was: If we decide to leave Jira, there will almost certainly be an importer from Jira to the new system. Which does seem to be true. We came to Jira from Fogbugz, and I spent the better part of a month writing tools to import our tickets and wikis. Jira had a Fogbugz importer, but it was horribly broken.
Looking at Linear, there is no such escape hatch, or indeed, searching the docs I saw no "export" or "backup" capability at all.
You're right we need add a page in the docs. It was only mentioned in few pages in passing. For now I added a section here until we can make a full page about it: https://linear.app/docs/workspaces#export-workspace-data
If you are looking anything else hit cmd+k to search the docs.
Does everything I used to use Jira for, but feels more modern and lightweight. Also, it has dark mode.
Linear has none of these issues. I've been super impressed with it.
How peculiar that the biggest active outage in the history of this company is not mentioned once in this "article".
I'm left to assume that PR teams can plant whatever they see fit in the WSJ at a moment's notice. I guess that's what passes for journalism these days.
https://www.wsj.com/articles/atlassian-puts-easy-to-use-codi...
> As of Wednesday, she said, services were back online for just under half of the companies hit by the outage, which may take up to two weeks to fully repair.
Source: https://www.wsj.com/articles/atlassian-puts-easy-to-use-codi...
Atlassian's SLA page says, Premium Cloud Products 99.9%
That's 43 minutes of downtime per month.
That works out to, Atlassian can't have any more downtime for the next 14 years. Are SLAs even real?
I'm being slightly facetious. From the page text it's just a threshold after which I think you're entitled to some money back for that month.
Except...I don't even believe that.
They are typically limited to the amount that you actually paid, though, so basically they don't charge you for the time when you couldn't use the product. You usually won't get more than that.
This outage alone has spurred conversations in slack about how terrible JIRA is and why we should replace it. If this kind of shit was pulled, I can guarantee we'd be on shortcut, linear, or something else in short order.
Atlassian absolutely can in enterprise settings. In my company (a large cloud company), if JIRA goes down, large swathes of the business will also stall, including code deployment (deployments are tracked through change management JIRA tickets). We also use the DC version of Atlassian products, so presumably we aren't be at the mercy of Atlassian cloud engineers.
I've been on-call during a total infrastructure outage whose root cause was a service my team owned [1]. Our CEO was aware of it. Customers and business partners were aware of it. Other CEOs were aware of it. The media, you name it.
Some outages can be "business ending" or "business damaging". That's why we made a practice and process of performing regular disaster recovery exercises, had exceptionally well documented runbooks, had monitoring attached to everything, and engineered for resilience.
Though I'm not familiar with how Atlassian runs, I think this is an "engineering culture" thing or can be mitigated with a proper approach.
[1] The company has only had a few of these in total, and no member of our team was culpable for the complicated failure.
SLI: Some metric you use to measure a thing (e.g. uptime, latency, etc.)
SLO: Some objective you try to hit, as measured by the SLI (e.g. "99.99% of requests are processed within 3 seconds)
SLA: A promise to a customer that they will meet some SLO, and consequences if they don't. If there aren't consequences for not meeting the SLO, then measuring and tracking the metrics is a pointless exercise.
The SLA is "real" to the extent Atlassian is adhering to any listed consequences.
SLAs are mostly aspirational.
If the maintenance costs exceed the margins on the cars you lose money. Do that on too many product lines for too often and you’re looking at bankruptcy. But some makers clearly are more risk averse than others, so a 6 year warranty from maker X does not translate to a 7 year warranty from maker Y.
* - their larger customers will have negotiated SLAs.
edit: to be clear, I expect Atlassian will offer concessions beyond their SLA obligations. I'm only responding to the comparison.
Customers that demand service level agreements often fail to recognise that they cut both ways.
And these consequences usually just amount to getting some percentage of your service fees back. I'm sure the affected customers will get their entire monthly Atlassian Cloud fees back. Since this is so severe maybe Atlassian will even give them credits for some # of free months.
But there's no way the amount they'll get from Atlassian is going to come close to what they're losing in productivity by not having access to Jira & Confluence. At my company, getting an entire free year of Jira wouldn't be worth Jira being inaccessible for a week.
> That's 43 minutes of downtime per month.
we need a better default way to communicate SLOs than "number of 9s", which are more human. how the status quo has stayed this way can only be attributed to intentional dark patterns, imho.
And a couple of percent discount on services for the extra downtime isn't really a meaningful consequence.
Offering you a free month or whatever doesn't acknowledge all the person-hours lost.
Tommy: Here's the way I see it, Ted. Guy puts a fancy guarantee on a box 'cause he wants you to fell all warm and toasty inside.
Ted Nelson: Yeah, makes a man feel good.
Ted Nelson: But why do they put a guarantee on the box?
Tommy: Because they know all they sold ya was a guaranteed piece of shit. That's all it is, isn't it? Hey, if you want me to take a dump in a box and mark it guaranteed, I will.
That is exactly what SLAs are.
There are just a lot of people applying the wishful thinking that SLAs are a goal or metric of uptime.
Consider the AWS S3 page on the topic: https://aws.amazon.com/s3/sla/
"Reasonable efforts"; if not met, you get some fraction of the money back.
S3 has worse uptime than my desktop PC over the last years, but affected users got some fraction of their spending back.
That's sacrilege on HN
One alternative I thought of is the Charity SLA. The service provider pledges to give $5,000 to charity for every minute of downtime. Now everyone within the company knows "if we're down, we're losing thousands of dollars a minute!" and thus will be motivated to ensure the services stay up. But even if the services go down, the company's making tax-free donations, which isn't really bad for anybody. The company could even have a specific downtime goal every year, to make sure their monitoring/alerting/runbooks actually work, and to ensure they donate every year.
Suits think a crummy Flash quiz on PII is enough to stop leaks. The automotive industry couldn't stop airbags from acting as claymores. It's even harder to get good code approved in tech.
But outsourcing one's whole engineering environment to a SaaS on a cloud is just freakin lunacy. Not only do you have things like this outage, but what about simple things like features and versions of the apps changing all the time with no ability to control that. What if they remove or change a feature you use?
And expensive vendor-locked-in closed tools have no place in a modern software workflow anyway, on-prem let alone SaaS. Look at the rug-pull for the on-prem Atlasian Server product.
Just to put this in perspective. These executives would have left on a Friday afternoon to start their weekends without bothering to publicly address an ongoing outage that was by then 5 days old.
This is mind boggling. Like did some C-level exec say something like "Let's just park this whole outage communication discussion until Monday, have a good weekend everyone."?
I still don't understand the strangehold JIRA has on some clients. I can't quickly think of another SaaS product that could be down for almost 2 weeks and not have most customers leave.
Secondly, there are a lot of individual competitors to Jira, Confluence and Bitbucket but which competitor can offer all three under a single invoice? May be Microsoft, can't think of anyone else.
Also for such an extended downtime the customers are entitled to a discount or a credit note which a lot of CXOs consider in their decision making.
Is there a Jira replacement/offering in the Microsoft 365 suite?
But it's got what plant's crave...
For companies of size, the cost of tools being down for 3 weeks can easily be in the multi-millions of dollars.
- Integrations with things like the source code repos, incident management systems, confluence or other wikis, Slack, etc. Moving away from Jira creates a bunch of dead links.
- Internal dependence on complex workflows and state transition rules that are implemented in Jira.
- Various very customized reports that leaders depend on to make decisions, despite the often dubious value and/or accuracy.
// if we don't toggle bit 7 here 10% of transactions will fail on Thursdays
// see JIRA issue BIGPROJ-12654 for detailed discussionThis is actually the least difficult thing, i would say ;)
But sure, we can break all of our users to avoid the possibility of you having to write some legal briefs and us paying a small fine for keeping data 7 days instead of three.
I'm still kinda salty about it. I understand why big services can't retain data indefinitely, but like... it's just a few KB of text, and that text happens to have a lot of sentimental value. Besides, OkCupid knows that I deactivated my account because I am a success story -- why not hold onto those profiles a bit longer? Or better yet, how about emailing an archive of those messages immediately when you click the "I'm leaving because I'm in a happy relationship now" button? /rant
If they had obtained the data without you supplying it freely then that would have been an entirely different matter, especially if it was used in ways that you did not consent to. But since that does not appear to be the case here the GDPR applies like it does to all data that is directly related to a data subject but continuing to store it on behalf of the user(s) that supplied it is not a problem.
Note that the user here is disappointed that their data which they consented to be kept is no longer there. This is a pretty clear indication that as far as they are concerned their expectation was the even with the GDPR up and running that such data would continue to be preserved as it is in almost every service that existed prior to may 2018.
It is precisely this kind of panicky thinking around the whole subject of the GDPR that gives these irrational responses, companies that suddenly no longer dare to mail you but you have to log in to their portal, which is secured by your email address and more of these totally weird constructs.
If they wanted to delete this data the better way would have been to positively contact the user (so that you know that they have received your message) to ask if their data should be deleted or not. That's good stewardship, just tossing it isn't.
Compliance? The contract has expired, so there’s no legal basis for them to keep your data?
> Second, the script we used provided both the "mark for deletion" capability ... (where recoverability is desirable), and the "permanently delete" capability that is required to permanently remove data when required for compliance reasons. The script was executed with the wrong execution mode and the wrong list of IDs. The result was that sites for approximately 400 customers were improperly deleted.
> To recover from this incident, our global engineering team has implemented a > methodical process for restoring our impacted customers.
[https://www.atlassian.com/engineering/april-2022-outage-upda...]
Anyone else find it disturbing that they are able to restore data that they deleted permanently for "compliance" reasons? If this is true, how were they ever compliant? I guess data is only permanently deleted when the engineering team is following their typical, non-methodical process...
Going back and purging things from the backups as part of the delete process would be overdoing it to a ridiculous degree.
That can mean as much as "you have to encrypt everything with a separate key, so that you can destroy the key for the given (say, personally identifiable) dataset making its retrieval irrecoverable"
I'm not saying that's the particular compliance reason they had here, or that the analysis you're giving is wrong, either. There is an interpretation where either of these ideas could be the correct one.
If this weren’t a concern, regulations would demand immediate deletion of data.
there's a whole lot of people in here who are way too quick to assume that just because one part of a permanent deletion process was inadvertently triggered and then caught while they still had backups, their whole permanent deletion process is a lie.
You seem to be right-ish, while the gdpr in certain circumstances allows you to keep backups of data that should have been deleted it seems like they are trying to discourage it in the future.
> ...It is, however, important to note that where data put beyond use is still held it might need to be provided in response to a court order. Therefore data controllers should work towards technical solutions to prevent deletion problems recurring in the future.
It's your job to use technics which allow you to do this like using encryption on your backup and deleting the keys for it, for example.
Exactly. I'm not aware of any laws saying "you must delete this data immediately". More like "within X days or months". The permanently delete thing presumably skips some cooling-off period in the online database but not the backup, which seems perfectly appropriate, provided your backup retention is compliant.
Google has a nice page describing out their deletion process. [1] It doesn't go into product-specific technical details/steps (like marked as deleted within the product, row deleted from Bigtable/Spanner, major compaction guaranteed to happen, backups guaranteed to be deleted or unusable) but it says this:
Google> We then begin a process designed to safely and completely delete the data from our storage systems. Safe deletion is important to protect our users and customers from accidental data loss. Complete deletion of data from our servers is equally important for users’ peace of mind. This process generally takes around 2 months from the time of deletion. This often includes up to a month-long recovery period in case the data was removed unintentionally.
This is a best practice.
Delitio> It's your job to use technics which allow you to do this like using encryption on your backup and deleting the keys for it, for example.
If they'd thrown away the encryption key immediately, this would have been much worse. Instead of "we're down for 2 weeks?!?" (already quite bad) it'd be "our data is gone forever?!?". You never want to delete anything too quickly for exactly this reason.
[1] https://policies.google.com/technologies/retention?hl=en-US
Also, modifying backups is a great way to inadvertently hose your backups.
> If a business stores any personal information on archived or backup systems, it may delay compliance with the consumer's request to delete, with respect to data stored on the archived or backup system, until the archived or backup system relating to that data is restored to an active system or next accessed or used for a sale, disclosure, or commercial purpose.
If you make backups, you are, almost by definition, unable to perform a full 'Compliance Delete' before the oldest backup in the set has expired.
Compliance-based deletion, if it is offered as a service, is almost always something time-based, like "we guarantee the data will be deleted 7 years from now". And then that deliberate deletion step is baked into the backup process.
So, i.m.o. at best they misrepresented the nature of the compliance deletion process. It never did what it was designed to do.
For example, if you delete an email or document on Google it moves to the "Trash" folder for 30 days.
When you manually empty the trash or the time window expires, most likely the next step would be a soft deletion for a few days where the data is still on hard drives but hidden from the application. Soft deletion is mainly protection against coding errors, since soft deletion is easy to undo if you've caused an incident but hard deletion (removing the data from disk) is not.
Then most likely a garbage collection process comes by a few days later and hard deletes the data from disk, leaving it only on tape backups
Finally, maybe a month or two later it disappears from the tape backups as they get rotated or otherwise disposed of
This addresses the needs of:
- Giving a good user experience (user "oops I made a mistake" undelete)
- Protecting against incidents due to coding errors (software engineer "oops I made a mistake" undelete)
- Making sure data disappears from both disk and backups within a certain time window, like maybe 30 or 60 days (comply with regulation and user expectations of data being cleared)
An overarching theme with these things is “legitimate business need” and “no indefinitely retained customer data. Having backups, system event logs, etc are all legitimate business needs. Based on the data type that business need may be days or years with things like financial and legal requirements.
Youre conflating permanent, immediate, and irrevocable. These are usually handled in different aspects. Think of accounts having multiple states like active, suspended, closed, terminated, purged. Some examples;
suspended: credentials/authnz immediately disabled, all data online, charges continue to accrue, can be restored in minutes.
Closed: credentials disabled, data online, processing stopped, charges stopped, may take manual intervention (hours) to return to active.
Terminated: creds & account irrevocably unavailable, online data deleted, offline data (backups) remains available.
Purged: all online and offline customer data irrevocably unavailable. This generally happens after a defined retention period for things like logs, backups, etc.
You can apply similar concepts to individual resources more granularly than the account.
Disclaimer: principal at AWS but the above is my own opinion/observation and does not represent my employer.
On-premise isn’t a magic pill guaranteeing 100% uptime and 0 data loss.
While on-premise may be a good choice in many cases, it’s not like running on-premise business tools has no risk associated with that choice.
Remember that the goal of a company is to sell the most product possible (output) with the lowest cost possible (input).
Any Joe off the street starting their own business can pay Atlassian $0/month for up to a 10 users. On-prem doesn’t compete with that.
[1] https://www.atlassian.com/enterprise/data-center
[2] https://confluence.atlassian.com/enterprise/jira-data-center...
[3] https://confluence.atlassian.com/enterprise/deploying-enterp...
(You can argue how successful it was when people are still using the old style in 2022).
It also makes more sense, since Jira is not an acronym, it’s a truncation of Gojira, inspired by Bugzilla/Mozilla.
1. Different root cause. There was a bug in a refactoring of gmail's storage layer (iirc a missing asterisk caused a pointer to an important bool to be set to null, rather than setting the bool to false), which slipped through code review, automated testing, and early test servers dedicated to the team, so it got rolled out to some fraction of real users. Online data was lost/corrupted for 0.02% of users (a huge amount of email).
2. There were tape backups, but the tooling wasn't ready for a restore at scale. It was all hands on deck to get those accounts back to an acceptable state, and it took four days to get back to basically normal (iirc no lost mail, although some got bounced).
3. During the outage, some users could log in and see something frightening: an empty/incomplete mailbox, and no banner or anything telling them "we're fixing it".
4. Google communicated more openly, sooner, [2] which I think helped with customer trust. Wow, Atlassian really didn't say anything publicly for nine days?!?
Aside from the obvious "have backups and try hard to not need them", a big lesson is that you have to be prepared to do a mass restore, and you have to have good communication: not only traditional support and PR communication but also within the UI itself.
[1] https://static.googleusercontent.com/media/www.google.com/en...
[2] https://gmail.googleblog.com/2011/02/gmail-back-soon-for-eve...
> Atlassian claims the customers impacted were “only” 0.18% of its customer base at 400 companies.
From https://jira-software.status.atlassian.com/ :
> The team is continuing the restoration process for the ~400 impacted customers.
If this is truthful it implies implies more than 400 impacted customers.
prevents data integrity issues in relational databases, makes debugging easier and prevents disasters.
ideally also include a timestamp, both for bookkeeping and safe tools that only remove things that have been soft deleted for some time and are safe to delete without compromising integrity of anything that is not deleted (this is especially important in relational data models)
Perhaps an effective measure would be to create a key that encrypts a customer's data, and give them a copy of the key, and let them know that after a certain point your copy of the key will be deleted, and if they want a restore past that point they'll need to provide the key.
not sure if the right way forward is some sort of innovation in operating system and software design where people write and run apps that feel like single tenant apps attached to dedicated per tenant datastores where os and framework magic handle per tenant encryption and segmentation (tenant id as an os level concept)
or... if it makes more sense to encrypt at the record level with keys that only the customers hold using (assuming it's up to the task) homomorphic encryption for things like searches and other backend functions.
either way, for now, soft deleting and following up with an automatic daily hard delete of things soft deleted more than x days ago is a totally reasonable approach.
ops scripts should require typing "yes i know what i'm doing" if someone attempts to hard delete things that have not yet been soft deleted.
if it forces additional fail safes or backups to be able to do so safely, then that's probably a good thing to have anyway, no?
They may be scared. But are they scared enough to reload every single backup they have, purge the desired records, and resave each and every single backup they have? And not also worry they will corrupt/break the backups in the process.
GDPR compliance is a mess of contradictions and unreasonable asks which all seem to amount to "depends on who you ask."
that's okay though, queries that reference the timestamp can be slow since they're housekeeping.
But a nullable timestamp could also be viable, just can't tell two different delete, at the same time, apart
We use JIRA. Not impacted.
If this had hit us.. we would just switch to excel or something for a week/month?
But maybe we are a very light user of JIRA. Nothing in there can't be replaced. It's "nice" to be able to go look up a 3 year old bug and which client reported it, but not really crucial for day to day ops.
A spreadsheet may be sufficient, but it's not as good as a system designed for development workflows.
(This comment sounds like I have a speck of love for JIRA. I don't! :)
Today it would even be more troublesome as we have a lot of integration rules dependent upon the workflow. I'd probably just recommend everyone uses a few weeks for self improvement and only address critical production issues.
Right.
We used excel for the first year. Trello for the second.
I know what we need.
This time.
At some point the logo on the engineer’s badge doesn’t really matter.
If your Oracle DB or Cisco router has a software big, you can always restore/rebuild but that doesn't guarantee you won't hit it again and in both cases you're still at the mercy of the company producing it.
Even if you're on OSS are you able to fix a data corruption bug yourself?
You get more control over maintenance windows and backups, but it doesn't automatically guarantee better uptime.
This even is not even showing on any financial news site. I'm still hoping it does and the stock goes down because I place an option order yesterday betting that it goes down by next Friday. Seems like it won't now but the risk was worth taking in my book.
Oh well, better luck next time.
Title should be "Inside the longest outage of all time", without "Atlassian" word in it
Slack is generally much more critical than JIRA in order to keep working.
You can do without JIRA for a week or two as long as managers understand and you all have a good concept of what work needed doing anyway. Then it starts getting dicey unless someone becomes a human JIRA to connect temporary manual bug tracking systems with everyone involved.
Other IM platforms wouldn't solve that just by existing. Sure, in principle one could set up such channels elsewhere, but that takes time, and the communication about it takes considerably more time.
Then if it does go down, you don't have to waste the first day arguing about the plan.
“We build to static HTML, deploy to S3 and Cloudfront, and it’s f*ing bomb proof”
I personally stopped using Jira a couple of years ago in projects I lead.
I am curious if anyone can provide any more insight on this simplification.
I've worked at companies like this. Originally a core of motivated creative individuals make a cool product. As the business grows rapidly, Pournelle's (Iron) Law (of Bureaucracy) takes over. For a variety of reasons, the very capable creators depart and are replaced by less motivated/aware individuals who are glad to have a job and easily compelled to do things to the product that probably should not be done.
My guess is that while Atlassian may have originally been one of those cool founder places, it has probably morphed into the more incompetent version that comes with scale all too often. But I don't know. Thus my question if anyone can speak to the true current tech capabilities of this company.
Note that I'm pretty naïve and armchair on this subject, I'll see myself out.
To be complete open though I don't know how much DevOps overhead is involved in maintenance or feature updates. I hated the app and used it for less than a year so I didn't have much exposure. I guess my point though is simply that you may not need to use their SaaS option if you have a decent DevOps team already. After the initial setup time I doubt I spent more than half an hour a month managing the internals and updates.
I did spend more than that on configuring the system for use, which you'll need to do regardless.
Which is even more amusing when you realize Server has been Datacenter with a fake mustache for years now.
- 500 users (Jira Software, Confluence, Crowd)
- 50 agents (Jira Service Management)
- 25 users (Bitbucket)
https://www.atlassian.com/migration/assess/journey-to-cloud
https://www.atlassian.com/migration/assess/compare-cloud-dat...
But I also get the impression that you may just be expressing a preference, not a rule of DevOps? If so then I definitely understand. Custom solutions to integrate or glue disparate systems together is often not the most interesting work. My area... A single word doesn't encompass what I do, I'm a generalist in my domain with one or two specialties, but glueing data together (not the same as a full integration, I know) is a big part of my job, and usually the least interesting.
Though in this case from other comments on prem seems a a dwindling option anyway for Jira. I worked with it about 7 years ago under one of their free licensing programs and disliked it enough that I didn't bother following them after that.
Maybe time to throw a few chips at some long term puts?
Disclaimer I have puts that expire 4/22 (purchased yesterday) so I hope they go down in the short term. Seems like a total loss now after being up 50% yesterday.
That might look like incompetence, but I think it’s confidence. They know the switching costs for large orgs are so high they can treat these people like trash and few if any will leave. I wouldn’t be surprised if the total number of seats among affected customers has gone up in a few months. By failing to acknowledge the problem they’ve kept it out of the mainstream media and financial press.
They have their customers by the balls and don’t respect them. That’s a short term bullish signal to me.
Mine expire 4/22 but I have more calls open at the moment anyways so if I had to choose between this going down or the market up I'll take a full loss on these puts (seems likely at the moment)
The real key lesson here. Your business is important to you. Not so much to the service provider.
My guess is many people new about the problem inside, but corporate taboos made it impossible to discuss. I'd bet a fortune on this being the case.
There are some Python scripts that will back up Jira and Confluence. I whipped up a quick script that gets a list of all our bitbucket repos and then it clones those daily as well.
>Open company, no bullshit >Don’t #@!%the customer
Naive approach, replace delete with select and see if you're surprised at the results.
More mature approach, especially in an environment where engineers are running bulk changes against the database, you don't do bulk deletes. You change that delete into an update that marks things for later collection.
One tactic I've seen that worked, assuming you have straightforward relational tables: you add a "marked for deletion" column whose value is an identifier for the single run of the bulk job you just did. Then you can query rows with that value in that column to ensure it had the desired effect. If you're satisfied, you run another bulk job which doesn't re-run your original query.. it just deletes rows with that "marked" value.
Lots of places rely on schema-enforced foreign keys and cascading deletes though. In that case, my recommendation is: don't.
It's not clear I'd the issue affected all tenants where the script ran--which it sounds like it did. It wouldn't be as effective if it only effected certain tenants (maybe with a specific config)
OK, so you restore backups to a separate system, and selectively copy the stomped accounts data back to production. Simple concepts aren't that simple at their scale, sure, but I suspect this is skimping details on some truly horrendous monolithic architecture choices that they're trying to hide.
Not that I ever thought using their products was a good idea; to be clear about my position... But at this point anyone continuing to rely on them for anything is asking for the suffering they'll get. Signing up for their crap for a vital business function is like offering your tonker to a snapping turtle.
Is it not a good idea to spin up separate db instances for each client/company?
It depends, really. There is a trade-off in terms of software and operational complexity vs scalability/perf and isolation. And probably a bunch of other factors.
If you have separate databases for each customer, schema migrations can be staged over time. But that means your software backend needs to be able to work with different schemas concurrently. You can also benefit from resilience and isolation guarantees provided by the dbms. On the other hand, having a dbms manage lots of databases can affect perf. Linking between databases can be a minefield, especially w/r/t foreign keys and distributed transactions.
https://docs.microsoft.com/en-us/azure/azure-sql/database/sa...
Not if you're truly multi-tenant and each customer has their own app servers. Then your code and schema version are always in lock-step.
1. An account/tenant_id field for each table
2. A schema for each tenant wrapping all of the tables
Option 2 gives you cleaner separation but complicates your deployment process because now you have to run every database change across every schema every time you deploy. This gets more complicated as your code is deploying in case the code itself gets out of sync, there's a rollback or an error mid deploy due to an issue with some specific data.
The benefit of the approach is the option to do different backup policies for different customers, makes moving specific customers to specific instances easier and you avoid the extra index on tenant_id in every table.
Option 1 is significantly easier to shard out horizontally and simplifies the database change process, but you lose space on the extra indexes. Plus in many databases you can partition on the tenant_id.
Most people typically end up with option 1 after dealing with or reading horror stories about the operational complexity of option 2.
Business wants to run a query across customers? In most DBs you need either custom code or to create a stored procedure to iterate across schemas.
Every table that you create is multiplied by the number of customers. This has implications for some database systems (like PG's vacuum).
Your migrations will take _forever_ to run.
Etc.
I've also had to restore partial data from backups on a few occasions when customers fat-fingered some data and asked pretty-please to undo. If someone on staff understands the system well, it's not hard. I suspect Atlassian suffers from a complicated schema and a post-IPO brain drain.
Just wait until a migration doesn't run on 2 of your 400+ customer databases. Or multi-hour migrations.
At least it would not be the first time in history that a company has lost the engineering spirit. And instead the business people have taken over, so that details like disaster plans become less of a priority.
A business person and an engineer will always view risk differently, better disaster plans is a kind of insurance that is a lot harder to sell when too many business people run the company.
Would still be some maintenance, don't get me wrong. But far from impossible.
source: every single place I've worked at that poo-poos referential integrity has a database that is full of bullshit that "the application code" never cleaned up
Always use referential integrity. The people who are against it almost always are against it for superstitious reasons (eg: "it makes things slow" or "only one codebase calls it so the code can enforce the integrity"). All it takes is exactly one bug in the application code to corrupt the whole damn thing. And that bug will happen over the lifetime of the product regardless of how "good" or "awesome" the programmers think they are....
... I'll get off my soapbox now!
You're lecturing about table design. I'm talking about more general transactionality over any errors.
Don't do any of the above unless you understand the implications.
Oh, and just forget about allowing your customers to share their data with each other, which most enterprises want in one way or another.
Giving every customer their own table means you're going to need database administrators. For these folks their dedicated job was maintaining, operating, and changing their fleet of databases, but they where very technical and were amazing to work with.
Does this extend to services as well? We have a suite of (micro) services. Are they all segregated?
I (and many others) assumed they had to graft in data from backups since a full restore would clobber newer changes from unaffected customers.
If they're all isolated in their own logical per-tenant DBs, I'm really at a loss for what is making restoration take 3 weeks for 400 tenants.
I understand if you'd rather not venture into it, but care to offer any speculation?
For example, they have the main PostgreSQL data store. Surely that's easy to restore. But the users in that DB have a "foreign key" (in a logical sense, not physical) to the Identity service. This is a real life example that occurred while I was there. So now we have a mixture of multi and single tenancy. So perhaps the identity records are also tied to this app ID and deletes were propagated to that service. And perhaps there is an SQS queue and a serverless function to handle, say, outgoing mail from Jira. Where does this data go? I dunno maybe some Go-powered microservice with its own DocumentDB store. Do deletes propagate here too? Who knows. You can see how this gets complicated and how issues multiply with more services.
Again, this is only speculation. But "decomposing the monolith" was a big deal and it was coming from the top.
The best would be multiple separate database instances, which is not even hard to manage specially for qualified engineers like Atlassian surely has plenty of. The problem are business decisions of ignoring the tech debt, usually...
Automation wasn’t the issue here. It’s the symptom not the cause.
Really interested if you can share any details.
Edit: I know each wiki is on a subdomain. Does each wiki also have it's own server?
Some of this was open source before we unified all of our wiki products, which has a lot of the selection / db logic, at https://github.com/Wikia/app.
Source: worked at Atlassian, on Jira, 4 years ago.
I've eventually gotten a tenant exporter to work. Practically, this requires some deep and nasty digging through the information_schema to build a graph of tables and foreign key constraints. Once it had that, it generates selects with a simple where clause for tables with the tenant_id, and selects with weird joins all over the place for other tables to dump the tenant data.
All of that sounds complex, but that part took a day or two to hammer together to 90% completion, since it's just some graph handling. The other 10% were getting some weird date formatting questions right to produce a properly importable sql dump. And interestingly enough, it's working for more than just that one product.
But that's just where the journey started. After that, it took a weeks and months to sort out legacy tables, old tables, tables without indexes, tables no one knew about, tables that were important (but not), tables with inconsistent data, .... And it's just handling a single relational database. And compared to \copy in psql, it's slow. And at times, weird things happen if you import huge chunks of sql into a postgres with deferred foreign keys (because our schema has cyclical references).
Point is, I know how painful it can be to handle that kind of database schema, at a ridiculously smaller scale. I'm kind of happy to not work there.
I would really like to try working in an organization that uses something simpler, like Trello (although now that this is also an Atlassian property, maybe not exactly Trello?).
If JIRA didn't allow you to make it terrible, it wouldn't allow for some of the absurd things that people want it for and those companies might not buy it.
The saying is apocryphal and unlikely to be accurate, but the shape of the thing its describing applies to almost every piece of enterprise software whether installed on-prem or SaaS.
And as another comment points out, at Enterprise scale you can substitute "team" or "group" for customer. Every team might use a different 5%, and unless you standardize their processes, you have to buy the product that can accomodate all of their needs.
>The saying is apocryphal and unlikely to be accurate
Well its mathematically impossible to be accurate as soon as you have > 20 users.
But your statement doesn't make sense; there might be millions of features, and trillions of ways to combine them to make 5%.
It's probably in the semantics.
Text input and editing is clearly a part of functionality that's probably used by everyone (or at least most users), so it's not possible for "different 5%" to mean what you're alluding to, maybe the phrasing needs work.
In any given 5% there might be 1-4% of overlap with what others are using and the remainder of that is specific to the company.
If it's a uniform distribution of discrete features then each feature is equally "important" and worth equal resources and dev time. If 81/100 companies use the exact same 5% of features and the remaining 19 cover the remaining 95%, then all else equal you can probably drop 95% of your features and still do well.
Typically you do the most popular features first, but most Enterprise vendors end up working on a long tail of niche features that nevertheless are profitable.
There's a long conversation to be had about how this ends up being a trap where Enterprise software gets bloated and shitty and eventually gets disrupted by a small vendor that does "less," but in a powerful, transformative way that obsoletes the Enterprise "standard," which leads us back to discussing Atlassian :-)
They're a good example of this dynamic, because they have a "constellation" of products to sell. So if they build a niche feature that gets a new customer to buy Jira seats, having "landed" in the account, their salespeople can "expand" by selling OpsGenie and other related products very profitably.
However, if we assume there are, say, 100 features in Word (the real number is likely much higher), the number of combinations is orders of magnitude higher than 20.
75,287,520
+ 3,921,225
+ 161,700
+ 4,950
+ 100
------------
79,375,495To this day I still don't know what JIRA does so much better that other products don't which big corps are willing to waste months worth of manhours over. It's biggest selling point is integration with the remainder of the Atlassian stack, not exactly known for being great either.
The fact that it’s so feature packed and customizable is the point.
I think the complainers are not really investing the time in to change project settings to fit their needs.
My only complaint about the Atlassian suite is the performance of Jira and Confluence. The overall page load speed is too slow.
How do you eliminate all non-task ticket types in a Jira board and allow any ticket to be a child of any other ticket?
It’s hard to configure away complexity from a product if it’s designed to be complicated.
Re 2: Set the project's issue type scheme to one that only allows tasks and subtasks. That gets you one level of nesting. (And even though task and subtasks are different issue types, changing from one to the other is trivial since they have identical fields.) Allowing epics gets you another at the top level. That's a bit limited, but wouldn't arbitrary nesting be even more complex?
This here is the single most insane thing about Atlassian.
I have no idea why you would want this from a work management point of view, but you can just use issue linking to describe a parent <-> child relationship.
As the lone dev on the team I've been continually astounded by my leadership's willingness to commit more and more to tech debt laden paths. The notion that all software requires maintenance is anathema to them, and it's led us to be 'cornered' into decisions re: what software we can use / where we can invest our discretionary funding.
Moreover, we're constrained by the parent mega-enterprise's software purchase policies; JIRA's already approved (and run elsewhere in the enterprise), whereas off-the-shelf or SaaSy alternatives are significantly harder to get buy-in for. (No using corporate cards for SaaS, all purchases need to go through the quote/purchase-order process, etc).
Clients add their notes to the card, I check the boxes as I hit the notes, and I move the card further right as we enter different stages of the post production process. We then have a column of every completed project, which is incredibly easy to sift through if we need to revisit something. It’s literally left to right in the workflow, it visually is telling me where we are at all times.
It’s incredibly simple and elegant. For fast turnaround, relatively stripped down content (like podcasts) there is nothing like it.
The idea that the C++ committee are unthinking people pleasers it patently false.
C++ does have a lot of cruft, but mostly because it aims to: i) support new features ii) maintain pretty strong backward compatibility guarantees
In general the new features are actually pretty well liked, but in conjunction with (ii) it creates a big language. There's a reasonably decent subset that can be carved out, but it's also clear why newcomers without legacy baggage (e.g. rust) are making inroads.
EDIT: clarified wording a bit.
It’s like my taste in wine. I don’t want an overdeveloped sense of taste where only a $400 bottle will do. I’m fine with what we have because the work is what excites me and if people are documenting projects and managing workloads and committing code, we’re 90% of the way there.
Wine that costs 400$ is for fun.
You don't drink that professionally.
In the beginning, with me plus 2 engineers, I noticed it was slow but since I used it for 20 minutes a week, that didn't really matter. By the time I started using it for an hour a day, we had 10 engineers on 2 teams using it. I got to see a friend using linear, and I had some spare time that I was going to use to switch, but I couldn't get in the beta. By the time they let me in, the opportunity was over and I was too busy.
Other ticketing systems do not work nearly as well for this purpose because they are designed mainly as external brains or communication platforms for workers, and they assume a level of worker autonomy in moving tasks through their lifecycle. In Trello you cannot make it so that a PM has to sign off before a card is moved to the in-progress column, or that only in-progress cards can have code reviews associated with them. JIRA eats these kinds of requirements for breakfast.
EDIT: This is not to say you can't use JIRA in a workflow-neutral way, or that everyone uses it for this reason, but I would submit that it's JIRA's differentiated advantage.
> Free software has zero acquisition cost, but non-zero TCO, which can measure in millions USD
Often a primary driver is exactly the opposite -- for-profit companies are accustomed to paying money for a good or service, with a billing pattern and legal obligations. The company financial deciders do not want a setup that does not have a billing pattern and clear legal obligations. Meanwhile, Open Source Software went from niche to mission-critical in the 2000s via the Internet. For-profit companies (and their publicists) scrambled to explain it, and came up with that exact line repeated again today. I do not blame any person for saying it, it was in print in some reliable place. It does not capture the reality in 2022 IMO.
> The company financial deciders do not want a setup that does not have a billing pattern and clear legal obligations.
I haven’t ever met a CTO or CIO, who would make budget decisions like that, neither I do it this way myself. The reality in 2022 is the same as it was in 2012 or in 2002: when you choose a solution, you consider all long term costs. In 2022 TCO for the server software includes everything that I mentioned in my comment and more. There’s a lot of use cases for OSS in corporate environment, for sure, but not every OSS solution is cheap or even affordable. Running on-premise open source collaboration tool is certainly not cheap if you do it right.
The same exists for levels, enemies, quests and tons of other elements.
I would not be surprised if a lot of studios had similar workflows.
However, that's just not what most people go through in companies using JIRA. Worse, they have to toggle between pages multiple times, each taking at least a few decent seconds to reload. I'd like to give JIRA the benefit of the doubt here, but it sounds like the tool is just very easy to misconfigure and abuse.
And you generally do them both at a lower level than tickets, certainly commits, so you don't want to have too much automation between them as that starts adding constraints.
However, regarding ticketing systems, in team environments, it is very effective and helpful to have a system that manages the data about the work that has been completed, is being worked, and is planned to be worked on .
Part of that system might be defining restrictive workflows for some teams, not for control, but to ensure the agreed upon process is followed for quality or consistency.
One of the many problems Jira has is that if you don't have a Jira admin on your team, it's impossible to build an effective and efficient workflow for your team. Coupled with Jira making many things global by default (it takes a lot of care to make a change that only affects specific Jira projects) most configurations end up being a pile of garbage automatically inherited from changes an admin(that is not part of the team) made when intending to change something for another specific team.
That is what I mean here by "assembly line" and "control." Making sure that processes lead and individuals follow.
Citing consistency as a terminal value in the same breath as quality is also exactly what I mean by the middle-manager aversion to local differences.
I work with a variety of different environments, and depending on the environment I can either solve my problem in minutes and get it deployed in another few minutes or solve the problem in minutes and spend hours figuring out how to safely deploy it without breaking everything. JIRA is terrible if you do anything that it offers by default, but when used properly it can absolutely help with this.
Obviously not all work is this way. Sometimes you need to drive a migration that touches every team, and then the technologies of bureaucracy and process become important. But most work should be done in human-scale groups that can be more towards the self-organizing and trust-based end of the spectrum.
However some middle managers take offense to the idea that their different sub-teams have different operating models internally, and lean on technologies like JIRA to try to make them all the same. Middle managers at my company have tried this, not very effectively , so it hasn't hurt me too bad. But I've seen their vision and recoiled in horror.
Could you elaborate? What kind of fear? “You’re fired”? I wonder how effective it actually is because of the current job market and also because I (and others) react very poorly to this kind of tactics: “you want me to fear getting fired? Joke’s on you, please DO fire me, I dare you”
Counterpoint: software developers aren't necessarily well paid or highly regarded everywhere, since remote working for companies abroad hasn't quite gotten mainstream enough.
So it might just be effective against some people, or in cases where the hiring process itself has become increasingly unreasonable - the job being working on boring CRUD apps but the hiring process being multiple stages of Leetcode and complex interviews.
That's probably not applicable to everyone since plenty of folk can grokk Leetcode and find jobs without too much trouble, but i still recall "The Unseen 99%" article: https://www.hanselman.com/blog/dark-matter-developers-the-un...
It probably applies to the industries and companies where devs are treated as a cost center and since those companies aren't all out of business, plenty of people must be working in such environments, with sometimes sub-optimal conditions.
Isn't that just a more corporate way of phrasing "control"?
And that is not a useful way of thinking when you have real engineers writing software that people depend on.
JIRA then becomes a tool for enforcing arbitrary rules, e.g. control
> It sounds like you've been hurt by the some terrible management practices, I'm truly sorry that some managers think their job is to control their subordinates.
When we assume someone was hurt, and imply they hold an opinion only because they were hurt, we risk delegitimizing their position. The interpolated message we might be sending is "your experience is personal and not representative of the subject at hand, and so your thoughts are only applicable to your situation; so, after we express our sympathy, your thoughts can be dismissed." Or the message we might be sending can be patronizing: "you hold your opinion for emotional, rather than rational, reasons; I'm sorry that you are so unfortunate."
To be clear, though, I'm sure this wasn't your intent, and it makes me glad to see someone being compassionate (i.e. that you bothered to consider the experiences and feelings of the parent commenter).
A personal story: I was raised devoutly religious but left the church in my twenties. My family and friends assumed I left because I wanted to be free from guilt, had been hurt by a culture that belied the doctrine, and so on (and they said as much). My change of belief occurred after recovering from a few years of mental illness, and while it is true that I may not have left when I did were it not for the opportunity to reexamine my beliefs (while trying to piece back the fragments of my life into a sense of self), the reasons why I left were the result of a lot of research and thinking. It was mildly frustrating when people assumed my decision was made for emotional convenience, when in reality, the research was uncomfortable and contemplating an unfamiliar universe was scary.
I recognize the irony here – the issue I'm highlighting in this comment may be something that only I feel is an issue, born from a personal experience. But I think it's more common than that.
I think the point is that Jira is particularly granular in the way that it lets you do things with permissions, workflow rules, roles, metrics, etc. There's a fair number of places that use that granularity to create a weird digital sweatshop.
Meaning the complaint is more about really deep "micromanagement as a service" than what you might get with lighter tools.
And I wouldn’t assume you’re not one of them. The worst cases I’ve run into aren’t even the psychos that embrace micro management as part of their “management style”. It’s the ones that genuinely believe they aren’t engaging in the behavior. They’re not micro-ing, they’re “helping” their team because they are an awesome manager and their team is almost awesome, they just need to be monitored very carefully and given “suggestions” until they nail it. But they’ll never nail it. Because no one is as smart, experienced or does a task “just so”. They view themselves as a mentor to all. All decisions must be theirs to make. Jira becomes the perfect tool since the team effectively becomes little boxes that accept tickets or stories and return work both performed and delivered as specified.
For any managers reading this that don’t see a problem with this or see some of those behaviors in yourself please understand that you are sacrificing your team’s happiness and motivation at the altar of your own insecurities. No one can grow where they’re not trusted and no one can improve their skills when they’re never given latitude to make meaningful decisions. Your people will make mistakes. They will accomplish things in ways that are different from how you would do them. It might even be objectively worse. That’s ok. That’s how you grow into a strong team with confident members.
I worked on the line (Toyoda Iron Works) and used a real-life Kanban implemented by the plant engineers. It was used for quality control, to broadcast quality control and station output, and was checked regularly against their internal estimates and baselines and used also as a gauge for employee output.
Control is what it's designed to do. The very fact that Kanban is the tool of choice should support at least some of OP's points, objectively.
You can use Jira as a simple Scrum board, a Kanban board, or you can build enforced-process monstrosities. You can build customer-support / internal-helpdesk workflows, or even model internal work-item-oriented business processes, etc. Now, as you point out, just because you can doesn't mean you should, and many orgs fall into the trap of making issue workflows overly-restrictive. But most companies (I believe) choose Jira before they choose those hairy task workflows. Startups with zero process use Jira.
Also, you can integrate it all together to give good-enough dashboards/roadmaps, good-enough (for some, not me) docs integrations with Confluence, Git integration with Bitbucket etc. -- while there are big issues with these systems, I think it would be myopic to ignore the real benefits of working in one integrated stack where every design doc you write has dynamically-updated labels and auto-complete for each issue you type in.
For context, I use Jira for tasks and don't love it, found Confluence to be really annoying and so I don't use it, and prefer Gitlab to Bitbucket, but I think you have to recognize these unique selling points. If all Jira had to offer was the rule engine it would not be as widely used.
Each member can actually organize their sprint and create tasks.
Point assignment is not a big deal, it's just there so we avoid promising more than we can chew.
I've found Jira really pleasant to use for lightweight processes.
One JIRA is the project that's used for development of the core product, where there are no constraints— anyone can add a comment, create links, change assignee, add new tags, push the tickets through whatever state transitions they want, and so on. It works, though it is a little chaotic sometimes as subgroups of people have different preferences for how things should go (eg, for tickets requiring test team validation, should the ticket assignee remain as the person who did the original work so it's clear who has more to do if it fails validation, or should the assignee change to the test team person, so that it's clear that that's the next person who has it as an action item?)
The second JIRA is the IT team's internal support project, which is completely locked down— no one except them can close tickets or move them around, or even edit the contents, closed tickets can't be commented on any more, and so on. This is the one that gives me the vibes you are talking about. Every time I have to interact with it, I loathe it because every inch of it is transparently a funnel, railroading me along a path toward one of either DONE or WONTFIX. This is absolutely efficient, in the sense of meeting the goal of closing all the tickets, but I feel it introduces friction for the larger business goal of actually helping people resolve their problems. To the point where eventually most of the IT support activity moved away from the JIRA project to an informal Slack channel, which is way more accessible, but worse in basically every other way: it's harder to effectively search, impossible to properly link, bad for async, bad for dealing with more than one thing at once, etc.
That means the tool is often the wrong one for the job, but instead of picking something that's a better match out of the box folks stick with the easy choice (extend what they have).
JIRA does try to be all things to all people…and mostly succeeds. For instance, we use the same workflow and mostly the same nomenclature across our development and helpdesk teams. Some of our software projects use Kanban-style workflows, while others use sprints, but we can keep track of a project across multiple teams using the same tools. I’m sure other products also offer this, but we liked the integration and overall capability for the price.
There are definitely issues: some feature requests and bugs have languished in their backlog for years. But you can get started very quickly and we’ve had great feedback from users.
Maybe Asana or Monday would work for you.
Personally, I like JIRA. I think it adds a ton of transparency in our org, and while I've used Trello for personal and home projects, I don't see how it's good enough for business. Trello doesn't even allow for time estimates (last I tried), which for us is part of planning. Search in JIRA is also really good, so no ticket is ever just lost to the ether.
Sure, it's not perfect, and waiting for a board to load is annoying, but for distributed work and visibility, I haven't seen something as professionally useful.
Open to exploring though.
How good is ticket search? I have to be honest, JQL is the superpower that makes or breaks for me.
Then, you pay for JIRA, and that expert customizes it the way they like. It still doesn't work very well for most people. Nobody likes it except one stakeholder, and the engineering lead who acts as a admin on it. A while later, those people have left the company, and everyone else is out of luck.
Seen this exact scenario play out at two different companies now. Am witnessing it play out in real time at a third.
IDK if this was to cheap out on the licensing with a minimal number of users, or if it was to insulate the developers from the experience of using Jira. Perhaps some of both.
Clearly that usage pattern would only scale so far.
That being said, because it can do anything, it doesn't take much effort to make a workflow as painful as possible. Somebody with the "right" mind might make all kinds of checkpoints in a workflow, which makes a lot of operations a pain in the ass because you wind up hopping through a bunch of steps. Pretty sure in our org we just make our workflow "you can hop from any state to any other state"--basically a free-for-all.
Dunno my point, but there you go!
I don’t really understand what this has to do with “monolithic” or not.
Atlassian’s software is probably very complex and convoluted but from my experience it’s almost impossible to keep a clean architecture in a software system that has grown over many years and is used and customized by many customers so you have to avoid breaking backwards compatibility.
How would they lose committed data? Even after restoring the backups can't they run the logs so that everyone is caught up?
You wouldn't be able to do that without forcing downtime for all customers, for the duration it takes to restore the snapshot and then replay the logs. Not to mention the risks of the process failing somehow
You could narrow the window to just the "replay" portion, if you were able to stand up an extra database/infra, to switch over to when it was ready. But at some point you'd probably still have to go read only to checkpoint the logs and begin the replay.
It's of course possible to do something more complicated here and stream the changes then eventually enact a failover, but this would all be too complex and error prone to introduce in their current crisis mode. It's something I'd suggest considering when architecting their DR/BCP, but it's too late for that kind of elegance (and complexity) now.
I wonder if they're not doing this due to some tech limitations, to avoid taking the financial cost of running two systems, or to avoid having to reconcile the systems.
For just this sort of problem, we actually had three DB servers running all the time: active, passive, and hour behind with the ability to break hour behind's copying of the write-ahead log of active as the DBA's secret weapon for just this problem. If all customers had accidentally lost an hours worth of data it would have been embarrassing, but much less than completely shutting out hundreds of paying customers for two weeks, I think?
This seems to be exactly what they are doing, as described in the article. They don't have automated tools to do this.
It's true that nothing is simple at scale, but it's important to note that simple concepts are the only concepts that work at scale.
Perhaps they don’t have the right people on hand to do hard things like this.
They also apparently lack an incident response plan since a critical component of that is coms to affected customers.
They also lack good practices around preventing human error. It should not have even been possible to make the initial mistake. It certainly should have involved multiple steps of “are you sure” and potentially even review.
Sounds like an operations shit show. Glad it’s not my circus.
You don't think that's exactly what they are doing?
https://jobs.boeing.com/job/annapolis-junction/jira-administ...
As an Aussie I always wanted Atlassian to succeed as we have so few tech companies at that scale or larger. Now I view them as another Oracle. Now they innovate little, they keep ratchetting up prices, pushing deployments to cloud where they make more money. Nickel and dime you for what should be core features (SAML Auth?). They aren't coming up with anything new to keep the value in the ecosystem. They buy applications in, spend a little to make some cross integration and then drop down to a slower development Cadence.
If MCB is so interested in those things, then he should do the right thing which is to retire from Atlassian and go whole-heartedly after those noble causes.
But with software it costs nothing to spin up new instances, costs nothing to deliver half way across the world, and has no delivery time. How can you convince a manager to use a software solution provided by a local company when a company in a completely different country 600 miles away offers similar software with 5 extra features?
It seems like the internet is now perfectly set up to create, for each software type, a single company that has a global monopoly.
Asking for a friend.
That no one uses.
So we didn't notice.
Yes. Always.
vi your todolists on an ec2 box
The most inexcusable thing is not communicating with the paying customers who have been affected for over a week.
Atlassian's Global Head of Customer Success probably should have been fired but here she is promoting Atlassian Cloud on LinkedIn three days ago: https://www.linkedin.com/mwlite/in/gertie-rizzo-5b70061
Actually reading a bit more, it seems like their customer team was partying in Las Vegas instead of taking care of business: https://www.linkedin.com/mwlite/feed/hashtag/atlassianteam22
Priorities.
When do execs living it up at the fucking Wynn Encore while the house burns down start to not get another job?
They’ll keep pulling this shit until it cost money.
I’m still trying to kickstart a second act years later, because I’m trailer trash and it’s hard work when you’re that.
Also, GP's quote
> Engineering mistakes happen.
I don't like this statement because it offers consolation at the expense of unintentional normalization.
Sometimes manure will always hit the fan. Being robust means being able to handle that.
Human error is probabilistic, and the probability of making an error cannot be zero.
On the flip side, it’s infeasible to use only provably correct systems; not lazy, but literally not a practical option due to compute costs, developer time, what formal techniques can even be applied to the problem at hand, etc…
They've underbudgeted for engineering, and they're feeling that now.
Let’s say there’s a 10% chance of any given feature being broken. Write a test, (which has another 10% chance of being broken) and now it’s only broken if the test and the code are broken, and broken in the same way. So we’re down to <1% chance of failure. In my experience most bugs that make it past testing do so because you forgot a test.
Then add a backup / redundancy system. That has a 10% chance of failure, but if you test it regularly then the backup / restore process only has a 1% chance of failure.
Now we have a system that’s pretty reliable in practice, made out of pieces which are only 90% reliable. And no need for PhD level formal methods.
Just do the obvious robustness steps: Write unit tests. Run them with every commit. Have a backup system. Test it. Have redundant servers. Do stages deployments. Monitor your servers and have an on-call roster. Then when everything is working well, add a chaos monkey to increase the failure rate of all of these parts so your team & software gets practice dealing with problems.
The fact that this bug slipped past all of their reliability engineering - past code review and testing into production and in a way they can’t recover - that smells of sloppy work.
The trouble was the restore would set back everyone’s data to that point in time, whereas only some customers data was impacted.
It's delusional to think software can be flawless in the real world when it's used by an untold amount of people, on all manner of devices, possibly running different OS's with different versions on networks that can be configured all sorts of ways. Thats not to mention all the dependencies involved in creating high level software, from the third party libraries to external services like cloud storage.
You anticipate there will be problems and make sure there are processes in place to manage them when they inevitably occur. Thats the exact opposite of laziness.
You're never going to get perfect error handling in any non-trivial system.
Being robust means that you plan for particular states (like "deleting the production data"). That doesn't mean that your plan is any good, or that your plan will fix the problem, only that you have a sequence of steps developed in advance of the problem.
Sometimes the state in question is considered too unlikely[1] to ever occur, so is ignored with the caveat "too unlikely", such as planning for the case when the company files for bankruptcy and all software needs to be sped up by a factor of two in order to halve computing costs.
Not all possible future states need to be accommodated for in the tech stack - that doesn't mean stack is not "robust".
[1] Or if likely, is such a large problem that all the other problems are irrelevant.
The Negative fall out was not due to the deletion of customer data, as the Story and multiple customers have stated the negative fall out was the SILENCE / lack of communications, which is Sales / Customer Service not engineering
As the comment I was replying to noted while engineering was trying to recover from what might possibly be the biggest outage in the history of the company Sales was partying and not handling customer communications
That (the failure to communicate with customers) should be a resume generating event of all leadership customer service / sales. It will not be because sales will simply redirect their failure on to engineering in the exact same manner you just have
Atlassian famously eschews the exact sales teams whose job it would be to manage direct customer comms in an outage like this one, and to be the lightning rod for the understandable customer frustration. In the past, I've been the guy that gets the angry text message from the customer and has to carefully paper over the gaps in communication from higher ups. It's not fun being the neck that gets choked.
The complete lack of meaningful communication for so long indicates to me that Atlassian doesn't have a meaningful feedback mechanism from the field back up to the executive suite - the exact feedback mechanism that Sales and Sales Engineering teams fill in most SaaS orgs. Customer Success should fill that role, but IME don't have the same incentives, pressure or influence as sales teams watching half their yearly comp go down the pipes.
Without knowing Atlassian's org structure all I, an outsider, know is that engineering had an issue and remediation status communication has been lacking. If anyone should take "blame" it should be the bad engineering practices (usually a lack of funding or inability for upper management to prioritize engineering risk) which led to an outage of this magnitude and the status update mechanisms in place. Even then, I question how much better status communication would even get Atlassian given the sheer length of this outage, which is why I think the blame should ultimately land on engineering practices. Hopefully Atlassian learns from this but I'm afraid the damage is done. A SAAS system which takes this long to heal is unacceptable in 2022.
In the contest of this outage, the "Global Head of Customer Success" should absolutely be taking the blame for failing to communicate with their customers.
You can attempt to deflect the blame purely onto engineering side of things, as people in sales often to, but you are objectively wrong.
I have suffered long outages with vendors before, the key for me was always COMMUNICATION. They should be communicating with every customer impacted, they should be telling them on going progress, where in the queue they are, etc etc etc.
Trust me as someone that make purchasing choices and recommendations the length of this outage is not the issue, the lack of communications is. As someone that makes purchasing choices and recommendations I do not blame the engineering, I blame sales or "Customer Success" or what ever PR name you want to further deflect sales to be....
Nothing, and I mean absolutely nothing, that Atlassian has to offer is rocket surgery-kind of hard... yet, here we are... not being particularly surprised at all.
But the idea of doing a full restore to a new environment, then only enabling accounts for the impacted tenants, is a good one.
They're in control of the architecture: rollback, backup, and recovery should all be considerations
But,
You can imagine problems restoring one
individual tenant's data to an otherwise
active database with many tenants
any cross-tenant primary keys
Why would multiple tenants share a database? Sharing a database server, yes, but sharing databases and mingling primary keys and such?That's such a recipe for disaster; giving each client their own database seems like the easiest win in the world.
But I'm not intimate with Atlassian products. Maybe they have some products where that's not practical for some reason.
1) Its super common even in multitenant systems to have a common database with configuration information (for example) which serves all tenants, and tenant-specific databases used alongside that to host their private data.
2) Back when sharding started to be a popular scaling pattern, tenants were not always split up by the tenant boundary but by some other reliable key. Obviously this isn't true multitenancy and I think most DBAs would discourage the pattern today. However, given the age of the products at Atlassian (and assuming a fast-and-loose engineering culture, which has been alluded to elsewhere) its entirely possible that parts of these products as well as the entire product itself may use this kind of sharding.
Bottom line, we can only hypothesize unless and until someone from Atlassian actually details their architecture (which may have happened? I dunno, I haven't been paying that much attention to it…)
1) Its super common even in multitenant
systems to have a common database with
configuration information (for example)
which serves all tenants, and tenant-specific
databases used alongside that to host their
private data.
Yeah, for sure. This is definitely what I'd expect to see, but I would also expect that to make individual client restores pretty easy, assuming the individual client backups themselves weren't trashed.One wouldn't imagine that the shared config database would have a dependency on any of the individual client databases and that they could therefore be moved/dropped/restored at will, independently of the shared config database.
2) Back when sharding started to be a popular
scaling pattern, tenants were not always split
up by the tenant boundary but by some other
reliable key.
I guess that makes sense. I mean, after all, it does allow large/demanding clients to span multiple databases I guess.On a single Postgres instance you can (at least theoretically) have 4 billion databases per instance.
In that case, the tradeoff between isolation and ease of development is made. That said, having a schema per user (even if in the same physical database) seems like a nice approach, if you can stomach the overhead and added ops complexity.
And I don't think even pgbouncer would help.
If I was in customer success at an enterprise vendor I doubt I'd be let anywhere near the tools to get this back up and running. These guys are generally in the way rather than helping in a situation like this.
Head of engineering or some product rather than customer support? That might be a different outcome.
So I don't see why partying at the same time could be an issue. Thats the engineering which made the mistake. So even though they could have communicated better ( and we don't have all the details, we don't know). The true people at fault are Engineering and product leaders.
Engineering mistakes is not an excuse to blame others, and there is difference between a mistake and removing data of production customers
And the bit about Atlassian employees being in Vegas has nothing to do with anything - as if the entire company is supposed to shut down its planned celebration because of an incident that a small subset of the company should be handling.
These are the kind of incidents where parties and celebrations are put on hold and everyone does what they can to help, regardless of department or title.
They’re talking about the outage lasting for two weeks without any communication as to what’s going on.
> Yes Andy it was and I didn't realize that my scheduled posts were still going out. They have now stopped.
However it doesn’t imply there is vindictive drive
some people will not have a good month career wise, some people will lose trust, some people will be an consequences of regaining the public confidence
This was a conference they hosted, not just some Atlassian team members partying in Vegas: https://events.atlassian.com/team22
Thousands of Atlassian customers bought tickets, flights, and hotels for the event. It's unreasonable to suggest that they should cancel it all days before the event.
> Atlassian's Global Head of Customer Success probably should have been fired but here she is promoting Atlassian Cloud on LinkedIn three days ago: https://www.linkedin.com/mwlite/in/gertie-rizzo-5b70061
She responded that it was a scheduled post. This is normal on LinkedIn: Marketers write a lot of posts all at once and schedule them to come out periodically. They're not literally sitting down in front of LinkedIn and typing out their daily puff piece.
Please, let's drop the pitchforks. It's not a good look for us on HN, especially when the facts are being completely misinterpreted.
Also posting a LinkedIn event photos with people's real names and pictures in a top post on HN along with provocative framing like "partying in Vegas" about this outage is pretty shitty. Atlassian is a big company and has a huge engineering and sales department, you have no idea if any of those people at the Vegas event had anything to contribute with the outage response. For example in your link there are literally people talking about some new mobile app their team is launching at the event, I doubt any of those people are involved in the outage response.
No one posting on HN unless they work at Atlassian in a leadership role is in any position to even start assigning blame, call for firings or publicly shaming people (the later of which you shouldn't be doing even if you 100% knew for a fact that they were at fault).
Do better.
Can't blame support for something an engineer did.
whether this is cancel culture not it definitely looks like finding an easy target to shout at
> Atlassian's Global Head of Customer Success probably should have been fired
I don't know who carries the true blame here, but calling for the resignation of a top manager isn't really unreasonable or immoral on the face of it.
The OC is directing us to information made public by Atlassian team members themselves. The OC did not make the information public. Anyone could find that information with little effort.
I’m not really sure why you attribute such intensity to the OC’s rather benign comments.
That’s my perspective. Reading your comment it feels like between the lines you don’t think this is as serious as other people. If you have a different perspective, could you come out and say it? Can you elaborate on why a “Global Head Of” isn’t a senior leader? Do you think this is an unfortunate tech outage that is to be expected from any b2b tech company of this size? If you did, I wonder how many people here would disagree with you. Implicitly, the person to whom you are responding does. “Business collapse” and “vegas party” do not look good getting caught in bed together.
Your point would come across much better if it wasn’t mixed with moral outrage. If you have alternative opinions please share them and back them up. Then you will have earned a little more of the massive amount of social capital you need to tell someone, in public and quite rudely, to do better.
My previous company lost customer data. Someone deleted a previous employee's account which surprise contained a production customer website. We didn't halt sales and outreach.
I mean RSA lost all of their encryption seeds in 2011. Every account was compromised. They still do sales today.
> Yes Andy it was and I didn't realize that my scheduled posts were still going out. They have now stopped.
The most inexusable thing when engineering mistakes happen is something like: failing to bring back people's dead friends and family members.
Not that some customers of some groupware application weren't communicated with for a week.
https://www.linkedin.com/feed/update/urn:li:activity:6918235...
I'm still unclear where "Global Head of" for something like this fits in an org chart (Who do they report to? Who reports to them? Is it within a marketing, sales or customer support tree? Etc). Title inflation and whatnot considered...
FYI my ancestors fled oppression on both sides and I'm well aware that it's a miracle I'm alive.
Again, one bad thing leading to another is a common human behavior, and the Holocaust is just an extreme example that I ABSOLUTELY did not intend whatsoever. You make this connection, not me.
Suggesting that Niemöller's poem is about "one bad thing leads to another" is like suggesting that Anne Frank's diary is about "sometimes girls have really bad days." I understand you didn't mean any offense to anyone. But that's not a license to be offensive, and then duck for cover.
Many of these users are decision-makers who decide what tools to use, and will continue to use Atlassian out of inertia due to lots of existing documentation on the tool (this is compounded by not knowing about the outages, or not knowing the severity of the outages), and also because large, professional companies use their tools too.
I don't necessarily agree with the perspective to stay with it, but it uses a lot of political capital/innovation tokens/goodwill/etc. to change systems, when there are usually higher-priority things to do (than to get buy-in to switch).
Nope. From TFA:
> I asked customers if they would offboard Atlassian as a result of the outage. Most of them said they won’t leave the Atlassian stack, as long as they don’t lose data. This is because moving is complex and they don’t see a move would mitigate a risk of a cloud provider going down.
I mean, I thought they were text.
5 days to restore text?
They must be generated by a huge complex deep learning voodoo.
Atlassian is working on the bleeding edge of technology. This outage is understandable...