All HTTP-based services unresponsive
status.bitbucket.org
status.bitbucket.org
At a guess:
- Bitbucket Pipelines
- Webhook workers
- Front-end web servers
- SSH push/pull workers
Basically anything that's elastic to demand. Presumably the cost of AWS storage makes it not worth it for the Bitbucket team.
Is it a cost concern, is DC reliable enough that it's just an accepted risk, or is there some other reason?
I should note that we have both publicly and privately reachable resources in AWS. The publicly reachable resources have fail-overs built in for situations like these (it happens automatically), but the private reachable resources with our architecture depend solely on AWS Direct Connect. For example, our Bitbucket failure today was due to the fact that we rely on AWS Direct Connect to link between the Bitbucket Cloud components that we host in our data centers and others that we host on AWS. Bitbucket could continue connecting to services in our own data centers and the public Internet/AWS, but could not talk to the privately reachable resources in the Atlassian infrastructure hosted on AWS.
We understand the importance and the impact for our customers, and dedicated several teams to this issue as soon as it was reported. AWS has resolved the issue, but we will look into ways to help prevent and better mitigate these types of issues in the future as part of our incident review and improvement processes.
My side project, StatusGator, monitors something like 250 status pages and there's quite a spike in warn or down notices at the moment that I can see.
What are people doing wrong with how they use S3? Until AWS provides a cross-region S3 that is master-master or self-healing, the suggestion that people are using AWS improperly seems incorrect.
Hosted JIRA is down too (at least for me).
Interestingly, I can seem to be able to find it only via search, it doesn't show up on the frontpage at all.
I've configured my push such that it will deploy to both Github and on my Gogs meaning that I always have an up to date repository in two places.
https://github.com/gogits/gogs (if you're interested in setting up your own)
https://stackoverflow.com/questions/14290113/git-pushing-cod... (push to two remotes)
You don't have to use just BitBucket, or rely entirely on your self hosted git service.
But integrating those two repositories with all the automated workflows is quite a headache. Jenkins won't automatically switch over to another git backend if one fails. Other automated tools like code review are mostly relying on a centralized repostiory. I don't see any simple solution to solve these issues. Bitbucket, GitHub, GitLab don't have any easy fallback solutions when you relying core workflows on their services.
Self hosting is maybe one solution (we do this with bitbucket), but this requires major administrative effort to keep it running reliably (always available, no data losses on hardware/software failures).
I think if your primary repo that's hooked into CI goes down, you're still SOL. Having your own repo just enables you to continue local development among your team.
I don't see any great solution, other than making your app distributed to begin with and doing your build/deploy manually.
As a complete aside, I've fantasized about deploying to all cloud vendors (Azure/GCE/AWS/Heroku/DigitalOcean/misc.) each with their own specific build/deploy and having persistent state shared with something like CockroachDB. Having some load balancer managing state between all instances. Taking advantage of the free/basic tiers provided by all of the vendors.
https://status.aws.amazon.com/
If there are network issues then it wont matter if you want to read or write to bitbucket
Granted, there's no economical way you could self-host a Gittea instance to avoid this.
How are your git repos stored? Attached block storage? Seriously? Those aren't highly available at all, and it wouldn't even work in your proposed setup because each instance will have its own block storage.
No, that won't do. You'll have to implement some form of cross-region replication, possibly blocked by S3 or something. Good luck. Its not easy.
The only way I can think of accomplishing it is to make use of GCP's free f1.micro instance, so you can spin up two of those for pretty cheap in different regions ($3.88/mo). Have DNS hosted somewhere that can resolve to each of your two instances at random ($12/year? Lets's say $1/mo). Then you have instance storage; good luck finding globally redundant block storage for practically unlimited repositories, with backups, for $2.12/mo. Let's just leave network egress charges out of it, since those would be marginal.
And let's go ahead and say I value my free time at a conservative $30/hr. It takes me an hour to set this thing up and maintain it every year; horribly conservative. That's an additional $2.50/mo.
Maybe you can do it. It isn't laughably economical.
- Jira - Ring Central - Github
I suspected AWS but don’t see anything on the status page.
Isn’t part of the point of a DVCS to not be overly held up by an inability to access your server?
I don’t know about your situation but I know that a company like Github can probably do better at SRE than I can with a self-hosted server
GitHub maybe but for BitBucket, I'm not sure at all that it's true.
One former workplace had to bring in multiple Jira consultants due to poor performance and stability
The only situation in which self-hosting or ditching BitBucket/similarly large providers will help protect you from the fallout of historically-large, catastrophic attacks is if you do all of your development on the same local network as your hosted server (and don't rely on any internet services to access it).
And even if you diligently self-host every part of your own services (not as easy as just plop a gitlab/gitea install on a host you own and start it up), you have to deal with the fallout from internet-breaking DDoS attacks and other malicious activity if you want to use the internet to run or use your code: from congestion caused by compromised devices in your network "neighborhood" (same/similar ISPs or last-mile providers) to DNS outages to BGP hacks, we have seen time and time again that, if not exactly centralized, the systems that comprise the usable internet are certainly highly interdependent. Large-scale attacks of many kinds compromise them.
Instead of tantrums, it might behoove users to understand what kind of SLAs they can promise in order to operate their self-hosted services in such an interdependent environment. Some examples:
Do you need local power? A local ISP to be up? More than one? If you have more than one internet link, how do you pair the connections--if it's via BGP, what happens if the central authority on that has issues? If local power is down, does your local ISP's connection stay up? How long does it stay up (is there a node/amp somewhere on the line that cut to battery)? Do you need to access internet services by hostname? If so, do you do local DNS caching? If so, how stale can it get in the event of a loss of external DNS? Most importantly: how much (it's a nonzero number unless you're developing for yourself, by yourself, on your LAN) dependence on external services are you comfortable with, and how much time are you willing to spend eliminating the long tail of such dependencies?
But I agree, it's not ideal that Bitbucket has been experiencing issues. I am also looking at alternatives, most likely using AWS CodeCommit and Upsource. It's been at the back of my mind to move away for a while.
* sometimes new things break in unexpected ways
* sometimes things get changed for the Rev. B and those revisions don’t get done in Virginia because they’ve already completed the Rev. A rollout.
Also, there’s a secondary effect: because it’s the “default” region, it has a LOT more tenants, which means it probably has scaling and HA problems that none of the other regions do.
> Some component services are currently unreachable due to an upstream incident on a cloud provider. We're attempting to route as much traffic as possible away from the affected components, and are working with our vendor now.