Scaling Bitbucket’s Database
bitbucket.org
bitbucket.org
I’m not sure if I’ve missed something with their LSN work or if this indicates that the GET semantics horse has already escaped their proverbial barn. Within the call-response of a single request, which they seem to be talking about, none of this should be necessary. Right?
However a cascade of narrowly-spaced follow-up requests could easily catch you in this trick. Lie. Whatever you wish to call it.
Free private repos used to be the one thing keeping me (and from conversations, several people I know) on Bitbucket and once GitHub offered those as well, there was no incentive to stay. The Mercurial support was probably another unique feature they had from way back but they seem to have decided to remove that too.
Central login where you can never understand what login you are using and what information they are collecting, nor on which website you are not whether you’ve suddenly integrated Confluence into a Bitbucket instance, is a real downside. God I hate SSO’ed companies, they went all the way of Google into Youtube.
which isn't true anymore, as other pointed out.
Also offering Mercurial was a reason for choosing Bitbucket, but they are ending its support.
See "Want a more powerful issue tracker?" at the bottom of https://confluence.atlassian.com/bitbucket/issue-trackers-22... for the documentation of this policy.
It’s hard to take them seriously when that’s the tableau upon which they are advising us to build systems.
At some point their primary is going to get overloaded again by the writes. And they'll have added all of this machinery for nothing. They've also made the replicas an essential part of their stack. Whereas their previous stack could probably tolerate a certain amount of replicas going down or replica lag. They now have a system that will grow and be dependent on all the replicas being available. Until a heavy write load or heavy table alter, causes the replicas to become lagged at which point a higher percentage of traffic will go to the master and potentially cause downtime.
Wondering if eg cockroachDB would be a good fit
There's also hybrid setups of sync+async for durability reasons, sync to a single replica so that if your primary db suddenly catches fire you have a replica that's as up to date as possible, then read replicas using async replication.
Feels like there is something there tho in terms of approach, esp since they are keying it on user id.
The article doesn't explain how they know the current LSN of each replica?
"We have to query Elasticache and replicas, which adds latency."
For hg->git migration, https://github.com/frej/fast-export worked smoothly for me.
mkdir new-git-repo
cd new-git-repo
git init
fast-export/hg-fast-export.sh -r ../path-to-old-hg-repo
(probably want to translate .hgignore to .gitignore and commit it)
git remote add origin git@path_to:git_server
git push -u origin master
git push --all1) this afaik couldn't be done inside the BB project, so one would have had to create a new repo and lose bug history etc.
2) It's a trust thing. If BB kills repos once, who doesn't say they won't kill others tomorrow?
3) Comfort: Import to GitHub is a click away
If github goes downhill, at least I'll be in good company next time!