Encountering some turbulence on Bitbucket’s journey to a new platform
bitbucket.org
bitbucket.org
Gitlab went thru a similar journey from NFS to a high level git api called gitaly:
https://about.gitlab.com/blog/2018/09/12/the-road-to-gitaly-...
https://gitlab.com/gitlab-org/gitaly
There are some other projects like this one that seek to address the problem:
https://github.com/takezoe/gitmesh
Git is already good and synchronizing between peers, but it's not low latency, so does require an extra management layer to make sure everything is correct.
1) Locking. It's a pain in the ass. You're probably going to need to take it over or otherwise work around it, some how. I never did[0], but likely should have. At scale and with unreliable HTTP operations and all kinds of crazy stuff triggering writes to repos, you're going to end up with locking problems at some point.
2) Caching. Cache the hell out of metadata. Cache entire repo-wide metadata read operation output. Cache commits. Cache archives. Cache, cache, cache. Cache early, cache often.
[0] in my defense, I built the whole thing solo and there are only so many hours in a day, and that was not the only thing I was working on.
[EDIT] this was, like, 2011 or 2012 or something, so there was a lot less info floating around about how to do this, too.
Sometimes I just want to quickly merge a small change (maybe a small config change, 1/2 lines) and then pull on master, branch off and start working again.
I'm regularly having to wait several minutes for the merge and while I know that there are ways around this locally it just annoys me that something so simple is taking so long.
This feels less like an apology and more like Atlassian saying "Us changing a platform which has worked a certain way for years and that breaking your workflow is YOUR PROBLEM. It's you looking at this wrong, merges have been asynchronous all along." despite our many combined millennia of experience being entirely to the contrary.
If I could distill the message re: slower merges down to 2 essential points, they would be this: (1) we underestimated how impactful this would be for some customers and that's on us; (2) some, in fact I think many, users believed they need to wait for the merge to complete and we wanted to clear up that misunderstanding.
From your use case, where you merge a small change and then want to pull, create a new branch, and start working again right away, I understand this directly affects you. I am surprised merges would be super slow for you if you're just merging small changes, though; average merge times are still just a few seconds. Have you opened a support case?
For many other users, I do think the UX changes we're rolling out will make a difference. There are a lot of users who would click merge, maybe 5-10 seconds would pass, and they would assume something must be wrong so they'd refresh the page and then it would look like nothing happened. Today we pushed out an update so that if you refresh the page and the merge is still in progress, you'll actually be able to see that.
FWIW we do have some longer-term work in progress that will make merges (along with basically all file system I/O) a lot faster; but it's a ways off and represents yet another significant architectural project (though much less disruptive than this one!). I didn't mention it in the article because it will take a while.
I suggest you up the capacity on the queue so it feels syncronous and snappy like GitHub, as this is now a pain point.
That’s not the only reason merges are synchronous, especially in the Atlassian world.
At scale, we have problems with build triggers from commits. Sometimes you have to fire manually.
At scale, tracking deployment failure is painful. The PR you merge may be in a different module than the deployment plan. So I need to merge a PR, then watch the dominoes fall. But the bigger issue is that Atlassian never finished getting deployments up to feature parity with builds. I have several dashboards that report build health, but deployment health is hard to track. It involves more vigilance and frankly it’s exhausting. Exhausting things get dropped every time people feel tired. To work around this, first you do alerts to team chat, then there are too many alerts so they go to their own channel, then people forget to check the channel, and we have a smaller version of the same problem and no solution.
At scale, I can only control whether MY code has enough tests to detect regressions prior to deployment. Breaking preprod impresses no one, even if automation should have caught it. That means a workflow of merge-build-test prior to moving on to the next task.
> our engineering teams prioritized their efforts to ensure that Bitbucket
It's not a coincidence that company that refers to "our engineers" rather than "we" is having basic engineering problems.
> Our focus was primarily on read operations
> In contrast, we accepted that there would be increased latency for writes
It didn't sound like they were throwing the engineers under the bus at all to me. They seemed to be highlighting that engineering executed well on the items management wanted them to focus on.
Seems to be expected. Things rarely deploy flawlessly.
Sadly BitBucket isn't what it used to be ever since they dropped Mercurial and Microsoft acquired GitHub (and introduced the super generous Actions / build servers). I see no reason nowadays to recommend them.
This is terrible, sounds like this is the new norm. If I have to merge in 10 pull requests I can't wait 5 minutes per PR to see if there are going to be any merge conflicts, etc.
Ever since this popped up I'm seriously considering migrating everything away from bitbucket because of how slow merging is.
I cannot believe this isn't a priority at Bitbucket and we're being told to get used to it and that the UI will hide it. I've been putting off migrating to Github because we use a lot of integrations, but Bitbucket has suddenly become a pain point where it was okay before.
The reality is that this migration is one of the clearest signals I can point to that the company is investing in Bitbucket. Our platform teams have been awesome and given us a ton of support, including features that didn't exist before (without going into too much detail... you can probably imagine that Jira and Confluence don't have nearly the same requirements around file system access that we do).
Yes, Bitbucket Cloud and Bitbucket DC are two different teams and code bases; but we work closely together and are talking a lot about both products' roadmaps and future vision. And we even have engineering teams working together on a shared project and may do more of that in the future.
https://web.archive.org/web/20160303204710/http://www.itwire...
Atlassian makes me long for Trac, which at first was an epithet, but I’ve come to see it more recently as a distilled product that doesn’t distract from the job at hand. There are a lot of things that get done in the tools because the tools have features, not because the tools are good at it, or because it’s a good way to accomplish a goal.
Mostly what I use are histories, and linking between docs, commits, and builds. If confluence had a way to start a task list in the docs and move it to Jira, that would be one thing, but it can’t even do that. It’s all feature factory work instead of workflow-centric, which is what most of the users actually need.
For many companies, it makes sense to use AWS (or a product based on it). For a company like Bitbucket, that is almost like if AWS decided to stop self-hosting, and use Google or Microsoft's cloud products instead, with their "product" just a wrapper on top of that. What is the point of Bitbucket, again? It's now a wrapper around a wrapper (Atlassian) around a wrapper. Except, it doesn't work that well.
It's a good thing people don't use Jira to write softw— oh.
https://confluence.atlassian.com/bbkb/what-is-atlassian-s-bi...
You would have to think of some way of federating git repos onto local storage (local to the git process), to avoid network latency inside loops, etc. But maybe that's what Atlassian did and it still didn't work out, or maybe they thought of that but something else (have to fit into the Atlassian cloud architecture?) outweighed it.
Point being, you really need to design the cloud architecture around the way your hosted programs were designed to work.