As a very simple example, if your CI takes 10 minutes, your CI time budget is 6 merges per hour.
This is because if you merge two things in parallel without validating CI for the combined changes, your main branch could end up in a broken state.
Merge queues run CI for groups of PRs. If the group passes, all the PRs in the group land simultaneously. If it does not, the group is discarded while other group permutations are still running in parallel.
This way you can run more "sequential" CI validation runs than your CI time budget allows.
In our monorepo, we get a volume of 200-300 commits per day with CI SLO of 20 mins.
Without a queue, our best case scenario would be getting capped at ~72 commits per day before seeing regressions on main despite fully green CI (in real life, you'd see regressions a lot earlier though because throughput of PRs is spiky in nature)
That is a way of handling even higher volumes than GitHub is talking about, at the cost of a system that is a bit harder to think about. From the article:
With GitHub’s merge queue, a temporary branch is created that contains: the latest changes from the base branch, the changes from other pull requests already in the queue, and the changes from your pull request. CI then starts, with the expectation that all required status checks must pass before the branch (and the pull requests it represents) are merged.
Uber's[0] implementation, for example, does some more sophisticated speculation than just picking up whatever is sitting on the queue at the time.
Queues come with quirks, e.g. small PRs can get "blocked" behind a giant monorepo-wide codemod, for example. Naturally, one needs to consider the ROI of implementing techniques against aberrant cases vs their overall impact.
[0] https://www.uber.com/blog/research/keeping-master-green-at-s...
I scrolled down to the how does it work section where the first sentence is:
> Merge queue is designed for high-performance teams where multiple users regularly commit to a single branch
Half of the how does it work section is buzzwordy fluff.
This is why we built merge queue. We’ve reduced the tension between branch stability and velocity. Merge queue takes care of making sure your pull request is compatible with other changes ahead of it and alerting you if something goes wrong. The result: your team can focus on the good stuff—write, submit, and commit. No tool sprawls here. This flow is still in the same place with the enablement of a modified merge button because GitHub remains your one-stop-shop for an integrated, enterprise-ready platform with the industry’s best collaboration tools.
The problem is probably whoever wrote the blog post (who is likely not even the named author, depending on how their marketing team does things) tried to add a lot of high-level stuff to make it make sense to them without really needing to understand the details, and then dolled it up with a bunch of useless vapid quotes from customers and what not, because that is what marketing people think matters. Maybe it does make sense to have mealy-mouthed corporate speak for the overall product, since some executive is probably deciding whether to use GitHub as a whole and they might care if a big company uses it. I don't know that it makes much sense for specific features like this, especially in a fairly technical product like GitHub.
The problem is that if you have multiple branches going into the same (mono)-repo, then they might all pass a localized CI-check, but fail if they are all merged. This is because the branches have an interaction between them. It can lead to a stall in commits and because everything hinges on the repo, work is going to stall as well.
So you serialize the branches, and impose an order on them: [x_1, x_2, x_3, ...]. Now, when running CI on one of these, x_j say, you do so in a temporary branch containing every branch x_i with i < j. This will avoid a stall up to branch x_j, if you started to merge the branches in order. If CI fails on branch x_j, you remove it from the list (queue) of branches to be merged and continue.
Beyond that, their API docs prior to the acquisition were some of the best in the industry, readable and concise. Now they are just a complicated mess.
Comms teams are really terrible in this regard. They insist on a singular 'voice', which means that every article is going to go through their review and get rewritten to their standard - that standard may involve removing technical content and instead making it more layman/ marketing friendly.
It's an incredible mistake that I see made everywhere after companies hit a certain size. It then becomes up to engineers to build their own engineering blog with less oversight and then guarding it from the comms teams, which most engineers aren't interested in doing.
Yes. The first useful line in the article is
"With GitHub’s merge queue, a temporary branch is created that contains"
and to reach there you have to skip fluff paragraphs halfway down the article.
What I think it is: instead of you trying to merge into the main branch, you try to merge into a branch where all pull requests before you are already merged in.
That way any pull request before you can't cause any merge conflicts, because they are already taken into account.
At least that's what I deduct from all the marketing fluff. Maybe I'm completely wrong.
Merge queues address the problem of how to (1) merge in a lot of changes (2) while guaranteeing no breaking/conflicting changes are merged.
[1] https://docs.github.com/en/repositories/configuring-branches...
This explanation is actually a lot better
Let's say you're an open source maintainer with 3 pending Pull Requests to merge: [1, 2, 3]. Each of which is based off `main`, has passed CI and has been approved.
If you merge all 3 at the same time, there is a chance to break the build: Your CI is testing `main <- 2`, but you're merging `main <- 1 <- 2`. A common example would be when (1) is a user-supplied change, and (2) is a dependency/localisation change, which don't cause merge conflicts but they do break the build/tests.
To do this safely, you need to re-run CI on (2) after merging (1), which is currently a manual process: you need to know that (2) is next to be merged, then rebase/pull + rerun CI for (2).
(There used to be a manual step of 'merge once CI is passed' here, GitHub has recently improved this workflow to allow automation)
Merge queues fully automate the safe approach: it merges (1), runs CI on (2) which fails, then runs CI on (3), which passes and gets merged.
This fixes that. It removes the race condition that exists because of the gap between testing a branch and merging it.
The solution is very simple - have a queue of PRs and automatically test & merge them one at a time.
There are some optimisations you can do to speed things up a bit, e.g. testing a bundle of PRs all at once, but that's the gist of it.
It is basically essential on any repo that has a high rate of PRs. I'm surprised so many people here haven't heard of it.
Gitlab has the same feature but they annoyingly called it something worse - merge trains, and it's only in Gitlab Premium.
There are many out there but I don't know about the "latest". GPT-4 itself says it has only a 10% chance of having been generated by a LLM.
These detectors are really unreliable. I've fed them content that I generated from GPT-4 and they never detect it as AI-generated.
I pity the students whose teachers will use them to detect plagiarism.
1. get a corpus of real text
2. generate a corpus of AI text
3. train a model until it can tell the difference
The problem is step 2 is semi-expensive and step 3 is really expensive, so everyone is trying to shortcut the process, and of course it doesn't work.
* Normally, the big green button says "Merge pull request"
* Now, the big green button says "Merge when ready"
In a large project with lots of activity, a stampede of people pressing "Merge" at the same time will cause trouble. "Merge when ready" is supposed to solve this.
It seems to mean:
> "GH, please merge this, but take it slow. Re-run the tests a few extra times to be sure."
[1] https://docs.github.com/en/repositories/configuring-branches...
I’ve read docs several times and never found them very clear about the details.
> The result: your team can focus on the good stuff—write, submit, and commit. No tool sprawls here.
The good stuff? Tool sprawls? Is this written for teenagers?
> Merge queue is designed for high-performance teams where multiple users regularly commit to a single branch.
I think you meant "highly active". High performance means something else. But I can kind of see it emerging from your awful sales person brain.
Who, exactly, is it necessary for? The original commenter getting their rocks off on insulting someone else’s job? Others coming in and laughing at someone insulting someone else? Critically necessary.
It’s not like the original article author is going to come in, see this comment, and reflect deeply on themselves and their work.
So you sit next to them and therefore know what their assignment was and how well they executed on it?