Making GitHub CI workflow 3x faster
github.blog
github.blog
The builds for acceptance/prod we still do on remote CI/CD, but those happen much more rarely. Also the advantage of a local CI/CD is that it is much easier to setup caching for things like Node Modules, NuGet etc.
We use Azure Devops which has a pretty good local CI/CD story, installing an agent was quiet easy for us.
0: https://github.com/iterative/cml_cloud_case/blob/master/.git...
0: https://docs.github.com/en/free-pro-team@latest/actions/host...
Why not simply disallow the PR to merge unless CI passes? Why stop all integrations? There must be something I'm missing.
Edit: I think I get it. Normally PRs to Github.com non-enterprise go through as normal. 45 minutes later, after that merge, the 2 expensive Enterprise CI tests complete and that's when the developer gets the 72 hour timer. The dev has the option of reverting their change and trying again.
It’s like when you merge a documentation PR while the tests are still running: You’re pretty sure that the tests won’t break. With this solution however you’ll still be pinged if they do break, later.
Read the article. It says:
> If the CI job remains broken for more than 72 hours, all deployments to GitHub.com are halted
Edit: it even creates the perverse incentive of trying to get your own stuff in quickly before the 72h window closes that s/o else caused.
You don’t even need to hit a 2% success rate to get people agitated. Consecutive failed builds that happen every few weeks will eventually happen when someone really needs to get something out Right Now and that incident will come to define their experience, especially if it happens two or three times. Even if it’s just to people they know instead of themselves.
Anything that happens once a day happens “all the time.” For some people, that’s true for once a week. For others, if it happened twice in two weeks and once every six weeks thereafter.
If significant numbers of devs found the whole process stressful, I wonder if it would help to make the reversion happen automagically after, say, 24 hours, if the dev hasn't committed anything in that time
Devs who have called in sick or over the weekend or attending a wedding or long-planned vacation don't have to think about it.
Stopping the world seems like an automated process. So should reverting be. In fact, there should be no stopping the world at all, just reverting if the dev can't fix it within 72 hours. There might be cases where automatic reverting doesn't work, but those should still be managed through organizational processes, not by stopping the world.
I'm an unusual American, in that I consider healthy work-life balance to be top priority, at least for everyday employees (founders and such have a different path).
I do try to make room for the idea that my understanding of this is likely superficial, and that there may be things that I am missing, here.
I would guess that this wouldn't introduce more or less build breaks. People will do as they've been doing. The psychological argument can be made that people could be either less safe with their commits, since they have 72 hours to fix any problem breaking a long-running CI job, or they could be more safe with their commits, fearing they'll break a long-running CI job. Personally, I don't think that either position has a lot of strength.
Or does "deploy" really mean "merge"?
GitHub takes continuous deployment very seriously, with dozens of deployments every day. My understanding is that they try to avoid having code sitting around only deployed to an internal environment any more than truly necessary. They want to be able to fix code issues discovered in production within a shockingly short time frame if the issue is not complex.