How long should your CI take?
graphite.dev
graphite.dev
I've always wondered if it would be possible to design some proof of work concept where you could hash your filesystem + test artifacts to verify tests passed for a specific change.
FWIW in yeas of development I've never had an issue where "it works on my machine" doesn't equate to the exact same result in CI.
I wish it was more nuanced
cmnv = commit --no-verifyI've been impressed by Pants (Python build tooling) which manages this really well https://www.pantsbuild.org/docs/advanced-target-selection#ru...
Most projects I've worked on with multiple contributors on have had a merge queue where you simply tell it to merge if/when the CI passes, then you can move on with your day.
of all the things a human should be in the loop for, merging to main seems like one of them
That's what the review is for. And I'd argue that the review should be decoupled from CI, because you're not reviewing if you think the tests is passing before/after merge, you're reviewing the code and docs themselves.
And once the review is done, it should be fine to merge at any point, today or tomorrow, barring any merge conflicts of course.
Merging to main should be the most mundane task ever, and even a robot should be able to do it with confidence.
Sometimes there are considerations like QA/infrastructure resourcing that dictate when something can go into master. Ideally that's rare but I have worked in a place where it's a consideration for every single merge.
To expand on this a little: we had a simple branching model. To/from master for work, branch from master to cut a release. Being highly regulated, as a matter of (unchangeable) policy, every change needed testing by QA team. QA team == QA guy. Among feature work without hard release dates, there was also bug fixes with some urgency, and regulatory work with hard release dates. Given we traded flexibility in our branching model for simplicity, we instead had to consider if something could be merged to master without impacting QA load of anything already merged but not released. It worked fine for us but were were a team of <5. My remark about trading flexibility is really the crux of it all though. There's compromises that need to made and this one made sense for us at the time.
And if your changes have been QA'd in staging and/or they aren't that hectic (and are backwards compatible) then I think auto merge into prod can be fine too.
I basically do not think "auto merge into prod" is ever OK, but then again, I let GitHub Copilot write 40-60% of my doc comments, so...
Funny how different people can be :) I'll always try to make the main branch so clean that it can be deployed at any moment, and production will always use the latest commit as soon as possible.
Basically, I'm doing CI + CD, but I understand it's not for everyone.
yes
(in my experience)
Note that my CI system runs multiple builds in parallel (as yours should too!). Any individual build might only take 10 minutes, but it is 2 hours if I try to run them on my local machine. Odds are if my code builds for one linux variant on x86 it will build on all the other ones and arm so I don't expect someone to run those before starting a review.
> but we have a lot of static analysis that we run as part of CI, some of it standard some of it custom for people who violate our code style in some way.
Yeah, but why would issues that static analysis find stop you from doing the review? I never review things that automated tools can find, I care more about reviewing things only a human could review. "Does this make sense here?", "Is this decoupled enough/too much?", "Does this work well with the overall architecture?", "Does this test test the right thing?" and so on.
I wouldn't spend my review time on spellchecking people or pointing out issues like that, that's for the tooling to do. And even if the CI fails because of some analysis, PR author fixes it, every review comment should still be applicable, otherwise you're just doing robot work.
If you have something that's easy to test locally and the CI checks are just a backstop then the initial PR should nearly always pass, but if your CI checks are much more thorough than a developer can do manually then a PR initially being in a broken state may be a routine occurrence.
Sublime Text works pretty well for me. I think you need to turn on spell check in the settings (off by default?) Then you can choose which syntax highlighting scopes you want to spell check. I have mine configured to spell check comments and string literals.
I've also worked on too many projects without a merge queue.
CircleCI dutifully polling the ~750k-line stdout for the find command was taking up about 3 minutes of our 5 minute process. Was done so someone could debug a failing build... supposedly.
Doing it for a one-time CI run is fine. Forgetting to remove it is less fine. Doing it on CircleCI who have had builds you can SSH into since forever is a lot less fine, it should never have hit the config at all.
[1]: https://www.computer.org/csdl/magazine/so/2023/04/10176199/1...
The CD ensures that the code is only deployed after all tests passed, it was approved, merged, and ensures a deploy isn't skipped accidentally. Depending on what your doing it could take a while to fully build new binaries, build for multiple platforms/versions/architectures, etc.
Unless you take great care in writing hermetic tests, testing on your local machine won't catch problems hidden by non-standard versions of libraries and tools in your local development environment (e.g. you installed modern versions of bash and python on your Mac, but your users are using what the OS comes with).
And for large enough projects, running tests on a single machine is simply impractical: a full test suite run on a single machine would take hours; the only solution is to shard.
And you could set up a trigger to build it when you make a commit so that you don't need to manually trigger every machine separately.
And then you don't need CI!
I see this class of issue often in early startups that don’t have robust CI implemented yet. Like literally once a month or more per company.
[1] Unit tests, maybe, but not integration.