Slow CI: real problem or easy excuse for developers?
codingnagger.com
codingnagger.com
- Not parallelizing tests. Microsoft runs something between 100k and a million tests per deploy, and that's only possible if you can run your tests in a distributed fashion.
- Not bundling tests by stage (re-doing setup/teardown, re-doing nearly identical tests, tests at random times). Similar to the above, you have to group your tests into stages where they can be parallelized together efficiently. Example: you're a network device company and you only have a couple racks worth of gear you can test on, so you have to optimize for every device and stage, or your poorly coordinated tests will tie up infrastructure someone else could be using. (In the cloud with dynamic infrastructure, this wastes money as well as time)
- Not optimizing tests for speed. You don't have to "micro-optimize" to occasionally do build and test profiling and replace the really slow bastards. If you never do build or test profiling, you will end up with slow junk.
- Too many/unnecessary functional or end-to-end tests. You don't need to do flyway migrations on every build, you need to do it when you've changed the schema.
- A gigantic monolithic codebase. If your code is huge and in one repo, you're going to end up with 100k+ tests, because a one-line change might affect anything in the codebase, so you better test everything every time. If you design your apps well, you can segment your application into discrete components that do not depend on each other. This way you can do unit testing on only the component you changed, and then move on to functional/e2e testing for the whole shebang.
- Using sub-optimal tools. Changing a build or test from using one tool, library, or service to another can give exponential performance increases in things like builds and tests. Lots of people have gotten dramatic speed increases just by changing a tool. If profiling doesn't make a big dent in speed, this often can.
If you design your apps well, you can segment your
application into discrete components that do not depend
on each other.
And if you don't have any full e2e tests in CI, you can test whether you did the segmentation right IN PRODUCTION!In my experience, this is spot on. In the CI system I help run for testing graphics, a significant amount of the CI time is spent:
1) fetching source code from external locations for all dependencies that need to be built for each test, and for the project itself
2) fetching source code from internal git cache to builders
3) compiling projects (ccache helps a LOT)
4) syncing built artifacts to shared storage (NFS)
5) testers sync artifacts (rsync)
6) testers run tests
7) results sync'd back to CI master
8) process results and display to developers (the processing and display of results also takes a few minute, since it's often over 1 million tests results to process).
Many of the steps here are repeated multiple times, since we run multiple test suites across nearly 200 individual testers. Our we average around a 30 min turnaround time for results.
Steps 2, 4, 5, and 7 take a surprisingly long amount of time, since internal network utilization depends on CI load (e.g. it could be running other CI jobs at the same time). We are constantly looking for ways to improve it.
For some of the tests that run, the syncing of artifacts and results take longer than the tests that run on the tester.
And if you have segmented your tests in fast and slow, you can always invest in the proper infrastructure to only require fast tests to pass before CI and have some degree of automatic rollbacks (or supervised) when slow tests break.
You can explain how much your test should take time, like: short(60secs), moderate(300sec), long(900sec), eternal(1hr, not really eternal hehehe).
Then you can specify the special rules for testing, like:
- exclusive - run no other test at the same time
- external - test has an external dependency; disable test caching
- large ... etc. etc.
For example small, and fast unit test can be executed in parallel in machines (containers, whatever) with much smaller requirements, and compartmenalize better, to the point where it would be reasonable to set some strict rules - your "super-fast" unit test, should not eat that much memory, or you have to qualify it as such, and then it'll be treated differently.https://docs.bazel.build/versions/master/test-encyclopedia.h...
If you have suitable infrastructure (this is not easy, but think tracing system calls, filesystem accesses, and function calls), you can automatically track the dependencies of tests as they run, in a similar way to automatic tracking of dependencies when building things.
If you do that, your discrete components will emerge naturally and don't need artifical separation to reduce the amount of testing.
Even better, if it is properly robust, you can safely cache test results and automatically skip tests that don't interact in any way with things that have changed since the last time.
That completely changes the scaling of test times for large codebases.
In principle, this can be more reliable than separating components into separate repos and pretending that testing them separately is enough, because changes in one repo can affect functional test results in another, and automatic tracking may pick up on that.
All that said, it's difficult to set up the infrastructure to trace things at the right granularity to limit false dependencies (for example, referenced functions that aren't called in a particular test, linked libraries that aren't used, etc.)
Later on developed more advanced way with dynamic tests allocation across parallel CI nodes not only for Ruby to get CI build as fast as possible. To give you some idea how it works check https://docs.knapsackpro.com/2017/auto-balancing-7-hours-tes...
or watch video https://www.youtube.com/watch?v=hUEB1XDKEFY
I rather not have people on my team waste hours writing tests that break the moment props change, or styling changes, or a child component renders breaking the test because they used mount instead of shallow. IDK, I think I just hate Enzyme mostly...
With the backend you should have 90% unit tests (1000s of them) all completing in less than 3 minutes, integration tests (7% of all your tests) competing in under say 10 minutes, and feature tests (those shouldn't be in the build but in a post build step that can reject a change)
IMO, the correct amount of testing is the amount that lets you refactor quickly and with confidence that you haven’t broken anything, and that tests pathways/edge cases in complicated logic.
If you have passing feature tests, the software sort of works. If you have 100% passing unit tests with 0 functional tests, the software might not even work at all.
Nope. The only thing it does is pass the test under the conditions that you made the test. There is no validation of the code within the project for correctness, safeness [is validation working or not], or even operation. When you only have feature tests, you have a lot of code that you're not sure that works or not.