Show HN: FlakyBot – Identify and suppress flaky tests
flakybot.com
flakybot.com
Most CI systems leave it up to teams to manually identify and debug test flakiness. Since most CI systems today don’t handle test reruns, teams just end up with manually rerunning tests that are flaky. Ultimately, tribal knowledge gets built over time where certain tests are known to be flaky, but the flakiness isn’t specifically addressed. Our solution, Flakybot, removes one of the hardest parts of the problem: identifying flaky tests in the first place.
We ingest test artifacts from CI systems, and note when builds are healthy (so that we can mark them as “known-good builds” to use while testing for flakiness). This helps automatically identify flakiness, and proactively offer mitigation strategies, both in the short term and long term. You can read more about this here: https://ritzy-angelfish-3de.notion.site/FlakyBot-How-it-work...
We’re in the early stages of development and are opening up Flakybot for private beta to companies that have serious test-flakiness issues. The CI systems we currently support are Jenkins, CircleCI and BuildKite, but if your team uses a different CI and has very serious test-flakiness problems, sign up anyway and we’ll reach out. During the private beta, we’ll work closely with our users to ensure their test flakiness issues are resolved before we open it up more broadly.
Another way to think about it is, whether Flaky tests are worth keeping? At some point if the tests fail often, do these really add value. And we think - it does. If you are able to identify flakiness from real failure and reduce noise, you can still avoid real failures.
There may be a point where the cost of ownership for a specific test exceeds its utility, but the way to resolve that is usually to reevaluate your code and supporting tests. Suppressing flaky tests seems a very unwise choice.
Perhaps under extreme circumstances and with unhealthy code bases there may be a case for this, but I struggle to imagine it.
Fixing flaky tests can very commonly take longer than writing new tests.
Once the test is skipped, a domain expert can come back and take a look and figure out why it was flaky, and fix it.
If it's urgently broken (e.g. there is real impact), we treat it like an incident and gather people with the right context to fix it quickly.
As long as everyone agrees to these norms, it's not a huge burden to keep this up with thousands of tests. People generally write their tests to be more resilient when they know they're on the hook for them not being flaky, and nobody stays blocked for long when they are permitted to skip a flaky test.
In another case observed, devs just got used to rerunning the entire suite (the flakiness here was about 10-20%)
Yes, I'm mostly agreeing with you that the tests should be fixed, but I have seen ones that were perfectly fine (given the constraints) and what should have been fixed was the CI.
They won't be fixed until they start actually preventing commits. If somebody deletes a test, that is on that person. I don't want a tool automatically suppressing testing.
A related capability we are working on is to also rerun the identified flaky tests X times so they pass. This depends on the capabilities of the test runner, so it will work with specific ones first (cypress, pytest, etc). That way you still make sure that flaky tests pass instead of supressing.
This is tricky to implement, for several reasons.
Log output is normally timestamped, making every line unique. Those parts of log lines would need to be ignored when comparing between runs.
Log output ordering is often indeterminate, particularly when a test has multiple threads, or interacts with an external service. Often the order of events logged is an essential feature of the difference between a successful and failed run. But some or most order differences are just incidental. The number of logged events may vary incidentally, or significantly. Explaining all these differences in detail to the test system would be too hard. So, the system needs to discover as much as possible of this for itself, and represent these discoveries symbolically. Then, allow a test to be annotated to override default judgments about the diagnostic significance of these features.
For people that prefer to minimize the number of moving parts, CircleCI now also has built-in Flaky test detection:
https://circleci.com/blog/introducing-test-insights-with-fla...
Of course this is coupled to our CircleCI platform, so if you want to stay agnostic for sure check out Flakybot :)
NB I work for CircleCI
- Network connections failed
- Run out of memory
- Duration timeouts
- Setting up the test infrastructure failed
All these can contribute to flakiness, and could be detected and reported by a bot like this. Making this a useful idea!
"Your test failed, but it was consider a flakey fail [Network Flakiness]" - FlakyBot
This is by no means an exhaustive list, but our goal with FlakyBot is to get better at identifying root causes as we identify flakiness across the systems.