They won't be fixed until they start actually preventing commits. If somebody deletes a test, that is on that person. I don't want a tool automatically suppressing testing.
They won't be fixed until they start actually preventing commits. If somebody deletes a test, that is on that person. I don't want a tool automatically suppressing testing.
This is tricky to implement, for several reasons.
Log output is normally timestamped, making every line unique. Those parts of log lines would need to be ignored when comparing between runs.
Log output ordering is often indeterminate, particularly when a test has multiple threads, or interacts with an external service. Often the order of events logged is an essential feature of the difference between a successful and failed run. But some or most order differences are just incidental. The number of logged events may vary incidentally, or significantly. Explaining all these differences in detail to the test system would be too hard. So, the system needs to discover as much as possible of this for itself, and represent these discoveries symbolically. Then, allow a test to be annotated to override default judgments about the diagnostic significance of these features.
A related capability we are working on is to also rerun the identified flaky tests X times so they pass. This depends on the capabilities of the test runner, so it will work with specific ones first (cypress, pytest, etc). That way you still make sure that flaky tests pass instead of supressing.