Waiting on Tests
etodd.io
etodd.io
Such a juicy premise, let's dig in. In my opinion the primary "speed up" to development comes in allowing changes - especially sweeping refactors - to happen without worry. As a corollary, they provide "primary source documentation" on the inner workings of the code so that less-experienced developers can approach such tasks without worry. The value of these are exactly the same regardless if your tests take 10 seconds or 10 minutes to run. The only difference is that we've sped up our critical path development work by not focusing on hyper-optimizing tests. And to start out the gate with such a premise but then fail to provide any data related to the effects on developer cadence, in a data-heavy writeup!?
I couldn't help but notice you run a "most common" test suite, implying that you have more tests that don't get run for whatever reason (the change doesn't affect that code, it's too slow, whatever). We end up having to do something similar to a small degree, and it bugs me.
What I would really like to see (and this might become a personal project at some point) is a way to optimize the tests at the step level. I believe this would accomplish several goals:
- It would reduce redundancy, leading to faster execution time and easier maintenance
- It would make the test scenarios easier to see and evaluate from a high level
Inevitably when you write tests you end up covering scenario X, then months or years later when working to cover scenario Y you accidentally create redundant coverage for scenario X. This waste continues to build over time. I think this could be improved if I could break the tests up into chunks (steps) that could be composed to cover multiple scenarios. Then a tool could analyze those chunks to remove redundancy or even make suggestions for optimization. And it could outline and communicate coverage more clearly. If formatted properly it might serve not only as a regression test suite definition, but also as documentation of current behavior. (Think Cucumber, but in reverse.)
And there is no tool that looks for that really.
He was simultaneously one of the smartest and dumbest people I've ever worker with. During that time, I did the best work I've ever done, and got the worst review I've ever had.
- Ability of other developers to be productive on the project. Having tests tells other developers the intended behaviour and notifies them when they have broken it.
- Flow on effect of slow or large test suites. It demotivates developers to write new tests, run existing tests and accept failures if there are too many. This then leads to a lack of confidence, trust deteriorates between team members and development pushes towards going faster than being stable. Eventually, once enough bugs occur or a large incident, then you go back to looking at tests again.
If you have fast and reliable test suites, a developer wants to run and add to them. Developers feel that this is a high quality project that needs to be well maintained. This culture them permeates into other areas of your business.
I'd be curious to hear about the manpower and logistical requirements behind this post. How many LoC was the original codebase? What's the current coverage at 59k? How many people were involved? How did they convince upper management to suddenly allocate this much opex spend for 5 years running, or what were they doing before this increase in testing expenditure?
That's a great starting point for refactoring or just normal feature development. You need the confidence that your changes won't break the system in unforeseen ways.
If you have a giant bunch of code that has been written without testability in mind, adding tons of tests mostly just locks down the interfaces of every single segment of code, making refactoring really hard to a point where the tests are useless because they have to be rewritten as well anyways.
Good test are written for good code that is written with testability in mind. In that case, only the relevant interfaces are proplery tested and not every little internal thing, which allows for proper refactoring and expansion without fiddling with the tests. This however also yields a lot fewer tests, which is why I doubt this large number of tests can be high quality and/or beneficial.
Some of us call those pinning tests. It's pretty on the nose for your complaints, since it encodes a sense of FOMO right into the name.
What, in your opinion, is the purpose of automated tests?
Since the set of intended behavior is only a small fraction of all actual behavior, if your test suite pins down all actual behavior, rather than just the intended behavior, you end up with a very brittle test suite. That is to say, a test failing doesn't tell you "this code is broken", it just tells you "this code changed".
I don’t actually know how the article approached this but again in my own experience, when code is written without testing in mind from the beginning and you just go wild adding them later you instead get tests that enforce very specific coding styles, tests that enforce that the dependency graph is immutable, that internal data formats must never change, etc. If you can add tests without intimately understanding the business reason the code exists you’re more likely writing tests that just lock in implementation. This type of testing also means that every imaginable change will break dozens of tests and puts people in the mindset of just updating the test to pass instead of considering it a real problem because the tests stop giving meaningful information when they fail because literally everything you do will cause a slew of failures. That’s helpful in knowing the blast radius of your change but it multiplies the cost of change instead of reducing it.
Here's the kind of thing I mean: https://github.com/simonw/datasette/blob/45b88f2056e0a4da204... - pytest displays that as 5 passing tests.
I dont think it's a very honest way of putting it though. It'd be more accurate to say that youve got one test and youve parameterized it.
If you want to optimize the create/terminate case, perhaps also create periodic snapshots of test instances after a run to populate the caches (or use something like Packer)
The vast majority of build failures will not happen on the last test, and they also probably won't happen in the slowest set.
Running all of your tests in parallel to try to honor the responsiveness requirements of CI, is better than doing nothing. But consider that the information contained in a red build is much higher than the data contained in a green build. A build that goes red in 90 seconds is better than one that goes red in 10 minutes. And red in 90 is enormously more information than green in 10.