I'd be curious to hear about the manpower and logistical requirements behind this post. How many LoC was the original codebase? What's the current coverage at 59k? How many people were involved? How did they convince upper management to suddenly allocate this much opex spend for 5 years running, or what were they doing before this increase in testing expenditure?
That's a great starting point for refactoring or just normal feature development. You need the confidence that your changes won't break the system in unforeseen ways.
If you have a giant bunch of code that has been written without testability in mind, adding tons of tests mostly just locks down the interfaces of every single segment of code, making refactoring really hard to a point where the tests are useless because they have to be rewritten as well anyways.
Good test are written for good code that is written with testability in mind. In that case, only the relevant interfaces are proplery tested and not every little internal thing, which allows for proper refactoring and expansion without fiddling with the tests. This however also yields a lot fewer tests, which is why I doubt this large number of tests can be high quality and/or beneficial.
Some of us call those pinning tests. It's pretty on the nose for your complaints, since it encodes a sense of FOMO right into the name.
What, in your opinion, is the purpose of automated tests?
Since the set of intended behavior is only a small fraction of all actual behavior, if your test suite pins down all actual behavior, rather than just the intended behavior, you end up with a very brittle test suite. That is to say, a test failing doesn't tell you "this code is broken", it just tells you "this code changed".
I don’t actually know how the article approached this but again in my own experience, when code is written without testing in mind from the beginning and you just go wild adding them later you instead get tests that enforce very specific coding styles, tests that enforce that the dependency graph is immutable, that internal data formats must never change, etc. If you can add tests without intimately understanding the business reason the code exists you’re more likely writing tests that just lock in implementation. This type of testing also means that every imaginable change will break dozens of tests and puts people in the mindset of just updating the test to pass instead of considering it a real problem because the tests stop giving meaningful information when they fail because literally everything you do will cause a slew of failures. That’s helpful in knowing the blast radius of your change but it multiplies the cost of change instead of reducing it.
Here's the kind of thing I mean: https://github.com/simonw/datasette/blob/45b88f2056e0a4da204... - pytest displays that as 5 passing tests.
I dont think it's a very honest way of putting it though. It'd be more accurate to say that youve got one test and youve parameterized it.