Testing in the Twenties
tbray.org
tbray.org
I've been writing software professionally since '98, so I'm a little greener than @tbray. During my career, I've almost entirely been focused on building UI's, though I have done quite a lot of full-stack and embedded work.
I got infected by testing in ~2003-ish when I started reading Kent Beck and Martin Fowler.
FWIW - I had a 4 year stint at Google (10 years ago), where I led the team that built YouTube on TV.
Our product on the TV team certainly had it's own set of challenges, but quality wasn't one of them. By the time I left, we had over 2,400 unit/component tests that would run in ~30 seconds. This single application ran on many hundreds of SKUs (PS3, PS4, XBox 360, XBox One, Wii U, Smart-ish TV's, Set to boxes, Over the top boxes, Cable boxes, Roku, Chromecast, Blue Ray Disk Players, crazy, crazy stuff).
A given engineer could set up a file system watcher, that would run a much faster subset of tests in under a second. We deployed this application to production every week (which was revolutionary back then), and one of the years I was there, we had a single P0 production incident for the entire year.
We changed out the underlying UI framework 3 times over 3 years. We did this incrementally with side-by-side running frameworks and while delivering to production every week.
It would have been utterly impossible for us to accomplish this work without the testing infrastructure that we had. And by testing infrastructure, I mean:
1) Karma (or Mocha or Jasmine) for unit / component tests
2) Manual testing on problematic machines where automation was prohibitively expensive
The OP-linked twitter graphics are also triggering me.
Every time I've seen "Integration Tests" in a production environment, they've had the following characteristics:
* Slow: 10-40 minutes
* Brittle: Break mysteriously while working (if you can even run them from your workstation)
* Time-consuming: These tests chew a lot of time to author, and much, much more time to maintain
* Flaky: This is the deal-killer for me. They fail randomly in CI, and chew hours of the team's time trying to troubleshoot. You'll know you're here when you see random timeouts in your tests.
The main insight I came here to share is this:
There's nothing magical about UI. It's just software. It's probably (IMO) one of the more complicated areas of software engineering, but that's because of what UI is generally made of:
UI tool kits usually represent the user interface as a long-lived, mutable, wide and deep tree of interdependent state with a multitude of message passing schemes and needless hooks to global references everywhere.
Despite this, it's almost always possible to:
1) Disconnect and/or minimize the hooks to global state
2) Clarify message passing rules/practices
3) Break down the tree into manageable chunks
Once these ideas are in place, we can incrementally test UI as it's being built.
As an example for why we don't need "Integration Tests", I've often written "component" tests that instantiate a relatively complex node of my UI tree, poke the data, and verify some substructure was updated. This can and should be done without random timeouts and global nonsense.
We may not always be able to test our root node or main loop, but jeebus people, make those two things tiny and test everything else.
It's not magic, and it's super valuable.
It does admittedly get me worked up when people notice that UI is rarely tested, and then make claims that it probably shouldn't be.
It's not even mainly about the maintenance. Writing tests first(ish?) dramatically improves the design of my otherwise, much less clean code. Testing as part of development is an extremely powerful design tool that has ancillary benefits for maintenance and reliability.
As the OP said, these tests have to run incredibly fast (< 1 second at author time) and additionally, the unit under test cannot have a bunch of tentacles reaching out to manipulate global state.
Yes, almost all UI tool kits as they exist today are horrible to work with, but that's no excuse to just abandon the primary practice that almost makes our work bearable.
[Update: formatting]
1. Understanding different kinds of test doubles is not a religion. It's actually more like science. There's a taxonomy of different critters that behave differently. In general, names for anything test-related are not standardized. Maybe Bray is confused with Babel?
2. Kent Beck (highest, most exalted pope of TDD) does not even use TDD religiously. Nor does he recommend that. He just recommends using it unless you have a reason not to. It's a default that forces thinking. I'm too lazy to dig up links, but that was underscored in his many-hour "debate" with a petulant DHH.
3. If Bray is suggesting adding a test-only method to production code, then no, please don't. It bloats code and it adds attack surface. And he can treat that as religion if he likes.
Powerful type systems and compilers that catch your mistakes before you can run anything is the best kind of testing. It’s fully automated and will inform you about most things early on. Practically, this kills dynamic languages for projects expected to become larger than proof of concept (pydantic, typescript, etc. mitigate, but must be in and enforced from the start).
After that manual testing is the most important. Specifically because it can be done by hand with isolated inputs and outputs. This won’t catch the refactoring breaking something disparate, but will ensure you can reproduce production issues with the code locally where you can expect it.
After that, integration tests are the most important. Give the whole component/system an input, and verify the output. This should be as fast as possible but should mock as little as possible. This should be able to be done from a local machine with a debugger running and still be fast. If an integration test takes half an hour, it’s never going to get run except on a Thursday at 6pm when you’re trying to get out a change before the Sprint review. Then it will fail and make life hell, because running it is so slow that running with a debugger is going to be hell.
After that, logging is the most important form of testing. Every input or start of a workflow should get an ID that follows it through the system with logging to watch it go. Any problem identified should include that Id so that you can grab logs from all the disparate systems. Specifically, that ID should be attached to the inputs or each component so you can replay it in a controlled environment later.
Then you finally get to unit tests, because once your software works as expected, they let you lock down the implementation and ensure that any changes will get flagged by the tests. If you lock down the unit tests before the software works, then you wind up having to go back and redo the unit tests for every issue found, as the implementation wasn’t correct so the unit tests can’t test it. This is the big lie of unit testing. They don’t prove correctness, they prove implementation details.
After that you want to get to end-to-end that workflow the whole user experience with a product or system. But if you don’t get there, that’s really okay.
Most problems are caught by the compiler, manual tests, and integration tests during development. Any problems that do come up can be identified quickly and repaired by the highly inspectable infrastructure that provides you IDs to dig out logs for any input, and while you’re fixing the issue, your unit tests will let you know when the implementation has changed in a way that requires more thought (though the compiler probably did first and the unit test complaints are just noise).
Then I finally got an opportunity to work in a handful of those environments, of course the first was everyone's favorite whipping boy, Java.
Yikes!
Maybe I've just been unlucky or unsmart, but wow, I've seen (and even made) some really impressive messes with these languages (specifically: Java, ActionScript 3.0, Closure (.js), Typescript, Go, C++ and C).
In my experience, especially in the context of UI development, the static type systems I've encountered generally make the task of unit testing more difficult and time consuming, while adding little to reliability.
At the end of the day, I just don't bump into major type mismatches in dynamic languages that cause a whole lot of hearteache.
All that said, Go was probably the best static type system I've worked with until I (recently) discovered Zig, which has a really interesting type system (esp. related to comptime) and has testing built into the language itself.
While it might at first look like it, I'm not saying static type systems don't have a home in my heart.
I'm just saying that in my experience (just like TDD and Unit Testing), they haven't led to the bug-free panacea that proponents tend to imply.
I do find it endlessly fascinating that 2 people can walk through similar-ish challenges and arrive at entirely opposite conclusions about how to best tackle them. Thanks!