Why TDD isn't crap
hillelwayne.com
hillelwayne.com
Unlike the author, I absolutely believe that tests are about design. More specifically, they're about identifying coupling so that you can reduce it. The function being "awkward" to use is part of it, but code which is hard to test is almost always going to be hard to maintain.
If you do this for enough time, you should naturally start to write code that is less and less tightly coupled. At that point, the value of tests as a design canary decreases, but never completely goes away.
This is all broad generalizations. Individual vigilance and giving-a-fuck matter more than anything else. But if you show me some random code, and it's spaghetti, I'll bet every single time that the author doesn't test.
By that logic, if you write testable, but untested code, then you're still making it difficult for people to refactor later. This applies even if the code is well designed and uncoupled.
Isn't that what good git commit messages are for though? (Or any other source code versioning tool.)
I'm not saying tests aren't useful, but often times I find reading a file diff history is way more useful to understanding it than its tests.
IMO, tests should be there only to provide some level of feeling safe modifying files, as you can easily know if you messed up something.
Tests would also do this for you, but without the mental burden of brain compiling. It's nice to be able to set some breakpoints in code, then start a quick debugging session with the relevant test and see the data flow through the code - you can understand any function usually within 5 minutes.
This is probably one of the easiest things to quantify. I just finished working on a project that used BDD and TDD. The project was a rewrite of an older system that required manual testing. It took about 6 months to get a change into production. After the new system went live, they found a simple way to improve it. The code change was very small, but because all of the TDD and BDD efforts could be leveraged as regression tests, the change was deployed in a matter of weeks. That single change was able to capture an additional $5 millions in value that would have been lost if things had of gone through the previous 6-month cycle for testing the entire application manually.
That is exactly the point why I don't like TDD.
I mean, there are enough constraints that shape the code, why should something artificial like tests shape it too?
With mocking and everything you end up writing code for tests and not code for your problems.
For example, instead of creating one class that does the network call, and processes the JSON response, break these out.
One class does the network call (keep it extremely thin). Dependency inject stuff in if you want. But view this purely as an orchestration.
if (connectionValid) json = makeNetworkCall anotherObject.processJson(json) <- test separately end
Then, in the other plain old data process class, do you heavy off-by-one edge case testing there.
If the first case, the orchestration, I can write tests like "One make the network call if the connection is valid". But that's all that test would do.
I am at the point now where I actually don't even test that. I assume that is going to work, because it is so simple. That and I generally hate mocking. Way too much coupling. Way too hard to refactor.
Where I do test heavy is on the plain old object side, that has no outside ports or connectors. These are just plain old objects.
I agree with some other comments I have ready here. Mocking is a bitch and I generally don't like it. But it's there if I needed it, so I do use it occasionally.
But it has so many downsides I try not to. Instead I separate the orchestrations from the processing. Unit test the processing (in an integration test sort of way), and try to stick to testing the public APIs of my class, so I am free to change the internals without things breaking.
My 2c. Hope that helps.
> ...not code for your problems...
Right. It's not your problem now. It may be your (or your successors') problem in the future.
Is it always worth the effort? Nope. But experience, communication, and good teamwork will help with balancing short-term and long-term goals.
I'm saying, with experience, teamwork, communication (including through tests), you'll know when this proposition makes sense or not. It's not universally true that "we'll worry about it later" make sense.
Tests shaping your code is bad, because? Because you already have "enough constraints"? Sorry, I don't get it.
- One test to prove I can cleanly load the component or object in a module. This is about loose coupling and ensuring that each module contains only one object.
- One test to prove the component or object does what I what it says it does. This is primarily about the Single Responsibility Principle. If I need six tests to prove an object works, then the object might be doing too much and might need to be refactored.
Other tests can be added if bugs appear, but just two tests per object can help us keep things light. Of course, if I have a module full of utility functions, then that is a different matter. But I think the limiting the amount of code in my tests and the overall number of tests is valuable.
If you write the code, and then you write the test, then you haven't tested your test. You have no proof that your test will detect broken code. The only way to prove that your test can detect broken code is to run it against broken code. So you can write the code, then write the test, and then break the code and run the test. Or you could just write the test first, and run it. It should fail. Surprisingly, sometimes it doesn't: either because the test was bad, or the existing code didn't work like you thought, or sometimes the language doesn't behave like you think.
I've dealt with hundreds of thousands of lines of tests where I could go into the code and just start deleting functionality en masse, and no tests failed. If you don't test your tests, they're just bloat. They're the kind of pointless bloat people are complaining about in these unit testing threads.
For example,
TestSort(ISorter sorter)
{
var input = { 1, 2, 3, 4 };
var output = sorter( input );
Assert.IsSorted( output );
}
This test will fail if you run it before the ISorter is implemented. But then it will pass if you implement an empty ISorter because the test still has problems (ie it starts with sorted input).Cue the ever familiar: You still have to use your brain when using TDD.
If TDD needs brain usage to function, maybe the real thing that is useful is the brain usage and TDD is vestigial.
Forget that you, a human, thinks you know what an ISorter should do. What do the tests demand that you do?
Doing TDD, the goal, when coding, is to write the minimal amount of broken code that makes the test pass. If you passed me {4,3,2,1}, then my code would be "return {1,2,3,4};". If you wrote a second test that passed me {6,5,4,3}, my code would be "return (input[0]==4)?{1,2,3,4}:{3,4,5,6};". You've got to write a test that makes me actually code what you want. I use this adversarial approach even when I'm writing both the test and the code.
Usually the first test I write if I'm starting with a blank slate, is to pass null and check that I get a NullPointerException. That gets the class created. After that, its got to be randomly generated data.
What happens if you're implementing something like a unification algorithm, but you don't know about the occurs check. You'll progressively add test cases that break your code and eventually stop. Until an end user creates a query that contains itself and your algorithm fails to terminate.
The solution to get TDD to work is to know what all of the weird edge cases of your algorithm are. However, I assert that the thing that makes TDD ever work is that the practitioners who are "doing it right" already know all of the weird edge cases of their algorithms. And I assert that if you know the weird edge cases you can drop the test first part of TDD and still get working algorithms.
Random data is interesting because it's not something I've seen anyone mention when they talk about how they do TDD before now. I suppose if you knew nothing about your problem and had some sort of tool progressively feed you intelligently generated random data, that might work for algorithms that have few edge cases or relatively simple edge cases. However, it seems like you would want something more like quick check for cases where you have really esoteric edge cases. And I don't see how you avoid implementing bubble sort (or its equivalent in your domain).
I don't think TDD can't work, but so far nobody has described TDD to me where I have any reason to assume that TDD is actually doing anything worthwhile. It's always sounded like TDD was irrelevant and the real key to success was thinking about the problem carefully. Additionally, it sounds like there are problem domains where anything that TDD might bring to the table would be nullified (ie certain problems that exhibit a certain level of complexity).
So you're in a local maximum, and to get to the next, higher, maximum, you're going to have to go through a trough.
That being said, if anyone wants to fund me for six months (plus either some compensation to my employer for having me gone for six months OR a significant bonus for my own financial security) to try TDD, then I'll gladly do it.
That being said, I plan on putting it through one heck of a trail. At the end I'll either have a pretty good idea of exactly why TDD works OR I'll have a compelling argument for why it's gains are illusory.
As long as one maintains decent version control discipline, you can be pretty confident that the test is valid with this technique. (It does get a bit more difficult when there are schema changes, but then it also encourages good discipline and vigilance when you're making schema changes ;)
In my experience coaching teams on TDD, its quite easy, when writing code first, to add more code than is needed to solve the problem. And then only the problem is tested for. Meaning that there is code that is not tested. Maybe an if statement where a branch is never taken. Strict TDD means you don't write any code without a failing test, and you write the minimal code needed to make it pass.
Now quite often, a programmer has no idea how to tackle a problem, so I advise that they just write something that works. Just have a go and get something working. Most programmers (if they test at all) will then just go and write a test for it, usually for happy path. And that usually results in code with bugs. I know this because many times the programmer has done as I suggested and commented out their code, and reimplemented from scratch using TDD, and doing TDD finds a bug in the original that nobody saw.
It really comes down to the nature of the problem - if I foresee a chance that the final architecture is not immediately obvious, I'll start breaking down the problem into small pieces and TDD the units, then the integrations.
I guess I'm of the opinion that not everything needs TDD, but everything does need diligence - my code after doing years of TDD has improved drastically, whether or not it was strictly test driven. Merely considering whether it will be easy to reliably test the code improves it, IME.
This is a pointless equivalence to make. You're trying to suggest that code which is hard to test could be written differently, and better. This is not always true, and therefore useless to say.
STUXNET was very difficult to test. Reports are that it required an entire Israeli nuclear facility as a test environment. Would you suggest that STUXNET was written to a substandard level of professionalism? I doubt it.
He literally admits it's not always true in the sentence
But, I dunno, I haven't looked at the source code. It might be very difficult to maintain.
1) ... over-emphasise the importance of reduced coupling, and can actually increase the chances of integration failures. Why? Because components are tested in isolation rather than together, but are still considered "tested" - especially when using stubs and mocks. I've seldom seen mocking used well, it almost always over-specifies implementations.
2) ... increase the tested "surface area" of the code, making it much more costly to modify in a way that changes the "surface area". When two bits that normally only ever face one another are tested on both sides, you now have 4 places to update when you move functionality between them instead of 2. And that's just for a small refactoring.
3) ... encourage a spurious modularity that increases the overall conceptual complexity of the solution. When every dependency needs to be replaced for testing, they become parameters, one way or another; and thus the code becomes over-generalized and over-abstracted, more removed from the work it's doing and more concerned with the bureaucracy of communication and coordination. It's this spurious modularity that's the cause of Enterprise Java FactoryFactories and friends.
That sounds quite negative on unit tests, and actually I'm not negative at all. I think code can be modified in more than one way to accommodate tests; poorly, by introducing single-implementation interfaces everywhere; or better, by reducing the number of types of interfaces, converting control flow into data flow, and generally adopting more functional and higher-order compositional architectural patterns.
It's hard to get more specific without talking about examples, though, so I'll stop here.
After all, a lot of effort had gone into building all of the internal mocks, and while many of the early design decisions enshrined in the tests were quite dubious, to revisit those assumptions would have been backsliding. Even though there were no users to speak of yet, the existing tests had to pass.
And of course because of (1), the product didn't actually work because there were no integration tests at all. Because all of the unit tests passed, the ongoing assumption was that the platform was basically sound.
The most charitable thing I can say about TDD given that context is that it might be valuable, but isn't sufficient.
If you agree tests are a good idea, but think TDD is too extreme, consider that TDD simply makes sure you write testable code from the outset. When you have a test wrapping a method, and need to add a dependency, you actually decide to use DI (Dependency Injection) because otherwise your current tests will break / become integration tests. TDD makes you think upfront about about things like mocking, separation of concerns, etc.
When you have the mindset that you will absolutely 100% write tests at some point anyway, TDD is actually a faster and more fun way to develop than bolting on tests afterwards.
Whats DI?
Regardless if your code works your tests could still fail, likewise your code could not work and your tests pass so since the tests and code may vary independently if either is buggy I don't see how this can possibly be true.
The point I’m trying to make is that the verification of correctness is bidirectional. If I’m writing tests around existing code, I rely on the code to test my expectations. If I’m writing new code, I write tests to assert my new assumptions. All automated tests do are assert that I’ve written the same logic twice. Writing it a third time will also decrease the likelihood of transcription error again, but at diminished returns.
But consider this: what if your tests are poorly written and fail to detect bugs? What if they fail because the test code is buggy? Failing to detect bugs is a bug (undesirable behavior/output) of test code. There are some techniques to address this, for example mutation testing ( https://en.wikipedia.org/wiki/Mutation_testing ), which effectively become some part of "the tests of the tests".
The take away: no, the code is not enough. You need to test your tests.
Coverage tests how much of your code is exercised by your tests. Mutation testing is (one way) of testing how effective your tests are -- i.e. how sensitive they are to bugs in your code. It follows that good tests must have both good coverage and good effectiveness, but the two are not the same.
Of course they're not literally the same, but both are ways of measuring how good your tests are at catching bugs in your code.
It might sound pedantic, but this way of thinking shows you that it's not turtles all the way down: if they'd really be "tests of the tests", then you might also want "tests of the tests of the tests". And since we obviously don't want to keep doing that, you might conclude that we don't need "tests of our tests" as well.
But since they're simply measurements of test quality, we can see that both test coverage and mutation testing can be useful and not a never-ending story.
- Coverage measures how much of your code is exercised by the tests. It does NOT measure test effectiveness (it's not like mutation, only "less/more brittle"). This can be trivially shown by writing tests that exercise all of your code but have no asserts (or trivial asserts such as 1 == 1). Unsurprisingly, this kind of obviously ineffective test code is often find in the wild (mostly written by junior devs), but less obvious cases of ineffective tests are also common, such as failing to test border conditions. This happens and it's more common than we'd like.
- Mutation testing is a way (but not the only one!) to measure how effective your tests are at actually finding bugs. That this technique exists shows that there is indeed a need to "test the tests", i.e. measure test quality & effectiveness. Production code is NOT "the test of the tests", as someone up this thread erroneously said.
This is like measuring altitude and airspeed: yes, they both measure something useful about your airplane, but they measure different things!
I confess I did not understand the rest of your post.
Yes, it can be inaccurate, as your trivial example shows, but it does guide you towards unhandled test cases (i.e. where your tests are ineffective). Likewise, mutation testing guides you towards unhandled test cases (including ones that coverage can't detect).
The problem is that coverage alone is a very poor way of measuring effectiveness, which is why other techniques -- including, but not limited to, mutation testing -- are needed. This is nothing new: limitations of coverage are well known in software engineering.
This seems like tautological reasoning. Naturally, if you write bad tests then they won't be effective, but that seems like a poor argument against testing. That's like saying "Yes, bypass surgery could save your life, but consider this: what if the surgeon is reckless and kills you on the operating table?". Obviously, there is a minimum expectation that the surgeon will follow proper medical procedures when operating.
> Naturally, if you write bad tests then they won't be effective
But how can you tell if your tests are effective, i.e. if you wrote "good tests"? Production code alone is not enough; you cannot tell if a test is green because everything is ok or because it's buggy/incomplete. You must introduce some degree of quality control in your test code.
> what if the surgeon is reckless and kills you on the operating table?
Excellent question! What's this surgeon's track record? How many patients have died on her operating table? How up to date with current medical research is she? Is she daring and reckless, or does she play it conservatively safe?
Your points are directly addressed in the pdf. One of his general points being, in practice, the tests become the legacy system instead. And I'd add to that. Given that you've at minimum doubled the code (and doubled the bugs), it seems like a really bad long-term trade off.
Also DI does not reduce coupling. I've seen plenty of code with DI that's just injecting like 30 things, which is obviously therefore coupled. It just makes it really obvious, but DI itself has massive downsides.
If you've ever worked with bad programmers and seen it in the wild now, I'm sure we can agree DI and TDD doesn't stop bad programmers writing bad code. In fact, all it seems to do is make even more of a mess.
Not only do you have to pick apart the bad code, you have to start dealing with carefully moving methods to the right places because DI can make it hard to figure out what's being used where, and then on top of that tests break all over the place because they're entirely dependant on the implementation instead of the functionality.
* and to reply to your edit: of course tests break when you change the code they are testing! But after reviewing the broken tests, you see the intent and re-adjust the test. But what often happens, is you realise you didn't fully understand the code previously, and actually after reading the test you need to undo your refactoring as it didn't make sense in the first place.
As for throwing away tests; Unit tests are meant to be pretty simple - rule of thumb is you can run a thousand tests in ten seconds. Arrange, Act, Assert - they don't need to be complex, they just need to imprint the intent into the codebase. If the intent changes, by all means remove the test.
Not necessarily. If the rewrite is only addressing issues that would have been prevented with tests, then sure, this is clearly an argument for TDD. However, if the rewrite is going beyond what could have been provided by tests, then having a large testing system could actually make the rewrites harder, which means it's an argument against TDD.
It's hard to say which is the case without knowing the details, and I'd certainly err on the side of saying a yearly rewrite is not a good sign. But if a company is undergoing rapid growth, it's not that unusual for certain systems to be rewritten frequently as fundamental new insights are gathered about how to tackle problems that are hard to scale. And if the system isn't that large to begin with, periodic rewrites could be easier than writing a single version that's supposed to last 10+ years as business requirements dramatically change and expand.
1. Wow it made 10k! Lets rewrite
2. Wow it made 100k! Lets rewrite
3. Wow it made 1M! Lets rewrite but let's properly think hard about the future of this product (Introduce tests and more people).
ALWAYS DO TESTING should always be given the context (if budget allows).
In saying that, it was raw and I think I enjoyed my job more. There is no safety harness. You write decent working code and sign your name by it or you run home to your mothers nipple.
I definitely learned a lot about programming during that time. To take a terrible system and try to research ways to make it better for your own mental health, not the companies is a personal drive and golden age of discovery which plays a part of my programming to this day.
I get none of this from TTD. It is intensely dull and unsatisfying but gets the job done.
I'm effectively the tech lead for two startups, one getting about 750k users per year, the other 250k. Freelancer, but I do a lot of time for 2 clients at the moment.
Both were projects started by other people, but a significant chunk of the code is mine now (one of them I've virtually totally rewritten from VB.Net to C#) and the other I've done huge amounts of refactoring to fix big performance problems, reducing the "main" pages from 10 sec load times for complicated orders to 250ms.
I've also refactored a lot of javascript for both of them without any tests, significantly improving client-side compile times and page load times.
I played with unit tests a few years ago, one of my clients has some, but we mainly don't add any new ones and they never catch anything[1].
[1]That's a slight fib, one caught a bug last month for the first time in the 1.5 years I've worked for this client. It would almost certainly have been caught in testing though.
This matches much of my experience with DI on real projects. Dependency injection is, unfortunately, used to reduce the labor required to achieve massive coupling. Of course the intended and potentially useful application of DI is being able to rewire an application with different components for different purposes or different environments, and the tradeoff for this flexibility is that you lose explicitness. You can no longer see the explicit wiring in code, which is a big downside. Yet I've worked on several DI-heavy projects with seasoned engineers who get this completely backwards. They see DI as a labor-saving device that allows them to write heavily coupled systems without having to explicitly work out the dependencies between things. In fact, they even see DI as eliminating the cost of complex interdependencies because it reduces the work and cleverness required to create them.
Partly I think this reflects a desire to work on "real" "enterprise" systems. Instead of fighting to keep complexity below the level where DI actually helps, they embrace the flattering thought of, hey, we're doing big boy work here — it's going to get complex, so we'd better use DI from the outset. People who take a lot of pride in working on big systems can't help creating them.
I still stand by my point. 100 lines of code with 1000 lines of testing code means your codebase is 10x what it would be without the tests. That extra 10x had better result in a hell of a lot of benefit.
Additionally, over the years I've found that easily-tested code sacrifices readability, and that tiny functions isn't always the best way to go. I've gone back to liking larger, more monolithic functions as often I don't want to create a bunch of generic functions for hypothetical use for other projects, I just want some code that fits my needs to a T.
Adding more code != more complexity. Complexity comes from high coupling between code, that doesn't separate concerns, which make it difficult to untangle, aka. spaghetti. Unit tests are mean to be simple, and anything they test will be no more complicated that how the system will actually be used. In this way they clarify the code, by encoding the intent, available for all to see for years to come.
> I've gone back to liking larger, more monolithic functions as often I don't want to create a bunch of generic functions for hypothetical use for other projects, I just want some code that fits my needs to a T.
This isn't what TDD causes, it just makes you break down things into small testable units. It doesn't make you create super generic wrappers, which I also agree are useless because of their un-readability.
Err, not quite. We're running on computers, remember. A comprehensive unit test that doesn't go through every possible integer input should still include:
maximum and minimum integer values for that size
maximum and minimum expected inputs
0, 1, -1
So, there's 7 test cases for a single integral input. And since you bring up banking, let's imagine you're working with a function with two inputs. Since most bugs come about because of the interaction between two variables [0], you'd want to check out each combination. So, 7 possible inputs for the first integer, 7 for the second: 49 test cases.
Why bother? You're working for a bank... imagine facing a client whose balance shot negative because of an overflow error and started accruing "overdraft coverage" fees.
Back to the original point: when was the last time anybody but the AFL tool wrote the "proper" 49+ unit tests for an "add_to_balance(int, int)" style function, when one test would give you 100% coverage for that function?
[0] https://csrc.nist.gov/Projects/Automated-Combinatorial-Testi...
If you mean, "no look seriously, you'd need 49 tests for every method because you're a bank", well then, no, you actually need less than 49 (since some are subsets of the others), and in reality even this would get very tiring so you'd write a currency class that can't overflow, and yes, being a bank we would test every case and certainly use automated test generation while we're at it.
But if you mean that in general, for non-banks, the classes of values is too large to manually write unit tests for, then we disagree. That means that your method is too complex.
In a purely mathematical sense, yes, this is true. But with computers those subsets matter. For example, I wouldn't drop the "max expected" for the "maxint" test; an overflow exception or implicit type change in the case of the latter shouldn't be tolerated for the former.
> That means that your method is too complex.
Would you consider the Python `requests` library's `get` method to be too complex? It has 10 parameters to it: 4 strings, 4 dictionaries, 1 file-like object, and 1 tuple/AuthHandler option.
If you conservatively say there are 5 test cases which should be tested for each parameter (a very low count, especially with strings and dicts), there's approximately 1,125 test cases that should be explored.
Not quite a trillion, but it's still a lot.
I coach my teams to think of two kinds of methods: methods that do data processing, and methods that coordinate. To test a coordination method, you don't need to care about the value of the things it has to deal with. Maybe you care the they are not null, if your language has nulls. But other than that, what you are testing is that if your method A is supposed to pass arguments 1 thru 5 to method B, and 4 thru 10 to method C, then that's what you test. You don't need to test that argument 1 is a positive integer. The test for method B will do that, if it needs to. And finally you might test that if we expect the result of method B to be passed to method C, the test for our method A would do that using a mock B. Again, what mock B returned could be any object: we only care that it got passed to C.
So it looks like its 1,125, but its actually three.
If we have classes that are SOLID, then we tend to see the processing methods and the coordination methods end up in different classes. sessions.py isn't a great example of SOLID, and I would have a hard time writing tests for it. It is certainly not the kind of code you get if it was written with TDD.
I'd say you've at maximum doubled the code. The test ensures you write only what you need to get the test pass. Without them, devs get distracted and wander until the feature works. Usually distracting themselves with tons of YAGNI violations along the way. In my experience, untested code bases have a ton of unnecessary code.
I don't understand how it would double the bugs. The article has references saying it reduces them. But, even thinking about it, I don't see why you'd say that.
I feel that TDD is an attempt to impose engineering discipline onto something which is still largely at the craftsman stage. As an engineer I like the idea of pushing the field forward, but I don't think software development tools are ready yet to make TDD and the like broadly accepted practice. That means you need to decide case-by-case if they make sense.
Tests are the programmer equivalent, they define a contract you expect from your functions, which can then help you shape the function itself.
There's a number of reasons why you could try TDD and get frustrated with it because it ends up sucking for you:
* You're using unit tests where integration tests would be more appropriate.
* Your integration tests would take longer to build than the thing you're building itself because you have to set up some sort of elaborate mocking.
* Time is of the essence but quality isn't / your code will be run a limited number of times.
I've just seen a lot of it where a product is still in broad strokes development and developers get stuck between whether having tests written based on early assumptions are correct and the code should conform to those, or whether new ways of thinking that invalidates the early assumptions and tests are the right way and the tests should be changed.
From the article "We don’t actually know that much about what good software engineering looks like." sums the issue up nicely. There is no definitive playbook on whether a strategy like this is good or bad. It's a tool that is good if you use it right.
I've seen and heard this a lot, and it's often the result of code that does too much (not cohesive, too much coupling) and tests that correspondingly do too much.
But if the entire population as a whole is still having problems with TDD for as long as TDD has been around, then it needs to be a niche methodology.
Our decisions to adopt methodologies/technologies can't revolve around the perfect case of a crack A-team of developers who are good at everything they do.
Think of the mediocre developers!
However, they're terrible if you want to refactor your design and correct problems with the interface itself. Tests freeze the interface. This is really good for whole classes of problems but terrible for other classes of problems.
The debate is going to continue endlessly as those on either side are looking at it from entirely different frames of reference.
"Why Most Unit Testing Is Waste"
https://news.ycombinator.com/item?id=15591190
(just for the record, i.e. for future readers)
https://www.youtube.com/watch?v=DQBf6li1hww
is a case in point. Presenter takes what he believes to be TDD's four main points, some of them strawmen, and mocks them. He does make some good points, but here's the problem: he offers no alternative.
If you're not writing tests as you go, that you run before every commit (or before you move onto the next thing), then what is your standard for putting code into a production repository?
I wrote my first big API with many tests, hundreds of them. Then the requirements changed and all of them failed. People worked for weeks to get the tests passing again. So just writing many tests up front in a new project doesn't seem to help anyone.
Also, I went from feature to bugfix sprints. First I implemented some features, then they got tested by non-devs, then I fixed all the bugs.
Often this was faster than the whole tests up front stuff with refinement of tests afterwards.
I also saw that >90% of my bugs came from the dynamic nature of my language of choice (JS). So I could imagine, that this feature->bugfix->feature->bugfix cycles could be greatly shortened (especially the bugfix parts) by using a typed language instead.
Using a static language does help with some bugs (where it effectively serves as a compile-time test suite), but it can also increase coupling a lot. It's not easy to quantify that tradeoff in abstract. I'd be wary of people overselling aggressive compiler checks leading to productivity boosts.
There are other limitations though, such as not being able to treat static types as a hashset at runtime, which can be good or bad thing depending on what camp you're in.
In one that has better support for generics (i.e., reification, contravariance and covariance) and some form of mixin, you generally shouldn't have that much trouble adding behaviors to an entire class of types without having to modify any of them. This is the fabled open/closed principle that is oft lauded and rarely practiced.
To take it a step further, if you're using interface polymorphism and decorators to build your types instead of relying on subclasses, you won't be able to paint yourself into the corner you describe in the first place. The problem is, of course, that a language like Java that doesn't let you add new behaviors for types without either modifying the original source file or resorting to some Gang of Four awfulness, will tend to punish people for writing cleanly-structured code like that.
I agree. In my experience, it's how the median developer writes code, though. Even in more flexible languages they'll reach for type hierarchies and abstract base classes. I've seen people create these things in languages like Lua and Javascript that don't really need them.
I think TDD-like approaches make the case for other approaches more clear, for what it's worth.
I haven't looked at a textbook on programming recently, but I'm worried that the standard is still to actively teach new developers to program this way, even though we've _known_ for decades that towering piles of subclasses invariably collapse under their own weight.
That said, I still wouldn't lay this tendency for damaged design at the feet of static typing in general. Not when some of the most vigorous arguments for static typing tend to come out of language communities that don't have subclassing in the first place (e.g., Haskell), and when (as you point out) similar mistakes are just as often made in dynamic languages. Dynamic languages are certainly more forgiving about poor design, but whether that's a good thing is yet another fun debate.
Assuming it's a modern static OO language, your business logic should depend on a User interface, so that it never has to take a dependency on implementation details like that. Even if User was a class beforehand, you can easily extract an interface at a later date, when you find that you need to avoid some tight coupling.
Don't blame the gun for what happens when you point it toward your foot and pull the trigger.
That said, testing the input/output directly doesn't work well because now a single change the output data (adding a field to the JSON for example) will call tons of tests to fail.
I usually use or create some JSON serialization/deserialization classes and use those to generate fake data and ensure that the API returns 200 responses.
That way if I add a field to the JSON objects, my tests are unaffected and will generate the new fields in the tests.
So I may do something like this (Python-ish):
new_user = User.generate_user()
response = api.post('/users', new_user.serialize())
assert response.status_code == 200
response = api.get('/users')
users_json = json.loads(response)
users = {}
for user in users_json:
users.add(User.unserialize(user))
assert new_user in users
That way changes to the User JSON will not impact tests and I still have some basic sanity checks.Hope that helps.
Do you have any good resources on testing? Books/articles etc.?
I often have the feeling the only people who write about this stuff are the die hard TDD gurus.
I did read Clean Code by Robert Martin. He's definitely in the "die hard TDD guru" category, but there are still useful examples to be extracted from the book.
Besides that, I mainly learned a lot of testing tricks in the wild, reading through Github projects. My learning process now-a-days is mostly:
1. Discover a really cool trick or concept I didn't know existed in a Github project. 2. Research the hell out of it. 3. Try to use it in some personal example project.
It seems to me that that would be the same standard that you would otherwise try to write a test for.
Now writing a test isn't necessarily a bad idea, but I think it's important to realise that the test isn't the standard itself, but is itself an implementation.
That's awesome!
Then it dawned on me that I was already testing, the hard way, by opening up a console and manually setting up test conditions over and over and over again, and that I could do this much faster and in a reusable way by writing tests and running them. What an epiphany that was!
I still open up the console all the time. It's a really useful thing. I think ideally, my testing environment would dump me into a full-fledged console if something went wrong, but this is not something I've taken the time to set up.
It's a design methodology, not some new way of unit testing. In fact, I think the more you think of TDD as being testing, the more you're probably missing the point.
Modern OO languages are full of hidden dependencies and perverse side effects. The only sane way to write clear and maintainable code is to write the spec first, that is, you code by writing tests, then writing the code to make the tests pass. In this manner your code is always up-to-date with your spec.
Where is it a bad idea? Exploratory or academic code, for one. Startup code where there's no clear benefit to maintainability or even knowledge of what the app is supposed to do.
Pure functional code is another case entirely. Lately, I've switched to writing small pure FP in microservices, usually with less than 200 lines of code. Writing code like this creates very simple and small pieces of functionality with little hidden state or adverse side effects and a limited cyclomatic complexity factor. I don't see any reason to use TDD here, because there's nothing happening that isn't obvious. (It's a horse of a different color with larger pure FP projects, however. Having said that, one should pay careful attention to whether or not you need to build out huge pure FP execution units in the first place)
Weirdly it seems to be the loudest advocates of TDD who are the most confused about this.
I've seen many outsourced teams say they're doing TDD and when you look at the code it's obvious it's just the same old unit testing as before. I have no idea how vendors get away with this. It's no less than fraud, really.
ADD: I think the danger here is that, even with hardcore TDD boosters, they don't understand why they're doing it. It's a discipline, not an engineering skill (Choosing the tests is the engineering skill). Over time they tend to get lax. After all, the code always does mostly what I wanted it to do, right? So I can look at it and by inspection reason through the execution.
At this point, when you don't understand the rationale behind it and you've started to slip-up in your application, TDD has become nothing but some weird way of writing unit tests. Then, sure, you can use the terms interchangeably. But then you've missed the entire point of what you're doing, so might as well just call it "unit test ahead of coding" or something.
However, I think I reached my sweet spot just weeks ago. Here is my optimum workflow now:
1) Write test cases (the one sentences that say what is expected, in plain English) 2) Implement the code 3) Implement the test code
The test cases become neatly arranged in bullet-like layout in a test case file. I'm able to read through and be confident that I'm probably covering most if not all the cases that need to be covered. I'm always able to switch between the code and this file to make sure my code covered all that the tests need.
Then once the code is done, I come back to the test cases to implement them one by one, catching error by error and seeing my code coming to life. As I code now, I know I have the implementation I'm quite confident of, my tests that have business value are being covered one by one. It's been pleasure since then and I'm more confident of my code and of my time spent efficiently.
The whole idea of testing functions and/or classes separately means tightly coupling your test code to the implementation of the real code, while you should only care about testing the functionality.
Nowadays, I try to write tests that test a unit of functionality. And the tests should only change when the functionality changes, not after every refactoring.
That said, code coverage is a metric with no inherit value (well, unless it's 0, of course)
> How many of you have a code base that if you refactored, test would break?
> Most people raise hand
and it made me realize maybe I'm not as stupid as I think I am.
I've been trying to understand how to create unit tests that allow me to refactor for years, and have been completely unsuccessful. The only way I can achieve this is using the "classicist" viewpoint of creating many "unit test" and only isolating the architecturally significant boundaries. That doesn't make me happy either, though.
But, yeah, good to know (or bad to know) that I'm not the only one that struggles with this.
Back to the video I go.
What I do think is kind of crap is the holy war of methodologies that exists. In my experience, it usually happens that some new CTO or other manager comes in and says, "We are doing it wrong! From this day forward, WE ALL MUST USE TDD/AGILE/WATERFALL/YOURFAVORITEWAYHERE".
There is no latitude given, and folks who are not used to the methodology now take 1.5 to 3x longer to complete things, and deadlines slip. Then the push to complete work means that things get written sloppily and tests are not well thought out or people arent fooing their bar with the baz properly. Quality suffers in the short term, but eventually everyone catches on, and life goes back to normal.
Then the new CTO shows up...
I'm not a fan of TDD (in the sense of writing the test first) but I do think that some sort of testing-as-you-go is important for back end API work in particular. I don't think that 100% test coverage is a good idea for 99% of business use cases.
And really, we’re playing a probability game with ourselves, trying to reduce the probability of typing the wrong thing. Writing it twice is one way to do that.
TDD means here
Test Driven Development
writing a full test battery after that is probably overkill, but making sure everything can be taken out of the running app and tested singularly is essential to be able at a later date to get a user bug report and convert it into a testable case to narrow down the root cause.
a good strategy for that goal is to make use of dependency injection at each layer separation and to make sure that every user generated event can be also triggered pro grammatically - that's especially useful as while relying on something like selenium do work, it's exceptionally costly and aggravating in the long run.
being able to isolate the bugged behavior and responsible component is the major advantage that comes with the full TDD implementation, but that doesn't mean you can't have enough of that with a lighter approach
I laughed out loud
In other cases the cost of refactoring with TDD is very high if any significant design or architecture changes occur, so in those cases I find it is useful to stabilize the broad strokes/patterns a bit before investing to much in tests.
https://i.redd.it/lwin56fisdsz.png
It doesn't look like a joke to me. It only works over integers so the code is absolutely correct.
It also strikes me as the kind of convoluted logic that someone took a really, really long time to come up with, before it finally worked. (As indicated in the comment.)
A test can hardly capture what's wrong with this code. But any human can see it instantly. (And it's kind of weird that the programmer didn't.) I think most people can think of braindead decisions that are not really captured by testing.
Someone mentioned DI not getting rid of coupling and I agree. DI is a tool you might use but the way to get rid of coupling is not one simple thing but a process of a bunch of different tools and techniques. You can't just slavishly fallow some process and expect it to fix all your issues. You have to think and do work yourself to fix it.
Code like this is basically only possible through testing, because the programmer doesn't understand why it is working: it says so in the comment, which I believe.
So if it weren't written against testing (manual or automatic) it simply couldn't be written like this.
There are a lot of really broken designs that "work". Testing gets you "working" code. Often you could do much better starting with correct code and then adding testing afterward - which I believe is not the essence of TDD. TDD is about driving the writing of code by tests, and I believe the example I shared is one of these bad outcomes. (There are others.)
With people like Capers Jones and others doing piles of studies 30+ years ago, I'm confused why someone says there are no studies on TDD.
My bet is the author doesn't have access to the relevant historical papers, and doesn't know they exist.
I'm on my phone so I can't link a lot, but I maintain that the author should do a lit review of software quality in engineering, and they'll get better conclusions
I'm not sure if we're taking away the same things from these studies, as their whole conclusion is it actually takes less time and costs less, due to early wins in quality. The papers claim it's not a case-by-case thing.
https://medium.com/southprojects/tdd-a-business-crud-is-it-w...
tl/dr: In theory TDD sounds great but in a real example, TDD is not magic with a real limited coverage.
First, it's probably not true. Linus Torvalds is not a legendary human being who can write critical systems without a single flaw. He relies on legions of human beings to carefully check and review every line of code before he even looks at it. There are discussions on mailing lists. There are arguments and disagreements. There's a process there. He doesn't just flit his fingers across the keyboard and output amazing, error free code. It probably has tonnes of errors.
Linus' philosophy is that errors aren't the end of the world and someone will patch them when they are uncovered.
For some use cases that's fine. However there are plenty of applications where a more proactive approach to correctness is necessary: real-time systems, safety critical systems, and yes... even security.
Maybe TDD is a misnomer. I think we should call it specification driven development. Unit tests and integration tests are just a weak form of specification. They provide theorems in the form of examples that we try to prove with an implementation. Property based tests give us more examples to quantify assertions over. Model checking can test liveness as well as safety in our high-level designs... how much you need to specify and how thoroughly really should be a factor of the risk and complexity present in the requirements of the system.
To use an analogy: blueprints. If you're just building a shed or a small footpath then it's enough to sketch your idea on a napkin. If you're building a house you need to have a more specific and detailed plan that passes by a civil engineer. And if you're building a sky scraper then you need to be thorough and able to convince others of the validity of your designs.
(credit for the analogy should go to Leslie Lamport).
I think most software projects are at the house level in terms of risk. You could get by with using a dynamic language and a few unit tests if you value productivity more than correctness. That just means you're willing to accept that you will have higher reported error rates and are comfortable with potentially losing customer data or a higher risk for security vulnerabilities. You can lower your risk if the project requires more sensitivity to data consistency or security by using a sound type system and encoding your assertions at the type level, add some property-based tests, and more integration and unit tests. It's a spectrum one should consider.
I know we all like to write code and sometimes we even hear ourselves saying, "Well if you wrote the perfect specification you might as well have written the software," but don't be fooled.
"Software engineering is the part of computer science which is too difficult for the computer scientist." -- Friedrich Bauer
I absolutely agree with your disagreement here.
I was using theorem and proof as an analogy to illustrate the separation of specification and implementation.
A useful distinction as you get further along in writing formal specifications.
Update: typo.
If one also writes a serializer, then one can additionally test for any input the property: `serialize(parse(input)) === input`. This means adding a new test is just dropping in more example inputs (say from bug reports).
Now one can go further, and define a set of mutation operators that can act on input to produce a new valid input. Reorder tokens, change data at leaves, delete and insert new data, etc. Now one can generate arbitrary amounts of new test cases based on existing input examples.
Other mutation operations can be designed to generate invalid inputs, which should always give an error (never crash, halt or unexpected exception).