An unexpected benefit of unit tests
matthewc.dev
matthewc.dev
I appreciate someone writing into the void how this helps them focus and be productive in their workflow. Who cares whether it 'fits the technical semantic definition of INSERT_ACRONYM_HERE'.
Unit tests are a literal investment - somewhat like buying insurance. Sometimes they're really great - almost necessary - and other times they're not worth the price. The article is offering a point of view of how tests are paying for themselves in an indirect way, which is useful to think about in the cost benefit equation.
The position of "insurance is useless" means maybe you just haven't been in a situation where it could really, really save your bacon. Likewise, if your rule of thumb is "buy all the insurance, all the time" - then BestBuy will gladly sell you a $5 'product protection plan' on your $10 purchase.
Did I not have a concrete idea of what TDD was when I wrote it? Yes. I had muddled test first development with test driven development as the two people I knew who were big on TDD were focused on TFD as they were refactoring large legacy codebases. I'm still murky on what TDD is (and it sounds like opinions differ) but I don't think I've discovered anything new. Just new to me.
1. Write a test
2. See that it fails (if it passes, you're testing a feature you already implemented or you're testing something you don't mean to be testing)
3. Make it pass (by changing the code, not the test, seen that too many times)
4. Refactor
5. Repeat
The catchphrase is "red-green-refactor". I don't know why people try and make it more complicated than that.
Regarding your post, just wanted to say that if you took my other comments as criticism of you they were not. Your misunderstanding of TDD is actually very common and I was intending to address that, specifically. I actual consider the fact that you both rejected the flawed TFD-style (and mass creation of tests that comes with it) and rediscovered TDD (the above process, more or less) both amusing and a positive thing. It shows you're doing what many others aren't, actually thinking and experimenting to discover techniques that work (at least for you and on this project).
Given that the author of the term wrote a whole book about it rather than just tweeting those steps, I think over-complicating may be part of its genetics. Others are just continuing the trend.
---
For the record: yeah, short-cycle red/green/refactor fits how I like to do TDD. I don't know that I'd claim that it is TDD though - by far the majority I've run across take it to extremes, though of course my coding-exposure circle is skewed and may not even remotely be representative.
In that way, TDD sits in a very similar brain-space for me as Scrum does. Useful foundations, utter nonsense if taken to extremes, unsure if a net-positive influence (globally, not personally) in the end. I have seen a looooot of negative-value tests, in ever-increasing quantity as time goes on, in the name of TDD.
Letting ideas like this move into the "similar brain-space ... as Scrum" is dangerous. You're actively choosing not to think at that point, you're choosing anti-thought over thought. Move too many things into that brain-space and you're going to become thoughtless. Don't let that happen, you have more to contribute to the world than becoming another mindless automaton.
Engage with the idea instead, understand why it does or does not make sense to you. Why it is or is not applicable to some area. Why other people have different ideas about it or disagree on its applicability. You'll be much better for it. And as much as I also dislike Scrum, it's worth engaging with its core ideas as well.
I've looked at those things. Found what I think are the core valuable bits, and how they apply to me and my situation. And I do not want to interact with them in a professional setting, because they are almost always distorted beyond any reasonable interpretation of their core, losing all meaning of the term and being used as an excuse to continue bad behaviors. They're quite similar in that respect.
1. Write a feature
2. Test it
When your feature is your whole app "I want to build a TODO list app", your test is "I can write a TODO list app".
When you inevitably can't do that (because you cannot write a unit test for this), you break it down "I want to open a window/app", test that "can I launch the app window", and iterate.
"and still somehow greatly misunderstood"
It misunderstood because its presented as a methodology involving some testing suite or library, then you're constrained to testing things that may have little to no value, like testing that classes have data on them, or that you call functions and they return data.
And considering (most) engineer's first contact is something like a "write unit tests for the class you wrote", its understandable that its seen as a mixture of useless (often is) and arcane (often is) instead of an extension of a methodology for creating and validating features.
Im on a team that is currently attempting shift left. One part of this has been an argument that testers should review unit tests. I think this is not a good idea. Unit tests tend not to focus on the behaviour of system as a whole.
It also goes well with the pomodoro technique, because that almost forces you to interrupt your progress while in the middle of something, leaving things unfinished, so they end up easier to pick up later.
> The best way is always to stop when you are going good and when you know what will happen next. If you do that every day … you will never be stuck.
Writing all tests up front has never been TDD.
I'm a distracted developer, but you'll never guess the unexpected advantage I discovered of pair programming...
For instance, one of the things they say beyond the "red/green/refactor" is "as the tests get more specific, the code should get more generic." This is an interesting concept in itself.
Edit:
To clarify, if you are writing the least code to make it pass, and do not have the concept that the code needs to get more generic, you /will/ think the whole thing is silly.
TDD commonly gets mischaracterized as a two-step process of writing a suite of tests upfront, then writing some implementation to make them all pass. It's actually much closer to what is described as TLD in the post. You write a failing test, do the bare minimum to make it pass, refactor if appropriate and start the cycle again. It's a development process that will (in theory) produce a high quality implementation and suite of tests at the end.
His TLD is pretty much TDD except he doesn't mention refactoring after passing the tests. Even leaving a failing test at the end of the day as a kind of "todo" for the next day, I'm pretty sure Beck mentioned using that idea in his TDD book.
EDIT: Found it finally. From Beck's Test-Driven Development by Example (page unknown, ebook copy, chapter 27):
> How do you leave a programming session when you're programming alone? Leave the last test broken.
> Richard Gabriel taught me the trick of finishing a writing session in midsentence. When you sit back down, you look at the half-sentence and you have to figure out what you were thinking when you wrote it. Once you have the thought thread back, you finish the sentence and continue. Without the urge to finish the sentence, you can spend many minutes first sniffing around for what to work on next, then trying to remember your mental state, then finally getting back to typing.
> I tried the analogous technique for my solo projects, and I really like the effect. Finish a solo session by writing a test case and running it to be sure it doesn't pass. When you come back to the code, you then have an obvious place to start. You have an obvious, concrete bookmark to help you remember what you were thinking; and making that test work should be quick work, so you'll quickly get your feet back on that victory road.
> I thought it would bother me to have a test broken overnight. It doesn't, I think because I know that the program isn't finished. A broken test doesn't make the program any less finished, it just makes the status of the program manifest. The ability to pick up a thread of development quickly after weeks of hiatus is worth that little twinge of walking away from a red bar.
It’s also sad to see how twisted people’s idea of both agile and tdd have become, usually because they never read the source material.
I'm not sure where to put the blame, but "Agile" does get a particularly bad rap these days. Mostly I'd say a corporate watering down of Agile concepts, money hungry consultants who didn't know much, and I h ave a personal dislike of how most of SCRUM tends to be implemented.
TDD in particular has fallen off the map in a way that I find very surprising. Generational amnesia I suppose.
I think inverting agile concepts is more accurate.
A key element of Lean is empowering the workers to improve the processes. Let them come up with ideas that improve things and run experiments (guided by management perhaps, but not directed by). But the way USAF did it, the managers would watch a process being done, identify "wasted" movement, and then rewrite the process/procedures to eliminate that wasted movement. It was clearly just Scientific Management but being called Lean because they "leaned out" the processes. Naturally, the actual workers did not like coming in every other week and having to learn their job all over again. After a while they still held "Lean Events" but by then it was for show rather than to actually effect change in how things were done.
Sometimes it makes sense, corps don't actually have any need to churn out new features, at that point you just have kanban, or even waterfall. Be nice to stop pretending, though.
Versus communism where people (e.g. Benjamin Tucker) were predicting its primary failure mode decades before it ever was instantiated at any scale.
This hinders freedom to experiment with new techniques and methods within development teams, but it doesn't stop it. A "trick" is to provide all the artifacts that they want as if you followed their process to the letter, but still do things the way you want so long as it gets the job done. The problem with that is that you have no evidence you did things differently than the defined process and so they'll continue to believe the defined process is perfectly fine, if not excellent. Then some exec will decide to write a book about it and become a consultant selling the (broken) defined process (I assume this is how SAFe came to be, an ironically very rigid "agile" process).
People are the problem.
What I mostly see it is people doing it to themselves. Someone thinks he has to be guardian of the process and refuses to let the experiment run.
The result: if you're not using safe you're lead is going to be removed. Zero consistency between project tooling.
It's like the worst of both worlds.
I also worked at a funded start up that implemented agile from the beginning, worked pretty great.
Do you really want to go back to the waterfall and V model style of development? Those are basically guaranteed failures in a fast moving industry like software development. If anything your comparison should be applied to the waterfall/V model because it is essentially central planning.
If you are doing things like building MVPs, iterated/incremental development with frequent changes and deployments to production then you are doing agile development.
Maybe it is not hyper formalized like Scrum but it is agile nevertheless.
page 148 in the physical book, at least my copy.
However, you leave out the following section titled "Clean Check-in" in which Beck argues the exact opposite when working on a team. Clearly it's the days before git and trusting code to sit on your PC overnight without checking it in (I, too, would not trust Windows98 with code that long).
But this is the general problem with the book and TDD. It's outdated and overrated. The book is not a well-written book, even for year 2000 standards, and I'm quite surprised people are still referencing a 20 year old book that is almost entirely composed of trivial examples using Java classes and objects. It is my least favorite tech book on my shelf.
> leaving a failing test at the end of the day
As far as this point goes, you pretty much have to. The TDD methodology, per the book, is to get to green as fast as possible. If you are testing for a function to have inputs of 5 and 2 and expect an output of 10 then the book literally tells you to do:
function() {
return 5 * 2;
}
What's going to happen is you write that code, get distracted or need to quit for the day, and you come back and your test is green. You forget that you wrote some total shit like the above and move on to the next Jira ticket.As a methodology or system it's just bad. Imagine instead of writing the implementation to pass the tests that you instead are writing tests for some AI to "implement". You set the inputs and the expected outputs and the AI goes to work and does the implementation. In AI this would be called overfitting. Yes, the tests pass. But only for the cases you wrote. There is no guarantee your code works for the general case. Now replace AI with you and the same thing will happen. If all you care about is green tests you're liable to stop thinking of what the implementation should be doing.
Sure, you won't do that. But I promise you people, in general, do. One code base I worked on had minimum coverage required. More than half of the tests were total garbage. They either tested nothing useful, or were false positives that could never fail (because no one knows to check for failure, ever!). This is the opposite scenario, but the same outcome. It's the difference between the letter of the law and the spirit of the law. Give the people a rule to follow and they will mindlessly follow it.
Agile and TDD are both too nuanced and leave too much up for interpretation that it's no surprise we continually see debate on "true TDD" or "true Agile".
There are useful ideas in TDD. But this would be better packaged and sold as "here are some ways to build software under X scenario." These are techniques applicable to a time and place and not a paradigm.
Regarding that code snippet:
If you forgot that you hadn't finished it, and then check it in, then that's on you. Good news, hopefully you and your office aren't morons and you aren't relying just on the tests from the TDD bits, because TDD itself doesn't directly address creating integration and end to end tests. So your integration tests will catch that. And if not, it'll make it to production and your customers will, rightly, call you a moron. And you'll be embarrassed, write a test to catch the error, fix it, and hopefully not fuck up like that again.
TDD doesn't aim or claim to cover all the testing needed to verify and validate a system. It's one part of the whole (if you use it at all). In fact, it doesn't address validation at all so that's something you have to cover another way entirely.
> If all you care about is green tests you're liable to stop thinking of what the implementation should be doing.
If all you care about are green tests, I'd say you're aren't just liable to stop thinking but that you have stopped thinking. You have become, in my more polite way of saying it these days, a fool. My advice: Don't be a fool, you have a brain, use it.
But when I did, it not only made testing better, it made my code better too. Not only because it’s more testable, but because it makes me think about the interface first, and the implementation truly as a black box as much as possible.
In Ruby/Java it is certainly a bit more of a chore to remember to do.
Link for the lazy: https://www.ncrunch.net/
If you skip straight to green you don't know, for certain, that the test actually tested what you expected. This isn't even a TDD thing. When you're working on an existing, deployed, system and a user finds a problem, you generate a new test (well, sensible orgs and people will). That test will fail, because you haven't addressed the issue yet. That is, it's "red". Then you make it pass by fixing the system, it becomes "green". That's it. If you fix the system and then write the test, do you know that the test actually recreated the original failure? Or is it merely exercising the new or altered code?
But when writing new code?
For new code, the reason it makes sense is that your system is bugged. It does not do what it's intended to do yet because you haven't written the code to do it (or altered the existing code to add the new capability). So the test detects the difference between the current state of the system and the desired state, and then you implement the code and now the test detects that you have achieved your desired state (at least as far as the test can detect, you could still have other issues).
Absolutely, and most prominently when writing new code. The red-green transition is absolutely essential for new code.
You should not write any production code except to make a red test green.
Think of the tests as the specification of your system.
If all tests are green, your system meets the specification. Thus there is no need to write production code.
So in order to make the system do something new, you need to first change the specification. So you add a test. When you add this test, it will almost certainly fail. After all, you haven't written the code to implement the feature.
Then you make the test green, and now the system once again matches the (now updated) specification. Commit/Refactor/Commit.
Having the test that is red also validates your tests. If your tests are always green, how do you know that you're actually testing something?
In fact, it sometimes happens that you write a test that you think should be red, because you haven't implemented the feature yet, but then it starts of as green. Meaning you inadvertently already built the feature. This can be very confusing... :-)
If I only see green for a given test, I have no way of knowing if it is asserting anything at all, much less if it's testing what I thought it was.
"New code" just means "I want a program to do a thing, and it doesn't do it yet. That's a bug." The difference between "bug" and "new feature" is more a matter of perspective than actual development effort.
Like if there was a virulent disease for which there was a 100% cure, but you can't get the cure unless you test positive for the disease. I give you a test and say "Yep, test says you are healthy". Ok. What if the test always says people are healthy? "Have you ever tested an unhealthy person and the test detected that they were unhealthy?" "Oh, no, we've done this test a thousand times and it always says people are healthy!"
You write the test first, because your code does not yet have the feature that you are testing. Your current code is a perfect test for the test. Anyone who has done TDD for even a short amount of time has written a test that should have failed but instead it passed. Sometimes the was just a simple error in the test. You fix the test so it can detect what you are looking for (i.e. the test now fails). Other times a fundamental misconception was discovered that blows everyone's mind.
extern bool run_hook(char const *tag,char const **argv); // [2]
tag is informational; argv is an array where argv[0] is the program name, the rest are arguments. Okay, how would you write a test first for this? You literally can't, because the function has to exist to even link it to a test program. Please tell me, because otherwise, this has to be the most insane way to write software I've come across.[1] LEGACY CODE!
[2] I go more into testing this function here: <https://boston.conman.org/2022/12/21.1>. The comments I've received about that have been interesting, including "that's not a unit test." WTF?
Unless of course you're only working no trivial programs (based on your write up, not the case) or an absolute genius you must have at some point or another encountered a failed compilation and used that as feedback to change the code. This is no different.
Yes, I've gotten failed compilations, and every time it's because of a typo (wrong or missing character) that is trivial to fix, no test needed (unless you count the compiler as a "testing facility"). That is different from compiles that had warnings, which are a different thing in my mind (I still fix the majority of them [1]).
But I'm still interested in how you would write test cases for that function.
[1] Except for warnings like "ISO C forbids assignment between function pointer and `void *'", which is true for C, but POSIX allows that, and I'm on a POSIX system.
At some point, I, and probably most people, operate under the assumption that we don't need to test (ourselves) that syscalls will do what they say they will do. Until they actually fail to act correctly and then I'd investigate it, and write tests targeting it to try and reliably replicate the failure for others to address since I'm not a Linux kernel developer.
It may seem cynical, but I assumed that anyone into "testing" (TDD, unit testing, what have you) wouldn't bother with testing that function, or with limited testing of that function (as I wrote). You aren't the first to answer with "no way am I testing that function to that level," but no where have I gotten an answer to "well then, what level?"
This may seem like a tangent to TDD, but in every case, I try to see how I could apply, in this case, TDD, to code I write, and it never seems like it's a good match. What I'm doing isn't a unit test (so what's a unit? Isn't a single function in a single file enough of a unit?). I'm not doing TDD because I have to write code first (but then, the testing code fails to compile, so there's not artifact to test).
People are dogmatic about this stuff, but there's no discussion about the unspoken assumptions surrounding this stuff. Basically, the whole agile, extreme, test driving design seems to have fallen out of the enterprise area, where new development is rare and updating code bases that are 10, 20, 40 years old are the norm and management are treating engineers like assembly line workers, each one easily replaceable because of ISO 9000 documentation. And "agile consultants" are making bank telling management what they want to hear, engineers be damned because they don't pay the bills (they're a cost center anyway).
Anyways, you never asked me "well then, what level"? and I thought I did answer it but here's an answer anyways (to your unasked question): I'd test it to the point that made sense. I wouldn't follow some poorly considered hard and fast rule (morons do that, we're not morons, we are humans with brains and a capacity to exercise judgement in complex situations). A hard and fast 70% code coverage rule is stupid, as is 100%, even a strict 1% rule is stupid (though for other reasons, like that it's trivially achieved with useless tests for almost every program). If I'm writing code and 90% of it is handling error codes from syscalls, then you'll likely end up getting 10% code coverage from tests (of various sorts, not just unit) out of me. I'm not going to mock all those syscalls to force my code to execute those paths, and I'm not going to work out some random incantation that somehow causes fork to fail for this one program or process but also doesn't hose my entire test rig. Especially not when the error handling is "exit with this error code or that error code". If it were more complex (cleanly disconnects from the database, closes out some files) then I'd find a way to exercise it, but not by mocking the whole OS. That's just a waste of time.
To reiterate my take: We have brains, we have the opportunity to use them. Use the appropriate techniques to the situation, and don't waste time doing things like mocking every Linux syscall just because your manager is a moron. Educate them, explain why it would be a waste of time and money and demonstrate other ways to get the desired results they want (in a situation like your example, inspecting the code since it's so short should be fine).
There is no goal post moving.
More likely, as we transmit information, we don't do it correctly, and knowledge/data gets lost. I found quite enlightening to always go back to the source.
But in this situation the test will fail regardless of what you wrote in the test code. So the supposed usefulness of the test failure showing that you are actually testing what you mean to be testing is inexistent and the exercise of making it fail before making it pass is pointless.
If this is actually the one thing that trips you up on TDD, then don't do this one thing and try the rest. This is the easiest part of TDD to skip past without losing anything.
I also like that Set of Unit Testing Rules. That is basically correct, external systems are a no-no on unit testing.
Usually, you deal with mocks through indirections, default arguments, and other stuff like that so you can exclusively test the logic of the function, which is more difficult in C, from what I've seen on your write up, than in other languages. But if you care about not having that on your code for performance reasons, then more likely than not, you will not be able to unit test. And that is fine. You have an integration test (because you are using outside systems). You can still do integration test first, as long as they help you on capturing the logic and flow. The issue is that they tend to be far more involved, and far more brittle (as they depend on those outside systems).
https://www.jamesshore.com/v2/blog/2005/microsoft-gets-tdd-c...
(Well, they popularized it.)
With code I'm getting paid for I'll write more tests up front and won't skip that step, but for code I'm playing with then as long as I'm just enjoying myself writing code I'll write zero tests for awhile and just code. Then as I hit the debugging/refactoring step I'll do the backfilling as a form of debugging. Often I'll write the code that fixes the bug, then write the test, then quickly and temporarily revert just the code to ensure that the tests fail and then proceed. That gets you the same safety check as doing it the test-driven way to validate your test actually tested the right thing.
I really tend to hate "thou shalt start by writing tests" as some kind of immutable golden rule. It does always feel good to me when I just naturally wind up doing it, but forcing myself to do it every single time just isn't any fun at all.
At the same time tests are absolutely essential when it comes to refactoring and debugging. When you get too far out ahead of yourself with code then you start needing to shore up the foundations and use tests to eliminate bugs in the code that you've already written. Some code though is obviously correct enough that in personal hobby projects I won't ever code tests for them (unless I do hit the point where I start to doubt their correctness due to some funny bug at which point the situation has changed so I add some tests to prove it one way or the other).
The whole point of this though is that the tests are always serving me and they aren't in the drivers seat quite the way that all the TDD 101 blog posts like to ram down your throat and which I suspect turns people off from that approach so much.
The end result is also that you'll tend to wind up with the tests that are actually useful, covering the code that is particularly hairy or essential and the edge conditions that you really need to make sure to get correct, and you wind up having a test suite which is composed mostly of useful tests instead of all the largely useless ones that infect codebases.
I'll also happily omit tests on lower level functions that are well tested at the level above them, because I don't need to test the same thing at 18 different levels (again, for professional use I'm more likely to include tests at every level if they're fairly mechanical to produce). I also have a flexible definition of what the system under test is, which often encompasses more than just the immediate object that I'm testing and I don't bother wasting mental effort thinking about how to mock the whole world.
I don't know what kind of TXX that is. I still wind up with tests, they're legitimately essential to have, I just don't get there via some prescriptive route. I wind up with good code coverage, but I don't necessarily wind up with it looking as comprehensive as rotely banging out lots of unit test. I typically wind up with tests that I know are useful because they were produced by hitting actual bugs or where I had real questions about the behavior of the code and needed to assert some invariants and prove the code worked.
The dogma most people see or claim to see is that TDD is meant to be used everywhere (or nearly). Which some fools, yes, believe. But they're just that, fools. People who use their brains (aka, non-fools) know them to be fools and do what works for them and the circumstances because they spend some time thinking about things instead of parroting a dogma (or an anti-dogma).
At some point the No True Scotsman fallacy kicks in pretty hard and that is just what TDD actually is.
And yet TDD preachers are never drawing the boundaries for TDD applicability. They are always extremely vague: "sometimes I see that TDD doesn't work for the problem that I'm solving and I don't use it". Well, how do you see it? What types of problems is it bad for?
If TDD works, they take credit. If it doesn't work, "it's just a tool" or "you used it wrong".
It is literally impossible to prove that TDD doesn't work. Which makes it a religion.
The only study I've seen on this shows only tests were correlated with more correct programs, but doing tests first or last showed no significant difference.
Step 4 is the trickiest and most important part. Refactoring transformations must be such that they do not invalidate the results of step 3. But if you don't execute step 4 properly you'll end up with crap code.
2. Step 4 is indeed the most important part, and yet TDD priests and scriptures don't cover it at all. TDD is actually distracting you from what's important, because it focuses on steps 1,2,3. Eventually you'll become disciplined enough to not get distracted, and you'll think that TDD works. But the reality is that you never needed TDD in the first place.
Such cases are relatively few.
Hardly a reason to downvote me.
As to the downvote, I guess you thought it was me, it was not. But I don't plan to upvote you either.
For the moment I will assume that I have a misconception that led me astray, and that I should correct in the future once I double check the thinking and the definitions etc., and I'll upvote you now for leading me in a better direction.
If I understood your posts right, it's the same misconception a few others have had in here:
In TDD, it's not write all tests, then red->green->refactor. It's write one test, red->green->refactor, then write one more test, red->green-refactor, repeat until done.
Real TDD is red/green/refactor/repeat in very small steps, writing a handful of lines of code at each step.
"I rediscovered something obvious, mischaracterised what exists already, and I make up a term and pretend like it is novel, and credit myself with it"
Hoping someday to be the new Fowlers, surely.
Keeps me focused.
In fact the theory is that this approach will produce a high quality design as an emergent property. This is an extraordinary theory that requires an extraordinary proof - one I haven't seen so far.
During "red", you're thinking about the design of your public interface.
During "green," you're focusing on implementation.
During "refactor," you're thinking about how to improve the quality of your implementation and how to improve the overall design, and making those changes.
If you believe that spending a lot of time thinking about and improving your design will produce a high-quality design, then TDD will produce a high-quality design. QED.
If you don't accept the axiom, then it's a longer discussion, but that's the proof, and my experience is that it does in fact work.
(If you're looking for a rigorous study and proof, you won't find it, because there are no rigorous studies that formally prove what creates high-quality design. Partially because there is no formal definition of "high-quality design" in the first place.)
"Refactor" is named this way to emphasize that changes you are making are closely related to the tests you already have and the tests you are about to introduce.
That does not leave enough space to justify a QED. If you choose to design beyond that, the process stops being TDD - at least as described by Kent Beck
1. It assumes your spec is good and rigid. In reality, most specs are shitty and fluid. And your first understanding of spec is wrong.
2. It assumes your first implementation of the spec is good enough to justify automated testing.
3. It leads to high test coverage which inhibits refactoring (despite zealots telling you otherwise).
4. Almost always it leads to obsession with testing, which leads to a ton of unnecessary complexity (e.g. dependency injection for the sake of testing, weird practices like "don't mock what you don't own", etc)
I always thought it's great that it forces you to think about public interface, but I came to believe that thinking cannot be forced with a ritual.
The major flaw with TDD is if you get your test wrong, you get the wrong code. The intent is to inhibit refactoring, because the assumption is the tests are correct, so any refactoring must be done within the constrains imposed by those tests.
OFC the tests and supporting design are usually just as flawed as the code. This is why I say the first step of TDD is wrong, write your feature first, not your test. If you don't have a feature, and can't logically reconcile it with your other features, then its not worth even writing the test in the first place.
Tests are just a supporting tool once you (believe) you have the feature written, which functions on one hand to protect the other features you have written (at least as well as they are tested), and to validate that you aren't wildly breaking the system expectations. A large number of tests is a measurement that something is wrong, but it doesn't tell you if the feature itself is wrong or your design is wrong, just that one of the two is true.
That's how it helps you refactor, a "good" design will add new features and few tests will break, as more tests are added and total test failures approach zero over time, you gain some confidence that the system is good. You never gain certainty, just the knowledge that your constrained refactor probably didn't break anything.
TDD has a wikipedia page [1] FFS, which is the #1 hit entering TDD into google, and which very clearly lays out that TDD is a test/code cycle (red, green, refactor anyone?). What the author claims is TDD is called TFD (Test First Development).
How do people develop this kind of hubris?
The absolute bare minimum should be for you to write those informal tests down and commit them to the repo. That's some golden knowledge I've gotten over time
It just requires the overhead of setting up the test runner
It's not really any different than in dynamic/interpreted/weakly-typed languages. "Writing the test" for a function sometimes just includes writing a method/function with the appropriate type signature that does nothing (maybe returns a dummy value).
Forcing you to view the code you're implementing from the viewpoint of someone calling that code from the very beginning is one of the advantages/goals of TDD. If you find that it's difficult to set up the objects/data you need to write a test, eg, your code has a bunch of implicit dependencies on other components being in a particular state or it takes a ton of arguments that all have to be constructed, that's usually a strong indicator that you should rethink the design. You're getting early feedback on your API design before you waste time implementing it.
Instead of just writing a class with a bunch of methods that seem useful, you write the tests to figure out what a user actually needs, and iterate on that. And as a wonderful byproduct, you get a suite of regression tests so you can comfortably refactor and add features in the future.
There's no fast and hard rule, I'm afraid.
IOW, you will never have your bottom layer done correctly or completely.
At some top layer you're going to think "well, looks like I won't be need that function", OR, "Well, looks like I am missing a function for $FOO"/
It's the premise of On Lisp, for example, http://www.paulgraham.com/onlisp.html
I'm not saying it never gets fixed, I'm saying it's extra work compared to top-down.
Top-down produces only and exactly what is needed to get the current layer to compile, run, and pass/fail.
Bottom-up requires that the bottom-most unimplemented layer be written to provide at least what the next higher level needs.
And since you cannot predict that perfectly all the time, you will have to come back and refine (not refactor) that layer: remove some functionality that turned out not to be needed, or add some other functionality that turns out was needed.
This step is never going to happen in a top-down effort.
[PS. At work I've almost always done bottom-up, because that suits the workflow constraints better: 1) Never leave commit a non-working build, 2) never do a PR for an incomplete module, 3) Always make PRs as small as possible, etc.
For my own projects/hobbies/experiementation, I used to do the same thing. Now I'm experimenting with the top-down approach and my velocity is a little faster because I am never spending time removing stuff that was written.]
We're using this approach in trading production system for about 5 years now and I recommend it, works very well.
There is also side effect where your dependency tree is more shallow.
> We're using this approach in trading production system for about 5 years now and I recommend it, works very well.
I haven't worked anywhere since 2004 that used any other approach; it's the dominant approach to development - commit code, starting from the lowest layer to the highest.
My experimentation thus far using "commit code, starting from the highest level to the lowest" has been more satisfying to me in terms of velocity.
Some might call that thing a 'narrow waist' but I don't quite equate it to that although it's a good example. Is the thing you want a stream of bytes, sequence of messages, unsequenced messages, or priority-ordered messages?
The problem with bottom-up is that you don't know the context. The problem with top-down is that you do know the context--and can leak it into what should be without. I find the latter works out better if you're aware of avoiding it. How do you avoid not having context without going top-down?
An analogy in UX would be Windows vs Mac. Windows builds things then puts the UI on what they've built--in many different kinds of settings places. Mac figures out what information & controls the user should have and how to name and group them.
Edit: I actually have a recent real-world example. Making a subsystem for querying and mutating things, what they are doesn't matter. We set out making an 'adjust' operation and a 'move' operation because we wanted to preserve paired decrement/increment amount for moves. We were also certain that we only needed to move between two things, or adjust many things but only of one kind. We got all the way up to the top where the public API was getting close to complete. We discovered use-cases that made sense to to more than those operations. We were able to shuffle things around and ended up with a fully-capable 'adjust' operation that can work on any number of kinds/things, and a fully-capable 'move' operation that can work on any number of kinds/things. On top of that we added checks to only do what we need now. The more general things were substantially harder to make work efficiently which is why we chose not to do it when we were sure we didn't need it. We were wrong.
You know the context. You have the whole problem in front of you. The question is about which direction you're going to solve it.
For "you want a stream of bytes, sequence of messages, unsequenced messages, or priority-ordered messages?" the answer is you probably want type parametrized iterator at this level if you can ask this kind of question.
To put it in other words you put more attention into trying to find underlying composition of algebras in problems than doing adhoc, inlined imperative constructions littered with if statements that you add every time some issue is discovered.
I've repeatedly, throughout my devlog, emphasised that I'm calling functions that don't exist, then I create stubs of those functions (a one liner returning an error), then I populate the stub with actual logic.
To me, this is the most natural way of writing code.
It's explained right there on the cover Test *Driven* Development. Not a testing method. A development process guided by unit "tests". In quotes as unit tests are at least 51% about forcing you to write small, isolated blocks of code with well defined and simple interfaces. Units. A large percentage of remainder is enabling ruthless refactoring (refactoring being a huge part of TDD's development practice/philosophy) by ensuring you have not violated those interfaces. Only a few percent is actually about having "correct" code.
And tests that describe how an API can be called are BDD. Hold the Cucumber.
But the one thing I can't agree on is that teaching them TDD is either necessary or sufficient for them to start thinking about that. I'm pretty sure that if you can manage to write code without knowing how it will be used, you will write tests that way too. Writing a test doesn't force you to think about your API any more than writing the API.
I am a new to development but I've noticed that lots of errors pump through code where "No method X defined for nil object" or something like that. I wonder if I could somehow make the case that we'd spend less time reacting to problems and bugs if we spent more time up front writing tests that could give us more than 19% code coverage.
Does anyone have advice for starting this conversation with my boss or during one of our standups? I know I could somehow pitch the value as "less problems later for more time spent now" but is there a more effective way to say it?
To get out of that situation, I'd recommend leading by example and showing how you have personally managed to use tests, in the space they are operating in, to make your life better and develop faster. Once you have something that has provably worked for you, you can start evangelizing it and onboarding other people.
And just adding tests isn't the only thing it takes to improve the software quality.
Thats the kind of thing a good programming language is supposed to help with.
The rationale being that until the API is fixed, I don't want to have to adapt the tests to the changing API all the time.
I totally understand where TDD shines. When you have a well defined problem to solve, unit tests are the definition itself, and the business logic therein.
The rationale behind test-first is that you won't know what the API should look like until you try writing some sample client code that uses your provisional API to access some data. So why not make those tests, iron out the kinks in your API and have a few functional tests at the end of it; two birds, one stone.
Since an API is a composition of multiple parts that can be tested (the unit in unit testing), it only makes sense to create these tests once the relations between these parts/units are well-established.
That doesn't have to always be after the API is completely done of course. Some internal functionalities are likely to be immutable. They can be tested quite early in the iteration timeline.
What you describe however seems closer to integration/system testing which is one approach I and other people in this thread also tend to favor. But I've been writing UI code predominantly so I am biased here as unit testing has lower value.
Writing code in pieces (functions) where each does one thing, even if that thing isn't reused elsewhere outside of the one place you need it, if one way to reduce the impact of that house of cards; because you only need to focus on the logic that's important "where you are".
Unit tests are another way. You can describe the behavior of your code in tests and, if you break something because you couldn't keep the entire model (and individual pieces) in your head at the same time, your tests help you notice that. If you're building your code initially, your tests can help you identify and focus on individual behaviors you need to work.
I think the above is why I think both pre-code and post-code tests are useful. It's helpful to write tests before you write the code, to guide your development. But it can also be helpful to write tests after; especially as you identify things that don't work correctly because you didn't realize they were requirements when the code was originally written. They are, effectively, regression tests... but they're more than that, too.
Sketching functional code (and/or pseudocode/stub-only functions) lays down a concrete hypothesis about what the code (or at least its key parts) should be. Writing docs fleshes out the aspirations and expectations (including helping you define non-goals and not-yets). Writing tests demands thinking through edge, corner, and other hard cases. Each informs one's understanding and intuition about the other two, and almost automagically drives gap analysis. "Whoops! Haven't thought about/implemented/documented/tested that yet! Let's go do that now! (or at least put it on the TODO list)" It's a strong tight loop.
Would also suggest you don't limit your "write tests" phase/work to pure "does just one thing" unit tests. You don't initially want full end-to-end tests that assume and require everything's working in order to start testing, but some of the tests can and probably should venture into the "requires multiple components interacting" space usually called "integration tests." I think of that happy middle ground / hybrid between pure unit and full integration tests as "functional tests" or "functional block tests" where the size/complexity of the functional blocks under test have more leeway than unit testing/TDD dogma usually allows.
Then I have to integrate a monster of a library that needs a whole battery of polyfills, and does it's own thing rendering modals somewhere, and I'm not in the mood anymore.
Mocking that whole thing? Hmnothanks.
I guess your case is something like this as well, that's why it's so hard to test.
Problem is, often you get an easy integration route, where lib and UI come in one package or a hard route, where you have to build a UI around the lib.
To save time, you use the full package, but then testing becomes a nightmare.
And the sad thing is, the stuff that's hard to test is crucial to test.
I guess in this case what matters is the estimated lenght of the product lifecycle.
I'm right now writing open source library in ts hoping that a ui will pick it up. I may need to write the UI myself at the end, but that will be much easier after I put every feature in my library I can think about :)
Test what can easily be tested. Architect your code such that business logic, etc is modular and testable. Don't worry too much about testing the hard stuff, especially if it's not likely to break.
Stuff that's easy to test is usually also easy to debug.
1) Helping you work through a complicated piece of logic.
2) Encoding some sort of requirement so future refactoring/bugfixes/features - perhaps written by a new developer - don't break that requirement.
3) When fixing a bug, ensuring the bug doesn't reoccur.
Tests that fall under (2) often feel the most useless, but I've found to be the most useful. They're typically the simple ones that don't feel like they need a test, but years down the line not every developer knows these requirements. Documentation is easily missed or ignored, but a test that's started failing? Sure there's still a chance they'll just remove/change/skip the test, but they can't just ignore or forget about it like with documentation.
Tests that fall under (3) are very similar to (2), except it's not an external requirement known from the start. These are ones that I've seen people occasionally write while they're fixing the bug, then remove afterwards so as not to clutter up the tests. Or do manually in the shell and never write a test in the first place (I'm definitely guilty of that). But whatever happened here was just complicated enough that the previous develop(ers) missed the conditions that caused the bug - so future changes to this part of the code have a good chance of reintroducing it or something similar. These are worth keeping.
Tests that fall under (1) are definitely useful while the code is being written, and typically people want to keep these because the logic is complicated (or even only write them because of the complicated logic, even if they didn't need it to write the code), but I'd say there's a further question here: How likely is this code to change (ever)? If you didn't write tests with the initial code, it all works, and it's something relatively generic that is unlikely to change... it might not be worth it. If it's likely to change it could end up falling under (2) or (3) in the future, so it might be worth a detour in writing the tests. If the tests already exist because you needed them for case (1), then it shouldn't hurt to just not delete them.
(I'm sure there's other purposes that don't fall under these three, but these are the main reasons for tests in my mind)
Stuff that has a test often doesn't need to be debugged.
Testing is already hard enough on its own even when following practices to make everything replaceable and mockable.
At some point I have taken the code so far apart that I am testing nothing useful.
Testing feels a bit like the sea shore problem.
The more detailed you measure a shoreline, the longer it gets, but what's the real length?
And this is not small companies we are talking about. They are banks, big educational institutes, etc. In the end the only thing most managers care is that the litle label in the pipeline dashboard that indicates the test percentage stais green.
Then you, or somebody near you, is doing it very wrong. I do not mean this as a moral judgment, I mean it as an engineering process diagnostic feedback. I literally can't count the number of times it has popped up a bug that I wouldn't have expected because of some change.
It is that very characteristic that makes me love them so much. No matter how carefully you code you can never get away from the problem of a small change over there causing a breaking over here because of something you couldn't even have anticipated, but you don't have to wait until some large-scale QA process or even production deployment to find out; you can find out 15 seconds later, and then fix it, or realize your new change is fundamentally untenable, or any number of things. I'd say "I don't know how people develop without these things", except I do; the code bases are treated like quick-set concrete and nothing can be changed once laid down. What a stultifying way to code. I would hate to work at a job like that.
IMHO thats the biggest problem with TDD. People often think of TDD as a set of requirements that all push for some set of benefits.
But it is really a development methodology where you as a developer heavily utilize tests to drive the development process itself. It is a mindset, not a checklist.
Perhaps easiest to explain (with an inaccurate comparison) as - it is like REPL-driven development but with persistent and sharable artifacts.
Behavioral driven development is IMHO stronger in this regards, because people tend to understand how the additional artifacts (beyond what you might see with a requirements list or use case document) are part of a methodology. People can better separate the methodology from the benefits/outcome.
The thing is, in the end most devs do shit tests becase theres no time allocated to that, and the test end up being just a number that needs to be met so the code could run to the dev ops engine and generte a new version of the software.
It's somewhat faster the second day.
It's a bit faster the third day.
It's probably net faster the fourth day, though by now you're developing noticeably more slowly.
It may still be net faster the fifth day.
It's net slower the sixth day, and the delta gets worse from that point on.
And I don't mean "per bug" on that, either. I mean, per project. By the sixth day of the project, you are net slower not writing any test code. And again let me emphasize net; by day six you are already losing overall.
Expecting test code to be written in the same amount of time as normal code is actually eminently reasonable, indeed it's the only sensible way to do it... if you do it right.
Also, you dont see it because those bugs were fixed before production.
This reminds me the saying of a manager arguing why do we need so many SREs since the system is working fine.
This was because it maintained an environment where I could replicate portions of the business logic (as assembled modules) outside a production environment. This made it much easier to do analysis/fuzzing of one component of the system, vs trying to replicate and do post-mortem crash analysis on the entire deployed system.
But I clarify this is TDD, not unit tests. A development methodology where unit tests are wedged in after the fact do not promote the sort of modular programming needed to be able to do this.
Oh man, those are my favorite. They're super difficult to critique in code review. "This test doesn't really test anything" isn't very helpful but writing a proper "do this" is often more work than writing a proper test yourself.
"This test doesn't do anything", then reject & request changes, "add substantive tests that cover the following invariants: ...".
Without these things, they're mostly theatre. With them, they're an incredibly valuable aspect of development that I don't intend to ever develop with ever again, by which I mean, if I do end up taking a job where this wasn't already in place, and I'm not allowed to put it into place, I will shortly be at a new job. Life is too short for the kinds of stupid debugging you have to put in to an untested, uncovered code base. (I don't mind debugging in general. It's a fact of the job. But that pull your hair out every single time, that's an uncovered code base.)
To most people what is important is the percentage of code tested, instead of test critical logic conditions for your software piece.
I work in a big bank in my country, that would insist us to do unit test in UI. Things like, if I set the button visibility to true, the button needs to be visible. Like... no shit Sherlock? What do you expect to happen?
Now, my strategy is to instead focus on writing a few solid end-to-end/integration tests. These tests often find just as many bugs/regressions, are actually testing the entire system, and much easier to maintain. Most bugs happen due to bad interactions between systems. I still write unit tests for some tricky code, it just isn't my first choice.
Tests are a business investment, of tech resources, to create business value.
It's good to focus on the Testing Pyramid [1]. High level tests are slow and brittle, but connect the low-level code to business features. Unit tests are fast and detail oriented.
In practice I write 1-2 high level tests (generally end to end, sometimes UI, or API/integration tests) to help focus development and have something that the business understands. Unit tests are helpful to "smooth the path forward" to ensure the code works as expected. Integration tests are great to iterate on, so that new code and tests actually work with real APIs correctly.
Tests are not free. However they create a lot of value -- they create (business and tech) confidence that the system is working as expected. Like you mentioned they assist refactoring, which makes the code much cheaper and easier to work with.
[1] https://martinfowler.com/articles/practical-test-pyramid.htm...
(disclaimer: writing a book about tech feedback loops, e.g. tests)
Covered code tells you nothing. It tells you this code may or may not have assertions made about it.
Tracking coverage is good because it shows you what code isnt tested. But once its covered, you have no idea. So instead of increasing coverage, you should be evaluating the new assertions being made to uncovered code. But virtually everywhere I've been has just used coverage going up to mean your tests are sufficient.
Code that you can easily write an automated test for generally is easier to maintain.
I find that unit tests actually make refactoring harder.
The problem is, unit tests need to be rewritten, or at least substantially shuffled around, every time you refactor. If you're making a "big bang" change for a major new feature, and that requires refactoring, then it's possible to justify that cost. But that's not how it normally happens.
Instead, a bunch of small changes over time usually make it apparent that the code would be a lot clearer and easier to maintain of the arrangement of responsibilities were different. It's already hard to ever justify the cost of that refactoring under any one small change. But add in the cost of updating all the unit tests and it becomes even more unlikely. Those updates also feels like demoralising busy work, whereas the actual change feels productive, so it also adds a human factor.
Overall, the net effect is that code with a lot unit tests ends up with a stagnant and confusing design.
As you say, a lot of the practical benefits can still be gained by end to end tests, while avoiding both the cost of writing so many tiny tests and the impact on long term design.
“But that’s TDD!”
“I dunno. I just wrote my tests first.”
“But you’re doing TDD wrong because you didn’t go and make it minimally pass before improving it!”
“I dunno. I just wrote my tests first.”
Doesn’t mean it’s not right to adopt a pattern when you identify that you’re implementing something similar to it. But you aren’t obligated to get trapped in its gravity well.
I guess I go by what Kent Beck, who coined the term said TDD meant.
https://www.goodreads.com/book/show/387190.Test_Driven_Devel...
This looks to me like writing tests around code that didn't have any before. It is crucial for refactoring to have tests; moreover, I would say that refactoring is impossible without tests. But it is not TDD, it is just adding tests where it should be in the first place.
> when you’re cobbling together a project from scratch, iterating and hacking away until something decent works, writing test code to throw it all away seems like a waste
This is exactly what TDD is for: to thoroughly think over the architecture and implement it right from the first take, without throwing away a lot of code. This is exactly the problem of people like me or author of the post: we just love coding as the process. Attention deficit greatly adds to that, making good architecture an impossible feat. It is easier for us to rewrite everything from scratch several times.
I like this same concept for top-down and bottom-up design.
Top down advocates might say if you focus on implementation concerns, you risk building what's easy to build instead of what's needed.
Bottom up advocates might say if you focus on what you'd like to have, you risk building something that just can't be implemented well when something else (that can) might have been fine instead.
I think, instead, you should ping pong between both. Start at the top, think about the high-level design you want, then work downward and see if it makes sense. Maybe even write some code to try out some ideas. (Like, if I use this design, I'm going to end up needing this database query. Can I make that run efficiently?) Then percolate those lower-level concerns and the results of that research upward and see what adjustments you can make at a top level. Then repeat until you find something that works well at all levels.
The important thing I find that helps with the mental overhead of testing is a long script of integration tests is better than nothing. Too often "TDD" is interpreted as creating perfect isolation of unit tests on each method -- which is fantastic! But is also a lot of work.
Worse is better. If just checking in some precondition data and loading it in and then sequentially running a script of automated tests that all have side-effects that could impact the next test is your MVP for a test framework, do that! Having a suite that you have to run fully end-to-end isn't great, but it's better than nothing.
When that frustrates you, refactor it into true isolated unit test.
I wrote about it a few months ago called it "1. Always go home with a broken build"
But, it's also just a tool. That has to be used for the right reasons and the right circumstances.
One unexpected use for them is to implement the "right way" to do something as a unit test. Instead of having to remind them of what we agreed on, making a test is a great investment, and not the typical use case for tests.
I describe my approach here[0].
I write tests and functionality together, and my tests are frequently harnesses; not "unit tests," as they are understood, today (In "the elder days," we used to call test harnesses "unit tests," but what did we know?).
[0] https://littlegreenviper.com/various/testing-harness-vs-unit...
Right. If you don't know what it needs to do, why in the HECK would you be writing any code at this point??
If you're just hacking / prototyping, there's no conceivable reason to involve tests.
In the precise place in code you want to resume work, add in free text whatever mental bookmark you need, the failing compilation will stick out well enough.
---------------------
1. Use a programming language with an expressive statical type-system
2. Write only type-signatures for everything that I want to do (using type-holes) and continously compile my code
3. Once I'm finished with #2 and it all compiles, I start implementing all the methods/functions
4. If I can't implement a method because I got the types wrong (e.g. my types say that I will return a number, but in some cases I cannot and have to return nothing or an error) then I change the type-signature and go back to #2
5. If I can implement a method but I'm not sure it will do the right thing (despite compiling) without running it manually to see if it works, then I write one or more unit tests for the method. If I believe that changing the method will easily break it I also write a test. Otherwise I write no tests.
6. Once finished I write some integration tests for what I deem necessary - both positive as well as negative tests.
7. Done
An example would be - oh wait, it just passed it, nevermind. (I wanted to quote a kind of numerical thinking test it fails at it but it just passed.)
If this had been a test previous GPT's failed at, but it was coded up in a unit test, then maybe more people would get it. At some point AI will just pass all the tests we throw at it. It would be nice to have a spreadsheet of tests to 1) know whether we're there 2) show people that we're there.
I've also heard it referred to as "weak TDD"
It's possibly the best testing-related Poe/trolling I've seen.
In a recent personal project I started using tests on a bit of a whim and found an unexpected benefit that I wanted to share.