Kent Beck: “I get paid for code that works, not for tests” (2013)
istacee.wordpress.com
istacee.wordpress.com
There are other techniques which can give similar confidence, but tests are the easiest one.
> I get paid for [code and tests] not for tests.
> I get paid for [code that works and is maintainable], not [more tests than are strictly necessary to achieve that goal]
I don't know this guy, so I can't speak for him in particular. But, in general, I wouldn't be surprised if the opposite was true: too many so-called software developers give "shipping" too much importance, leaving none for any other aspect of the job. Shipping is a feature, not the feature.
This is not some random guy, he is like the "father" of TDD. And that is why the guy that wrote the article thought it noteworthy to mention his quote.
If it came from someone else it would not be that important to make such a fuss about it. But when it comes from Kent Beck then it is worthy at least some discussion.
Literally the second sentence in the article.
> Kent Beck, respected authority, creator of Extreme Programming, TDD and writter of several great reference books, mainly at the great Addison-Wesley edition
Back in those days, there was a backlash against "big design up front", and very little respect (in general) for testing as a practice. Unit testing in it's modern form was reasonably rare.
After this Agile/TDD stuff caught on, many people ended up over-testing things. This is a pretty typical thing to do when you're learning about how much testing is sufficient. I've definitely done a good amount of this myself.
It can take a good deal of experience to know where to draw the testing lines in particular contexts. I think this blog post points at this specifically - that we should write high-value tests, and just enough of them. We also use feedback over the long-term to have heuristics of where we tend to have recurrent issues, so we can test a bit more in those areas.
Far from being focussed on "only shipping", he's underlining the fact that "just enough well-written tests" support working software - and that should be our focus, instead of thinking our job is to "write more tests" (or focus on test coverage, etc).
Unfortunately, there isn't a really great language that enforces working like this, while being simple enough to push onto a big team :/ (if there is, please let me know)
I've done a lot less statically-typed code than dynamic (mostly Python) in the last few years, but I was playing with Unity3d recently and had a chance to write a good deal of C# with Visual Studio.
I hit a hairy problem some time along and had to do a big refactor to support a new feature. I deleted a single line of code, followed an error trail for 10 minutes, and suddenly everything was just done.
It was a pretty interesting moment for me. I realized that statically-typed languages can really have the potential to be as or more productive than dynamically typed languages paired with a good enough IDE. (And Visual Studio with C# is about the best pairing you can get).
I've had to reconsider my thoughts on these things a bit. Sure, there are things that statically typed languages can make much harder to test or work with. (Want to stub an external provider? Okay, you're going to need an extra interface, then you'll need to create a new stub version that implements that. Want to read/write pretty-arbitrary JSON? Good luck). But there are other places where you get huge wins by bugs just disappearing by the boatload.
I'm still not on board with heavy OO/inheritance, and love the pattern of simpler struct-style constructs with just functions in functional programming, but the static typing can give a lot of wins.
I think something that gives an inherent advantage to OO languages in IDEs is that SomeThing.<tab autocomplete> makes a lot of sense and is easy to compute! I can take an object and know what I can do with it at a glance. I haven't seen a functional language with enough structure to support that simple feature yet (though maybe I'm just not looking hard enough). This is really where statically typed languages can make the most of an IDE. For some reason, that's the big thing I think of when I'm thinking about the downsides of functional programs I'm working with. The editors just seem to help a lot less (though I haven't written any FP professionally-speaking, so have less experience in general with tooling).
F# using Records with member methods looks like it might be able to get that sort of benefit though, I'll need to try that. It looks like they're just pure functions declared on immutable structs, which I think is the perfect middle ground.
A lot of object-oriented languages have taken tips from functional languages lately (map/reduce/filter is the new hottie), but I think there's a lot of benefit still to get in the opposite direction.
I'd point out, though, that is still is only part of the picture. You're holding on to a value and you want to know what you can do with it:
widget.<tab>
will show you everything of the form widget.someData
widget.doSomething()
but you're still missing out on other structures like freeFunction(widget)
handler.handleWidget(widget)
hammer + widget
widget[part]
// returns a widget, does this one count?
widgetFactory.buildWidget()
In nearly any language the first two will be common, and where available the others are critical usages, too. I want to be able to tab complete them!It's hard to continue hewing to the tab as activation with these other structures, which may be why IDEs and REPLs don't really try.
I actually dislike tab complete in usual forms most times. Auto import is nice, but i feel that auto complete is a form of searching the code base. And, when I am coding seems a poor tube to be searching for the answer.
When debugging, however, jump to symbol and quickly listing alternative methods helps. And sometimes I am just searching. So, good feature. Just not something i want to rely on.
F# (as well as OCaml) offers something similar in that you'll use a lot of functions that are within modules with the same name as the type you're working with. So you can write "List." and get a list of functions (map, reduce, etc.). I'd prefer something like Idris which will disambiguate functions based on the relevant types, but at least it makes IDE support easier.
> A lot of object-oriented languages have taken tips from functional languages lately (map/reduce/filter is the new hottie), but I think there's a lot of benefit still to get in the opposite direction.
Something in particular I wish F# would add it general non-linearity of definitions. All files and definitions in F# must be strictly ordered (either type A can reference type B or vice versa, but not both) except for specific, contiguous blocks. It presents a challenge for type inference, but I think just punting it back to you for the tricky cases would be fine (and it often has to do this anyway).
If you do need to get around it though, you can have mutually referential types in F# if you use the "and" keyword (although the definitions of the types have to be right next to one another). And in the next update to F# you'll be able to have mutually referential types and modules within the same file which is often good enough for most other things you might need that sort of thing for.[1]
[1] https://blogs.msdn.microsoft.com/dotnet/2016/07/25/a-peek-in...
Parsing JSON in Java was one of my worst programming experiences, so I have to agree with you here.
But in Rust, using the `rustc-serialize` library (and Serde, but I haven't tried that yet), parsing and writing JSON is really pretty painless. The really nice part is that you can declare the structure of your JSON data as a completely normal Rust struct (just with a derive annotation that makes it Encodable and/or Decodable), and with a single function call turn an instance of that struct into a JSON string. And in reverse, you can just parse() a string and it will return either an instance of your struct or an informative error if the JSON is malformed or doesn't match your structure. Makes JSON really easy to work with.
In golang recently I had to take some json (that I only knew part of the structure of), and modify just that small subpart of it without touching the rest. It was a really painful thing to develop, and the code ended up very messy.
There were a few golang libs for reading arbitrary json, but none supported writing to it that I could find.
Indeed, and C# isn't even a particularly safe typed language. It still has pervasive null, for instance. When you get into F#/OCaml/Haskell/Rust-type languages, it's a real eye opener.
> Want to read/write pretty-arbitrary JSON? Good luck
Not sure I see the problem. Just deserialize JSON into a JsonValue which provides dictionary semantics like JavaScript.
The break-and-follow-the-errors approach is very powerful. It's usually not hard to find the exact break that will show you all the things you need to change, and then just work through them. My record is 5 days without buildable code, working in C++; once I'd worked through all of the errors, the program worked, and without any non-obvious problems.
I miss this a lot when working in a dynamically typed language.
(Thing by Jonathan Blow that touches on this: https://web.archive.org/web/20140929232443/http://lerp.org/n...)
I'm surprised this myth perdures.
Reading arbitrary anything is trivial in a statically typed language: use a hash map.
There. You're merely emulating what a dynamically typed language gives you, of course, but it's trivial. And at least, statically typed languages give you the choice: you can be dynamic or static. You don't have such a choice when you don't have types.
Another is that some code reviewer asks for more tests. A third reason to spend time on tests is that they're required to maintain the same standard as the shipped code, even though they're only run in the presence of the developers, and their breaking only affects the developers, not any customers.
People forget the ultimate reason for our work oh so often.
I write tests, many tests, when I'm working in a dynamically typed programming language. I write tests even when I'm working in a soundly typed language. The only difference is that in soundly typed languages the type system guarantees many properties for me so I don't test for those.
Personally I like to write tests first but I don't believe that gives me any productivity benefits. It's just the way I think.
> People forget the ultimate reason for our work oh so often.
Tests are important because reliable code is important. The customers are important but so is the business. It costs quite a lot of money to support error-prone, poorly designed software. Tests aren't a silver bullet but they are a tool to alleviate the problem.
For example when I changed some code for which tests didn't exist, so I tested what I changed and wrote some extra tests while I was at it, and it was blocked in code review because my extra tests didn't report failure in any detail. What I did said "x failed" if a test failed, no details. The reviewer said much the same as you did now to justify that additional reporting was absolutely required.
It's a fine sentiment when it actually applies, and I wish it weren't applied quite to often to justify YAGNI and other rubbish.
The advice you received with regards to defect locality sounds reasonable - tests that don't give much in the way of isolation can cost a lot of developer time to hone in on. It's hard to say, not knowing the exact details however.
I also find it hard to reconcile the idea that "tests are an important tool for writing good software" with "justifying YAGNI". How would an "openness for developer testing" justify an attitude of "not writing things you don't need"? Those two concepts sound almost entirely orthogonal.
Specifically, if a test passes right away, then its error reporting isn't important today. It probably comes in useful if the test ever breaks, but will the test ever break? Therefore, spending significant time on the error reporting today is YAGNI, even if minimal version of the test is useful.
My complain is that even if unit tests are useful to a degree, people trot out the reasons for usefulness primarily when those reasons do NOT apply.
Not to mention it will literally take you 2 minutes to add the better reporting.
So if it takes you significant time to get proper defect locality, I think you should see if there's a better way to approach the problem. This should be essentially the default for typical/modern unit tests. Perhaps you're writing tests that are more like system level tests?
I'd also say that if you're essentially certain a test will never break, then (other than for documentation purposes) why are you writing it? To paraphrase Kent Beck - we should only test things that could possibly break.
You might be overgeneralizing what "people" say about unit tests - I'm not sure what your specific scenario is, but there are a whole spectrum of opinions on the subject. Perhaps this is just an organizational code-smell of the place you're mentioning.
Dogmatism/cargo culting in general can be annoying however.
Yuck. Imagine if you got a bug report that just said "X failed" with no details.
I write tests to check if my code works. And tests that document how the code is supposed to work currently are usually enough to prevent code from breaking in the future.
Anything related to privacy or security should be fully tested. But for the typical startup, I'd posit that beyond that test coverage should be more closely related to the number of users and level of usage rather than to the amount of code.
How often do you write code that's not related to privacy or security? As soon as you connect something to the Internet it's related to privacy and security.
The only situation where privacy and security don't matter a whole lot is if your code runs on airgapped devices with very limited tasks.
Sounds hard to believe. What kind of code would that be?
> There's a lot of code that's not written by web devs.
There's also a lot more than the web that has some form of connectivity with the Internet (even if it's not directly connected it may still parse data that comes from untrusted sources).
There is a widespread belief among many that "security is important, but doesn't matter for me". The most extreme example is obviously IoT ("Who would want to hack my coffeemaker?"), but there's a lot more. The unfortunate truth is: There is hardly any code these days that is not security relevant.
You do realize there's a bajillion non-networked apps in existence? Word processors, excel, editors/IDE's, system tools (esp monitoring/backup), media players that don't download stuff, compression libraries, MATLAB-style tools for numeric analysis, and so on.
All of them parse potentially untrusted inputs. They don't have to be directly network connected to be a security risk.
Just pick the first example: A word processor. It is not a security risk only if you can guarantee that you'll only ever open documents that you created yourself. If you ever use it to open documents you got from someone else it needs to take security into consideration.
Still not writing or patching a networked app.
"Started new area of research: software safety." That was Bob Barton in Burroughs B5000, Dijkstra on THE, and Margaret Hamilton on Apollo code. Maybe they mean first dept at MIT or just making status of sub-field more official. Then TCAS II. I recall reading that long ago as an exemplary work in formal specification & safety analysis but project was too heavy for me. Article says them too haha. Props to her for it & others. Article shares my view on scattered groups & methods. At least seen STAMP referenced once but unfamiliar with it. They wrote against N-version programming being re-invented... which I proposed for subversion resistance. Hmmm. I'm sure my variant is the one that works this time. ;) Also did SpecTRM at their company that looks a lot like state machine and modeling schemes I saw elsewhere in high-assurance. Not claiming a copy rather than inspiration or independent invention + convergence of multiple parties. Usually means a good idea.
Very interesting person. Thanks for the tip. Your sister is going to learn some wise things for sure given they've got sane methods and got results before. I especially liked how the article jokes about writing what she knew on high-assurance development then gotten wiser or more confused. I know the feeling where I'm redoing the foundations now with what a decade taught me. More slowly this time given I have more doubt than certainty.
Note: Just got to the last part. Wait, she was the one who wrote the THERAC paper? I just assumed it was some guy (male-dominated field) named Levenson since that name was all I saw in references to the report. Never saw it again. So, she wrote up an investigation we've been citing about software safety for decades, helped spearhead efforts to legitimize it as a field, did huge projects, and I basically never hear about her. Unreal. I'm bookmarking her stuff to go through it later.
> Sounds hard to believe. What kind of code would that be?
Device drivers, compilers, and some embedded systems come to mind immediately, there are plenty of others out there. I've worked on a lots of software where the only inputs were physical and sensor based, and the only outputs were to the screen. Device didn't even physically have network equipment.
There are few things where security doesn't matter. But they are extremely rare. The situations where programmers think their code isn't security sensitive are probably vastly more common.
I hope you're joking but I suspect you're not. Device drivers have the highest level of privileged access in many operating systems and code quality for drivers is so uniformly lousy (certain large vendors whose names begin with "N", "A", and particularly "Q", I'm looking at you) that attempting to break the drivers would be among the first things I'd consider if I were trying to root a device.
I jump up 1000 levels of abstraction from time to time, and when I do, I agree that security is extremely important (FDA class III device and HIPAA compliance is mandatory.) I'm also a lead, so I have to know enough to call BS when I hear it from a team member.
A good example is that I care a lot about making sure our search endpoint doesn't return private user data. But beyond that, I'd rather just know that the endpoint returns a 200 and let someone tell us if it's broken rather than have an extra three hundred lines of code to see if it's returning the correct results. If we get a ton of users then that will probably change, but for now the cost of writing and maintaining those extra tests wouldn't be worth the benefit.
That used to be true for the software in cars, but it no longer is. The problems that result are not the fault of the original authors; that belongs to the people who decided to bridge the airgap without thinking through the consequences.
For me the most important word here is 'supposed'.
All the time I read documentation describe how code works step by step (what each 'if' does, but spelled out more verbose). And test that only test that a function does by mocking out everything else.
But I don't care reading what code does. I can see that by looking at the code. I want to know what the developer intended/expected the code to do, so I can validate it against what the code actually does. Most of the time assumptions are made with those expectations. And with those expectations you have a much better idea why a trivial refactor of a piece of logic could unearth a massive 'undocumented feature'.
Why do you think that is the wrong approach?
When something doesn't work as expected, I now check my design and not the code.
Or you're tired. Or you're just not as focused as you could be. Or you have a deadline. Or you're trying something new. Or you're not fluent in the language yet. Or you're fluent in the language but not the framework. Or you're fluent in the language and framework but not the design pattern (if you use those).
Or a whole host of other things.
I'm not saying spot-checking the design isn't a good idea, but saying that it's the design more often than the code just doesn't match up with my experience.
Based on a true story
Out of curiosity, did you use any debugging on paper for your code?
1. Write tests to prove your code works, which is sometimes referred to as "test-driven development"
2. Write tests to catch any side effects or regressions when altering code
Funny thing is that if you write good tests, then the results are pretty much the same regardless of your motivation.
1. Have an overall idea of what your software will do before writing the first line of code.
2. Challenge and change any touched assumption from #1 during development when you refine that idea.
3. Test that the program satisfies your refined idea after it's written.
4. Create some assurance you'll keep #3 correct while you write any further code later.
However you fulfill those needs, if you got them all, you are good.
Automated testing only shows its real value when you go back to change code that was working before. With the manual approach you'd have to retest everything to have any real confidence. With the automated approach you just run a command.
I'm a big fan of automated testing, but if I didn't expect to have to ever change code, I wouldn't bother with it.
> With the manual approach you'd have to retest everything to have any real confidence.
This kind of implies that software development before "~tdd" was a complete disorganized gong show of quality especially where refactoring, but in fact that was not the case. There are ways of coding that are more conducive to quality than others.
> With the automated approach you just run a command.
Once you've written all that code, yes.
> if I didn't expect to have to ever change code, I wouldn't bother with it.
Usually you don't, in which case the extra effort on testing is wasted, usually.
The typical claim is that code maintenance is at least 10x as long as the initial write.
The question isn't "do we need to test." Testing can just take the form of running the code manually and making sure the output makes sense, but you do need to test.
The question is "do we test automatically or manually." It's the same question regardless of how good your practices are. Note that this is totally distinct from the question of whether or not to use TDD.
> Once you've written all that code, yes.
I've found that writing tests often doesn't take much longer than testing manually, and rarely takes longer than testing manually twice. Sure, if you obsessively try to test every possible case and input, you'll waste time, but well targeted testing doesn't have to be slow.
> Usually you don't, in which case the extra effort on testing is wasted, usually.
For non-trivial projects I have an average number of revisions per line of code much closer to 2 than to 1. Sure, some of the code only gets written once and never touched, but other code gets revised multiple times. If you're good at writing tests, the tests will focus on those often revised lines of code.
And, again, you have to compare the effort against the effort of manual testing, not against the effort of writing code you've never run and shipping it.
end to end tests, integration tests, regression tests, etc, yeah.
Unit tests though...usually no. Often if you're making any kind of significant change, the entire code paths may get refactored away or change too much and the test will get nuked anyway.
And it's that kind of test that usually confuses people, so it's worth understanding.
A unit test's goals are many:
It proves at authoring time that the code works. It saves you the time right away of having to go through the UI or spinning up a server just to validate a function is working. It proves that you thought about specific edge cases (and if you do testing consistently, the lack of test is your evidence of unconsidered edge cases. It's documentation of all of the things you considered when writing the code. It is an example of how to use the code with all it's use case.
And when you nuke a piece of code away, the failing tests are now a guide of all the cases you have to make sure are truly no longer necessary.
If, in the future, you do a refactor of an implementation detail (so the existing tests are still valid), then that's bonus as you get green/red validation. But in practice, that is less common than all of the other reasons for testing. That's why the "I don't expect this to change" thing isn't really a reason for or against writing tests.
That pretty accurately describes my experience with unit tests. I've been part of several projects where we had pretty comprehensive unit tests (I'd say small to medium sized projects) and I never managed to get as much use out of unit tests as I liked. After one or two big refactors most of the tests needed to be, as you said, nuked anyway. While seeing all the pretty green lights is reassuring, they are rarely working when you most need them - during large refactors which blows away big sections of code.
I picture unit tests as a row of black boxes sitting on a table in certain positions. Unit tests are great when you don't move the black boxes but do change the mysterious processes are running inside them. But refactoring is rarely ever that isolated in programming since you tend to move some of the black boxes around, remove some entirely, add some new ones, change the contents. To then expect the unit tests to give you back useful information on whats broken is rarely possible.
I've had more success with e2e testing using things like Selenium, but it's still frustrating as a developer to read articles about how great unit testing is (like this root comment) and never able to actually get a decent working version of it in your projects (because of the reasons mentioned by this parent comment).
Building up small oases of dumber/more verbose code that has unit tests seems like the best way to wrangle legacy code into something that anyone else on the team can understand and not mess up 6 months from now when it's their turn to have to touch it for the first time. Of course no one wants to touch that 200-line, 10-levels-of-indentation monster method, but bringing just a little bit more and more of it under unit test over time will help a lot. Every other benefit of unit testing that you listed besides the time saving aspect (e.g. why launch a big e2e test if you can test the same thing in an xunit-like context? even if it's not strictly a "unit" test) pales in comparison of the benefit of making crap code nicer to work with. A corollary is that if your code is already nice to work with, and you have a system to keep it that way, unit tests won't be very valuable. (Though other tests, which may or may not be in an xunit-like context, may still be quite valuable.)
"Because you should write code that works in the first place!" -- your boss
Reality, on the other hand, is usually suboptimal.
You should be testing input/output and results as opposed to testing how the internal gubbins works. That's the line we have to carefully tread when making a test. The test shouldn't force the item to behave in directly the way it expects; more that the I/O is correct.
it depends how brittle your tests are. You can write tests that make sure internal stuff works at a unit level without them being so brittle
to your point testing I/O or behavior is the way to accomplish this, but it can still be done at the level of internal functions/methods
Unit tests should be used for extremely small and isolated mission critical objects while functional tests should generally cover the entirety of the I/O chain. That's how I do it at least and it works extremely well for a fraction of the cost!
This is in no way a benefit unique to unit tests.
If your project breaks because of local changes I think regression tests with real data and bisecting is better and less work though.
1. Your unit tests fail basically every time anything changes. This is the scenario where your unit test is something like "the command line arguments are -abcd" and every time you add one you need to change the test. This makes the unit test worse than useless, but actually a source of extra work every time you change something.
2. Your unit test never fails. It just doesn't fail ever, at all, under any circumstance. It's so obvious that it should work, but someone wrote that test anyway. It's a waste to run it every time.
3. Your unit test fails when you refactor because it tested some internal functionality. You need to throw away your unit test every time you refactor. It's a waste to write one every time.
The only tests that ever show that a refactor broke something are integration tests. The 200+ unit tests in my project either NEVER fail. Except for that one that you have to keep changing every time.
You don't want tests for that - you want a type system and a static analyzer.
However, many in the industry forget that this is the underlying reasoning, as seen just two days ago here on this site. Read through the top-rated comments on this post: https://news.ycombinator.com/item?id=13119138
They describe an emergency situation where a single "3" needed to be changed to "4" ASAP or people would lose their jobs, and everyone's applauding the gatekeepers who insisted on significant refactoring and the creation of additional tests before the change could be approved.
I agree with those who say those improvements should maybe have been demanded immediately after the fire was out, but those who would have delayed the firefighting out of blind allegiance to the rules seemed, to me, to have forgotten that the rules are there to serve the programmers (particularly, their ability to quickly ship working code), and not the other way around.
A rule that's failing to do that should be changed or ignored.
But if people's jobs were truly on the line, I'm inclined to agree with the "screw it, push it through" approach.
If you are hired as a programmer, then yes, just do whatever we ask of you.
But if you are hired as an engineer, everything the business asks of you comes with an implicit: "and make sure it's done in a proper way that won't break anything, or slow us down, or cost us too much, or limit our ability to gain a competitive edge."
You don't just change a 3 to a 4 because the CEO wants you to. You have to make sure the change doesn't come with unforseen impact that would put the company at risk, and you have to make the change in a similar way. That's what the CEO expects also. If you did the change, and it had caused impact to the business, that you had not pointed out, and for which the business believe is more harmful then having waited a few more days, you and only you are to blame, and you will be. You can't say, but CEO told me, you're the engineer, you're the person they hired to know this stuff and prevent these issues from happening, not the CEO.
But if Scotty delivers on time at the cost of overloading an expensive piece of equipment that, after the battle is won, requires a week in drydock to replace, that's probably a successful execution of exactly the kind of call a senior engineering manager is expected to make.
In your example, Scotty knew what he was doing though. He didn't say, wow, what Kirk wants me to do could kill twenty redshirts in the process, I'll just take the gamble since he seems to want me to. He knew exactly the impact, and made it knowing he would easily be able to contain it.
Which is often not the case in Software and in practice. You have to do something to know the impact, because most problem we solve is always new. Its not something we did many times before. If that variable was often changed, then it would be completely different, because he'd known, just like Scotty, that its something they can do. In that case you can make the choice to say, lets change it, and later handle the tech dept of the less maintainable code.
Also, in software, its almost never the case that people can't wait a few more days.
In the example with the line of code that took 6 days, there was no dramatic emergency in production that required cutting corners. If it had been an emergency, of course the code refactoring demanded in code review should have been postponed; those changes increased the impact of the change, and therefore the resting requirements.
And if it really is an emergence that requires people to drop what they're doing, then someone with sufficient authority should be directly involved in order to override all the usual procedures.
But you don't just drop all procedures just because somebody claims somebody said something. That would be dangerously irresponsible.
If the CEO lacks the understanding of the technical consequences of a change that may blow up the company, the engineer should make the decision. If the engineer lacks understanding of the business consequences of not making the change - like losing an important client, or suffering a wave of negative PR, or facing a lawsuit - then the CEO should make it. Ideally, both sides should be communicating these consequences so that both of them have all the relevant information and would ideally make the same decision. Then the decisions can get made at the lowest level that has all this information, and the CEO doesn't have to get involved.
In practice, there are many cases where the CEO can't communicate all of the relevant business realities, eg. if you're facing a lawsuit if you don't make a change, it's often better not to worry the rest of the staff or make them subject to depositions, and simply to ensure that the change gets made. That's why the CEO is the decider by default in organizations, and also why it's usually expected that employees will obey a direct order from the CEO or be fired.
This is a distinction that is a very thin line and most people with "engineer" in their title would not sign up for.
It all my years (15) of professional experience, I've only worked under one PE, an Electrical Engineer.
In some states you literally can not have "engineer" in your job title unless you have a PE certificate/accreditation/whatever.
In my opinion, computer science and engineering is not about just making the code you're told, it's about questioning whether that code needs to be done in the first place, and if so, how.
It still depends.
David: It's for Philip. It we don't do this right away, we'll have to have a layoff.
and
Judy: OK, then I'll fill out that section myself and put this on the fast track. ----- 2 days later. ----- David: What's the status of 129281?
"It we don't do this right away, we'll have to have a layoff" and "2 days later", making an impression of this is an "a day or two" task, not a "fire/emergency"
I don't think if this change finished in 2 days anyone will be unhappy. But if you took this to production, and somehow failed, everyone would blame QA/testing
In theory, could informed, intelligent, rational actors without ulterior motive do so, sure. I'd sleep on a couch and eat ramen for the chance to be part of such a team, but I haven't met them.
The engineer did their due diligence, wrote tests to make sure it had the desired behaviour, and got it done in minimal time. Clearing technical debt in old modules should be done, I agree, but not while trying to put out fires. It adds considerable risk to a change which should not have any impact except for the request.
"We can do that, but it will take us 6 days, otherwise we risk taking the plant down and aggravating the issue."
I wonder if the CEO would have just said ok thanks.
In my experience that's the case. The engineer in that link got himself in a bad spot, because he didn't know what was involved for the change when he communicated his estimate. And most of his back and forth that slowed him down would have been avoided had he known beforehand how to properly do it. Even with everyone's feedback it sounds like a 1 day code change. That seems to me like the reason the change was slow is more ramp up time for him working on a code base he doesn't normally work on.
Then the boss can make an informed decision.
The important thing here is to provide the risk-assessed alternative in writing. This covers the asses of both sides! If something blows up and management was not warned about the possibility, they're well within their rights to rake the engineers over the coals for it. But if engineers warn management of the potential consequences and management chooses to take the risk, if it blows up, you have your CYA right there - they were warned in writing.
If you can't trust your manager to take responsibility for the decision, then it's better to make a decision you're going to be held accountable for than to let your manager make the decision and then hold you accountable.
I've seen other engineers in the position where they give a risk analysis and warning in writing, and then shit hits the fan and they get fired. Maybe it doesn't happen immediately, maybe the reason they were fired isn't explicit, but the change in the manager's attitude toward the engineer traces back to when they did what the manager said.
There are also mangers who won't follow through on the tech debt part because they don't trust their engineers even if their engineers are trustworthy. When they discover that they can bypass testing by pulling the emergency lever, they'll start pulling it all the time because they see it as a way to get what those lazy engineers to do their jobs faster. And when tech debt catches up with you, bugs abound, and development slows to a crawl, the engineers get blamed.
Maybe you have a boss you can trust to take responsibility for their risky decisions. Maybe your boss trusts you when you say that paying down tech debt is necessary. But maybe your boss and their boss don't have the same relationship, and the shit rolls down hill.
Yes, I want to work in a trusting environment where my interests are aligned with doing what's best for my company, but an at-will employment capitalist economy doesn't always work that way. It's every man for himself at a fundamental level, and exceptions to that are too rare to make a blanket claim that people should just do what's best for their company.
I work 35 hours a week, make $125k/year, have good benefits, like most of my coworkers, and am doing relatively interesting work for a company that isn't completely evil. Sure, my boss can scapegoat me if I let him make technical decisions, but I don't. That's a tradeoff I'm okay with. I'm always on the lookout for something better but I'm happy where I am.
And besides, most people who think their bosses won't scapegoat them if something goes wrong are naive. If it's your job or theirs and they have the power, you're screwed. It's better to avoid the risk and not put yourself in that situation.
To be clear, I'm not saying don't take risks. I'm saying wroth risks on your own behalf.
What economy does work that way?
Trade offs that are business related can be delegated and should be delegated to the business to make. But saying something like: "if we do A, it could make us vulnerable to X,Y,Z technical problems, but would allow you to have what your business needs 7 days early" can be dangerous, because it assumes the business stakeholders truly understands the dangers of technical issues X,Y and Z, which they almost always do not.
As an engineer, you should advise the business towards the proper way to do things, and you are responsible to make sure your advice is clear and loud. It is sometimes appropriate to suggest the less ideal scenario, but if you do, be ready to assume full responsibility and ownership on it.
You should never have an advice that goes like: "Yes we can do it, but it'll cause other problems." and expect for these other problems to not be your responsibility to fix and mitigate as they appear. The implicit of all your suggestions is always: "I can make this work." So better be sure whatever you suggest you can actually have it work.
Plan C is what the Strategy guys always want to go for, because they can't recognize when the repucussions happen. They just tell themselves it's those good for nothing engineers screwing up again.
But too often there's been one guy ithe engineering group who values accolades over stability and will offer to be a hero, only to wander off without finishing anything. Later in life I've realized I should have suffered through more of the issues at companies where the team at least had solidarity. I knew I wanted that in a team, but didn't know that I needed it.
Plan C doesn't differ from Plan B in terms of technical approach. It differs in terms of the consequences of failure - not the technical consequences, but the political consequences.
My first really interesting project that led me down the sort of DevOps/Agile trail was in the mid-1990s, when I helped write a risk management and contingency planning process for the mid-sized company where I worked. One of the features of the process was that management could not reject a risk analysis on any grounds except incompleteness. They couldn't say "I don't want that written down!" They had to sign off on it, in writing.
This turned out to be very popular with both teams and management. Teams felt they were finally getting the opportunity to cover their asses, and management finally felt like they were getting straight answers from the teams about actual risks. In the earlier lack of process, teams were at least passively discouraged from honest, written risk analysis, as naysayers who were resisting business opportunities. Worse, if a warning was given and then the problem happened, the people who warned were blamed for not preventing the problem!
When we beta-tested the analysis process, the first project that tried it actually turned down business because it was too risky. That had never happened before. We probably saved the company many millions.
Later...never arrived for tomorrow is another Very Important Thing that must be done now!
Somehow management continues to forget the critical resource: time to do things right
Eventually the software engineering department has a code base that is riddled with tech debt, to a point where changing a 3 into a 4 ACTUALLY takes 6 days.
Then management asks "WTF? What did you do? Why does it take 6 days to turn a 3 into a 4?", next comes frustration and SWEs leave the company.
I've seen this happen at 3 different companies in 3 different industries, a huge company (hundreds of devs), a medium company (50 devs), and a small company (2 devs) [1].
Every time it happened, it was because a SWE team was constantly delivering under pressure by management that disregarded the cost of writing code "that works" without ever addressing tech-debt, no matter how much the SWEs warned management.
[1] To be fair the small company only fell into that anti-pattern because neither the engineers nor the managers knew any better.
By all means, if you have additional resources, invest them in refactoring and improvements and additional testing and continuous integration and all those things that make our days enjoyable and our product quality high. But your first priority has to be making sure that your product solves your customer's problem.
At some point organisations forget that the process is there to serve us, and not us to serve the process. Moving variables to parameter files, renaming legacy variables, these all seem like much more risky things than simply changing the value of the variable.
I think it depends on the type of product you are working on. Certain domains require very strict adherence to policy - for very good reason. Just because your shiny new aircraft's deadline is tomorrow, doesn't mean you can skimp on the required testing.
A fire is something like, "the site is down for XX% of customers!" (for a large value of XX) or the "the software is routing product to the wrong place!" No doubt a real fire would have been handled differently.
I certainly didn't applaud the gatekeeper because I believe they opened themselves up to a great risk. If you are doing an emergency patch to production you should minimize the amount of changes. All the refactoring of code was an unnecessary, reckless risk. The refactoring should move to the next scheduled release as the top priority since part of it is already in production.
If the input data is complex enough to drive a non-trivial branching logic in our code - yes we do. Well, at least I do. The bonus is that it provides the fixed point for further changes at the same time.
- for catching regressions in the future
- to verify that your code works. How else would you know that your backend service reacts properly on some edge condition
- as documentation
- as a kind of quality mark, for instance to be able to pass a code review or when writing code for an external party
Unfortunately the last reason is also the most useless and still it is the one that seems to be the main motivation for many big enterprise developers.
Thank you for including this - it's underappreciated. So many times recently I've needed to write code against some poorly documented API (not always due to lack of effort, some things are hard to document well in prose / JavaDocs, or just due to constant change), but I took one look at the unit tests and it all made sense - and I knew it was up to date for the latest work.
OTOH, projects with excellent end-to-end tests tend to have lower coverage but better regression management.
I think this is the point of the article?
I think that full test coverage is nor sufficient or necessary for the working software. What matters is the sufficient covering of the edge cases.
I think that an additional point to keep in mind is that most probably Kent Beck talked mostly from the point of statically typed languages.
"We" write comments and use "plain" language documentation, or formatted but otherwise plain language for documentation generators to parse to document our code, and that includes "what the code does."
I'd argue that if your test code is what you're relying on to "document what the code does" then you are probably (almost certainly) over-testing, testing the wrong things, or some combination. Oh, and also using test code incorrectly (as documentation).
Which works great as long as your changes are shallow. If your changes aren't shallow then you have to change your tests as well and that defeats the purpose.
Automated testing is good for freezing an interface and it's behavior; allowing you to change implementation while maintaining the same outputs. Lots of technology requires this: Networked API endpoints, libraries, etc.
But most change I encounter is from changes in requirements that necessarily requires reworking and re-arranging code that won't be compatible with the test suite.
I've had to completely redesign entire subsystems; break down components into different pieces; move code between different layers; etc.
Honestly as a manager nothing bothers me more than developers who try and patch complex changes into existing systems without re-thinking how it affects everything else. I have to re-factor a project that's a total mess because code was added but not removed or changed over the course of some very big requirements changes. I'm sure all the tests pass but it's impossible to follow now.
The problem with unit tests is it adds an extra layer of friction on making changes that benefit the product. You are actively discouraged from changing your design from your initial assumptions! This change friction can be seen as a benefit if you need all your interfaces to be stable (like with a library). But it's a trade off and it's not appropriate everywhere.
I'm not saying anything about bad architecture. You have a good architecture for today's requirements that is a bad architecture for next year's requirements. Tests lock you down to whatever your first architecture is.
I prefer automated tests of larger units or subsystems, and testing against requirements, protocols and specifications rater than implementation details would change. Of course anything will change over time, but I believe this kind of tests provides higher value and lives longer.
So what methods do you use for that purpose?
My tests and practices are, occasionally, like a five point harness in a race car. Without it anchoring the driver in front of the wheel, they could never ever drive that fast with any safety at all. This one comes out as a counter to the 'straightjacket' argument some people like when the tools won't let them just write shit code that the rest of us have to babysit while they run off somewhere else to write more shit code. Which we are supposed to be grateful for... I digress.
But most of the time my tests are more like smoke detectors. The ones that go off for no reason get replaced or removed. The ones that go off after I can see fire and smell smoke? Why am I paying for upkeep on those exactly? What's the value?
But when changes can come from anywhere (lots of developers potentially making changes) and the need for changes can be urgent, there's a great deal of value in complete code coverage.
Not every developer thinks the same or has the same tendency for errors. I would caution people against thinking about tests as personal verification. With any luck, you'll be handing off that code base to someone else in time.
Do you allow your juniors to use their sense? No, you make them hit 100% coverage until they start to learn.
As Paul McCartney said: "Learn the rules like a pro, so you can break them like an artist."
Testing makes a lot of sense if you're doing data munging, but for front end or other kinds of code that are almost defined by their ability to create side effects, it's usually a waste of time. In my humble.
How often do you find, despite having written tests, that there is some bug in your software? And how many of those times did you think that you should have considered it beforehand, rather than that it would be impossible to foresee?
In my experience the most useful tests are the ones that came from some unforeseen bug, which was then fixed and a test case built around it, so that it wouldn't get "unfixed".
The least useful tests are the ones for cases you know not to invoke, because they are obvious. Like how you know when you divide by a variable, you know it can't be zero. So you make sure it can't be zero, making the test case a bit moot.
This was from some TDD course that my employer brought in where they emphasized that consts that are part of a specification or requirement should be tested against to prevent them from being accidentally changed or changed without full consideration of side effects/review against the program spec.
Is this what you're referring to?
"Oh yeah, that's going to need to be thread safe."
"Oh yeah, that might be nil. Rather frequently."
In my experience, testing has been much more beneficial as an exercise of my mental model of the code and its interaction than for refactoring. But to that end, I can think of quite a few times that I've been very grateful for a unit test suite while I did a large refactoring.
It's worth it's weight in gold when you can flesh out edge cases and logic issues before going all-in on an architecture, and being able to refactor without fear is a great bonus, but it's not always applicable.
Unless the majority of the tests are "external" (i.e. testing an API by calling it's endpoints and testing the return), a major refactor is going to require updated tests, and it's all too common that I see people take the "easy" way out and fix the tests to conform to the way it's working after the refactor rather than the way it should work.
> tread safety and race conditions
I feel like this comment wouldn't be out of place in an F1 thread.If you are, then I am as well, because this is where a huge part of the value of tests come from - validating assumptions early and quickly and cheaply.
For most of what I do (basic line-of-business type webapps & servers) the vast majority of the code works on the first try and is obvious enough that I never break it; during development there tends to be a few sections that break repeatedly while I'm changing other things, and those are the sections I write tests for. (Generally this is a few percent of the entire codebase.)
On the other hand, more 'computer-sciencey' projects tend to need a correspondingly larger proportion of test cases. The most dramatic example of this was when I built a Lisp interpreter as a side project: this was the only software I've ever written fully test-first, and I don't think I would have been able to get it working at all without a full test suite that hit every line of code at least twice.
- Write a few tests that fail
- Write code until your tests stop failing
- Repeat until the program is done?
In my experience, problems that require writing the tests first normally require writing all the tests first. If you start solving only a subset of them, you will have to rewrite most of the code once you start looking for the other tests.
Yes, that's what I meant by "fully test-first". I was following along with Paul Graham's The Roots of Lisp[1], so I had a convenient set of pre-made test cases. I would start by copying his examples into a test case:
it('exec atom on a symbol is true', () => {
var symbol = Symbols.get_symbol_named('foo')
expect(exec(empty_scope, [natives.atom, [natives.quote, symbol]])).to.equal(true)
})
it('exec atom on a non-empty list is false', () => {
expect(exec(empty_scope, [natives.atom, [natives.quote, ['a']]])).to.equal(false)
})
it('exec atom on an empty list is true', () => {
expect(exec(empty_scope, [natives.atom, [natives.quote, []]])).to.equal(true)
})
And then write my code: natives.atom = new Native(function (item) {
var i = exec(this, item)
return (i instanceof Array && i.length == 0) || Symbols.is_symbol(i)
})
Once I got through the paper that way, I had a complete Lisp interpreter working, and wrote a few small programs in it with no trouble at all.> In my experience, problems that require writing the tests first normally require writing all the tests first. If you start solving only a subset of them, you will have to rewrite most of the code once you start looking for the other tests.
This has been my experience as well. As I say, this is the only program I've ever written fully test-first; I think it was only possible because I already knew exactly what the inputs and outputs were going to be (having them well-specified by virtue of being a Lisp) and which ones I needed to implement to be able to get later ones working (with the help of the paper).
Generally I find it's not possible to work this way; my usual approach for solving problems where the code isn't immediately obvious is to write (and generally rewrite) in parallel:
- A set of usage examples and/or documentation, to clarify what exactly I'm trying to do and make sure the interface makes sense.
- The actual implementation (or, in early stages, some pseudocode).
- The tests (if any), to check that what I've written actually works. Often these are the same as the usage examples.
Frequently questions raised while writing one of these will influcence the others, often significantly; often I'll only think of an edge case while I'm writing the code and then have to go back and add it to the tests, or realize that the code could be simplified by changing the way the API works.
(it (atom 'foo) t)
If it fails, it just says something like: test failed: (atom 'foo) returned nil; expected t.
So no need to have a string there. Save the "blub-level" testing for things that are not testable through Lisp. (Or, really, things you are justified in not wanting to expose such that they are testable.)(Sure, it could be the case that both the interpreter are wrong, and the atom function are wrong such that the test passes. That level of breakage isn't likely going to pass much of a significantly detailed test suite.)
If this had gone any further than the three days of off-hours time I spent, I certainly would have rewritten the test suite in Lisp; however it was primarily an academic exercise to really 'get' how Lisps worked and push my comfort level; writing all of my tests in the language I was trying to write would have pushed things a bit to far.
In fact, I never got around to writing a parser, so even if I'd written them in Lisp it would have looked like this:
var Symbols = require('../Symbols')
var natives = require('../natives')
var foo = Symbols.get_symbol_named('foo')
module.exports = [
natives.tests,
[natives.it, [natives.atom, [natives.quote, foo]], true],
[natives.it, [natives.atom, [natives.quote, []]], true],
]I do think of myself as using TDD, but I use the tests to drive the model and often end up not needing the test - a bit like https://spin.atomicobject.com/2014/12/09/typed-language-tdd-... . With experience, more and more I'll skip past the writing a test and deleting it after, and go straight to encoding the properties that I actually need in my types.
The greater value of the tests is later on, when you need to make a change, and want to know whether you broke some existing functionality. Or, once you have a framework of tests around something, being able to quickly write a test that fails due to a discovered bug, and knowing that you fixed it by passing the test.
Maybe you know not to invoke them because you wrote those features, but what about the rest of the team? Or the poor schmuck who takes over your codebase after you leave?
Test cases help future-proof your code against the unknowns and the unexpected. I think of them as guard rails on a path. When the path is on straight terrain, they aren't necessary. When it is crossing the side of a cliff though, you're glad they're there.
That said, many of the feature requests from the ops' side end up taking an extremely long time to produce, if ever. Very basic requests that would save the team dozens of hours per week go on and on longer than they should. In these sorts of shops, the ops side ends up pretty miserable and builds up a lot of frustration for the development team, which drives the two groups apart.
More recently, my team's been asked the ops team is being asked to write tests for ops' prod bugfix automation. Categorically speaking, we're not developers, so it tends to take us longer to write tests than the devs who write them daily. Issues pile up and span across months, rather than days. The reasoning behind the tests is understandable since it helps build assurance that our scripts won't break the world. Though, the ops team is battle hardened enough to know that if we break it, we have to fix it - it's in our best interest to write code that's reliable as possible. I don't think any of us would be in our roles if we didn't understand it. Frustrating stuff.
one kind is the acceptance test, where you write a test that specifies correct behavior under normal circumstances (including expected error scenarios).
the other kind is a regression test, where you reproduce a discovered bug in a test environment, then fix it, and make sure that the test covers the fixed code to confirm that the bug is no longer occurring.
both types of tests are important and for different reasons. they aren't mutually exclusive and they aren't addressing the same problem.
I started my career writing iOS clients, and the obsession with TDD was baffling. 80% of my code was usually either UI or simple Core Data manipulations, while the last 20% was mostly API parsing and a touch of business logic. I wrote a few tests for parsing corner cases or business logic, but they never really gave me any confidence or helped with refactoring, instead taking up time and adding overhead whenever I made changes. I supposed I didn't have enough coverage to get the benefits, but what tests would I write for my UI? What tests would I write for simple Core Data queries (which is assuredly unit tested already)? What tests would I write for my parsing libraries (which are already unit tested)?
Then I started working on the (Python + Flask) API backend, and tests were self-evidently necessary. Python is dynamically typed, which can result in lots of corner cases when doing simple data manipulations. Python is interpreted, which means the compiler/IDE won't warn you about syntax issues without running the code, and you can't catch even the simplest logic errors without running the function. Most importantly, the API's entire job is translating data, inputs are in the API parameters or database, and output is the JSON. It's a perfect function, and tests were obvious. I wrote something like 600 in a week, then used them to make some major refactors with confidence.
What I learned from those juxtapositions was that unit tests and automation are invaluable in certain circumstances. Specifically _any system that creates machine-readable output_ like JSON, populating a database, or even a non-trivial object factory should be unit tested like crazy. Any system that creates human-readable output, like views or changes in an unreachable database (something like an external API or a bank account) needs to be human-tested, there's just no way around it.
People tend to be bad at forecasting these kinds of costs, its easy for prejudice for or against test automation to cloud one's ability to be impartial in the forecast.
Here's my personal thoughts/experiences:
- Testing job is underestimated.
- In General, Developers considered superior to testers.
- What makes Tester position difficult is 'repetitive tasks' . Yes you can automated tasks, but you still need to do some tasks that can't be automated. These are manual & repeated tasks, often boring.
- Some developers are so lazy. for ex: while testing we found 'python syntax issue!'
- Management thinks testers can be replaced once they automated everything. Obviously they push for this.
- I know for sure, there are projects with passionate developers but no-one can really take care of their testing side.
- Dev underestimate/avoid unit-tests & rely on testing team to find basic issues.
Automats have been around since 1897.
What I found was, most developers unwilling to do some basic testing for their modules. Even if you find bugs with logs & reproducible steps, they push their dev task to tester again like " can you perform 'git bisect' to find out the patch caused this bug"? or "Can you run 'tcpdump' from your reproducible steps?".etc.
worstcase is sometimes, software 'design flaws' (for example: arithmetic operation software missing division operations) are blamed on testing team for find it later.
Source: have worked in hospitals, raised by MDs, have asserted similar things in the past and have been corrected by trusted mentors.
I've seen testers, who find bugs and also fix them but they don't get credits they deserved.
That said, I do remember this one: "an ounce of prevention is worth a pound of cure." - Ben Franklin.
They were cheap and since they didn't know a thing about the software, they would mindlessly click through the test-plans and find many bugs.
Problem with them were, that they had a "half-life", when they learned too much, their bug-finds would decrease, because they would get more careful when using the software.
Often we would simply get 2-3 students of a non-technical field and replace them with new ones afer 1-2 semesters.
> for ex: while testing we found 'python syntax issue!'
Do you not at least have CI to run unit tests?!Exact conversion is like 'I tested code & after that made a small 1-line change & looks like missed something, sorry about that, pls don't tell it bosses :P'
When you get out a bit of the IT world, you'll find that people who demand software want to receive something: software. They bought an app, they want the app. Simple as that.
If you are good enough to have your code working without tests, good. If you don't need documentation, good. If you paint your walls with use cases, good. All that doesn't matter, if you deliver the app you were hired for.
And if your app doesn't work... well, everything you've done doesn't matter either. Because you were hired to deliver an app.
Of course tests are good, documentation is good, self-documenting code is good. But only for the IT. For non-IT people who's contracting you (you can be your company, too), he just want the app. The software. Working.
If you tell people "You told me you wanted software, not MAINTAINABLE software :-)" they'll say "Aren't you a professional? Shouldn't you just be doing this stuff? Isn't it just implied?"
So yes, they're paying for 'the software', but the tests are part of it, maintainability, as well as things like security and scalability are things that should be considered as well as just whether the software 'works'.
They'll go to the market and ask how much it costs for someone to include a new functionality. They'll get $100, $90... and your bid. You know you have all those tests and documentations properly done, then you can charge just $10 and win. And you can charge $70 and also win. It's up to you, because you know how hard/easy it'll be.
And if the software was made by someone else, who charged less in the beginning? You don't know how worky it'll be, so you have to charge a bit higher, like that $90.
So, from the client perspective, the difference from a well tested software and a barely-works one is just the initial price. After all contracts are over, they (remember, non-IT) can't know if the new functionality has a fair price or not. And who didn't developed it can't be sure about the maintainability, either.
Having tests (and everything else) was good for you, for your future. But, in general, the client didn't knew all this. He just paid for the software...
But for those of us who work on a team, it's far more complicated than that, and you have no idea who might be touching your code in the future.
I think addresses that point quite nicely?
But the collective mistakes of a team aren't a predictor of potential future members of the team. For example, we started hiring "junior" programmers, which increases the skill gap.
I very rarely have issues in problem space A, but I don't really test for A. Maybe a few here or there but it's a small minority of my tests and the coverage is probably just incidental. I test heavily in problem space B for whatever reason.
I get hit by a bus, and you're my replacement.
You have no experience with problem space A, so there are few if any tests around that space. All of a sudden the coverage is misleading because we might have 100% coverage or close to it, but the thoroughness isn't there and suddenly releases are less reliable and support ticket volumes increase.
That seems to cover the team element, too. Even if I'm not having a bad day, someone else messing around in my code might be.
Not only that, but you have to be able to test against yourself a month from now or a year from now, when you've completely forgotten some aspects of the design and might overlook something doing a modification. I write a lot of report generators at my job, so there's quite often a bug or a new feature in a report I haven't touched for months at a time.
Even if I'm alone in my team, When I refactor code after say 10 or 15 years I could just as well have been another person. I don't remember. I don't have the same skill set.
And the thing is: when you first write the code you usually don't know how many years you will maintain it. That might be a reason to postpone some test writing a while initially (if the code is scrapped in 2 months what whas the point of effort for maintainability?)
However there are upfront benefits with writing at least some tests for everything, in that it (usually) helps design a less coupled system.
There does seem to be a lot of resistance to the idea of handing problems to solo programmers right now though. Any hints as to how you've made this work?
That doesn't means that you should not have them. But you at least should be able to answer that question to be able to evaluate the value that they bring and how much tests do you need and where.
Right now the generated tests are pretty superficial and silly, but the key is that they are randomized. Because of this, we can run millions of variations, some of which will prove to be good tests. Right now I'm hacking an AI that will pick those good instances from many random test runs. If it works, we'll be able to simply take source code and produce (after burning enough CPU cycles) good unit tests for it. This will be a huge help in letting the human programmer only do "enough" test writing -- the AI will take care of the rest. Additionally, the solution can be unleashed on cruft code that no-one dares touch because of a lack of tests and understanding.
(Yes, there will be a business built around this, but that's for next year. :)
- 1. Test infrastructure is too complex. If I have to create a bunch of config files, obey a questionable directory structure, etc. before I can even write my test case, there is a problem. There should be very little magic between you and your test front-end.
- 2. Test infrastructure is too lacking. It is also a problem to have too little support. There should be at least enough consistency between tests that you can take a look at another test and emulate it. There should be clearly-identified tools for common operations such as pattern-excluding "diff", a "diff" that ignores small numerical differences, etc. depending on the purpose.
- 3. Existing tests should not be overly-brittle. Do NOT just "diff" a giant log file (or worse, several files), and call it a day; that means damn near any code change will cause the test to “fail” and create more work. Similarly, make absolutely certain when you develop the test that it can fail properly: temporarily force your failure condition so you know your error-detection logic is sound.
- 4. Tests should not be overly-large. Do not just take some entire product and throw it at your test, creating a half hour of run time and 40,000 possible failure points just because it happens to cover your function under test. It is vital to have a small, focused example.
If your test environment has problems like these, I fully understand the desire to balance time constraints against the hell of dealing with new or existing test cases, and wanting to avoid it completely.
And if you’re in charge of such an environment, you owe it to yourself to devote serious time to fixing the test infrastructure itself.
I think testing trivial code is a waste of time and does nothing but improve coverage numbers.
When you think about a non-trivial problem to write tests, you don't always know what the final code will look like. Maybe you forget an edge case or some small detail in the requirements that will cause you to restructure the code and approach the problem in a different way. In which case, you now need to re-write your tests. You might as well just write tests around the final version.
Maybe, but I'd wager on the whole that you probably don't. Your surface API should be the same, and you shouldn't be testing internal implementation anyway. If you write clear and consistent interfaces, then forgetting a requirement will mean adding a test – not rewriting them.
Unit tests are great, but they should still be testing at the boundary of code modules – if changing the internal implementation of that module breaks a test, it probably isn't testing at the correct level.
Why not? When my big black box stops giving me the expected answers I want to know exactly which cog inside that box that broke.
What kind of test could you possibly write that would break because of this? Something that measures time or memory usage, or uses reflection?
If you have clearly defined interfaces between significant internal components (which you should), why not test that they do what you ask? Public vs. internal is often more about what is exactly the client of the code, not the code's function or structure itself.
I agree that there's nothing worse than a test that says "your stuff broke" and provides no additional information why. But there's a balance between that, and testing the internals of an implementation.
If it's truly trivial code, you should be able to test it trivially, so I'm a little unsure of why this becomes a make-or-break issue for some people. Pretty sure more time and energy is wasted determining if code is trivial and needs to be tested versus just testing the damned thing ;-).
The DirectX12 api is non-trivial code; their testing has little in common with TDD.
Trivial tests may be trivial, but they are numerous: their need grows exponentially with code size. And they generate almost all the false positives you will get.
In some case unit testing is necessary, e.g. for ensuring that a hash function works exactly as defined. However, there are other cases where unit testing is absurd, and black-box API tests or automated tests could do a better job on error coverage. As an example, imagine the Linux kernel filled with unit testing everywhere: plenty unit testing religion fun, but no guarantee of getting anything better, but a risk of new bugs because of the changes and increased code complexity.
However, there are cases where unit testing is not suitable, or it is not a guarantee, or it is an additional risk, e.g. event-driven or low level stuff, multithreaded code, etc.
Ah, I get it. That explains the piece of s--- I'm looking at right now.
That said, the title might be sensationalist, but I agree with the holistic sentiment of the text.
I've seen dozens of GitHub repos with a "tests/" directory that only contains tests for the constructor and ignores all the parts that should be tested. You don't need to test a constructor, this is stupid. If your constructor is not working none of the other tests will -- BUT HEY, your constructor is working, it is not hard to see it.
When I am writing a module/function, I tend to continuously think of how this can be tested, which helps me design better abstractions.
For example if you're writing a class that uses a socket read/write, when testing you probably need to mock them. If you weren't planning on writing tests then probably you'd have ended by having the methods embedded in the class itself as read/write/close when those methods don't belong to the class and should be in another module called Socket that inherits a Socket interface. Now that you have a socket interface it becomes easier to test your class by passing a mock Socket interface.
> I get paid for code that works, not for tests, so my philosophy is to test as little as possible to reach a given level of confidence (I suspect this level of confidence is high compared to industry standards, but that could just be hubris). If I don’t typically make a kind of mistake (like setting the wrong variables in a constructor), I don’t test for it. I do tend to make sense of test errors, so I’m extra careful when I have logic with complicated conditionals. When coding on a team, I modify my strategy to carefully test code that we, collectively, tend to get wrong.
On the business side, you don't get paid for code at all. You get paid to make something people want. The fact that you're using programming to do that is inconsequential.
On the tech side, you're not delivering anything unless somebody, somewhere can test it, even if only one time.
So yes, you are getting paid for tests. In fact, that's the only thing you are getting paid for. The nub of the question is what the tests look like and how many you should have.
Depends. As a coder for hire sure. When I am a company employee though, I tend to think one level higher as a person who solves problems. There have been times that the business has thrown a problem on my desk thinking they need software, when all the business really needed was a process change.
Now, that doesn't mean you should not test. It means you should understand the limits of unit testing and test what is important as best as you can. Most every software engineering class at universities will cover this in-depth.
https://www.sqlite.org/testing.html
If you want rock solid software you need to spend the time to properly test it.
> Different people will have different testing strategies based on this philosophy, but that seems reasonable to me given the immature state of understanding of how tests can best fit into the inner loop of coding. Ten or twenty years from now we’ll likely have a more universal theory of which tests to write, which tests not to write, and how to tell the difference. In the meantime, experimentation seems in order.
Indeed, we still "don't know" how to test—more generally, and given the abundance of methodologies and their tendency to go through a hype and dump cycle, I would say we still "don't know" how to write code in the first place.
We'll get there eventually, but for now I would take whichever approach, methodology, tools, and language that I use as having a "best before" date, and invest in it accordingly.
I've long since learned the hard way that if you don't test the functions as you write them, the bugs get buried in the system, and become very hard to find. When you test your code as you write it and modify it (formally on larger projects, informally on smaller ones) this doesn't happen.
That's the advantage of tests, so I can get Beck's point: If the function is so painfully simple that you already know if it's right, (say, an accessor) just by looking, then it's not worth writing a test for it.
Put a sane, normal value test. This will pass unless shit's broken.
Then test edge cases. Test min, max.
Then test some impossible values. If they correctly fail, you pass.
I've always found that if you let the author of a piece of code decide on what value that code should be tested with, he'll test for the edge cases that he's thought of (and dealt with), not the ones that actually crash production.
Don't get me wrong, that's still valuable - if only for the non-regression aspect - but I feel property-based testing is a superior approach.
Write your code ("this is a function that sorts lists of ints"), write a property ("when given a list of ints, after sorting it, elements should be sorted and the list should have the same size"), let the framework generate test cases for you. Whenever something breaks ("lists of one element come back empty"), extract the test case and put it in a unit test to make sure the same issue doesn't crop up again in the future.
Usually, property-based tests are more complex, which means that the chance of an logic error in the test increases, and having something easy to verify hardcoded brings peace of mind.
>>I get paid for code that works, not for tests, so my philosophy is to test as little as possible to reach a given level of confidence (I suspect this level of confidence is high compared to industry standards, but that could just be hubris).
Is the humility in the parenthetical. There is a difference between arrogance and confidence.
In reality, there's a considerable amount of respect given to people who are great at what they do, but are also humble and without the boated ego.
But then again, it may be easier to just set a single round goal like 100% for test coverage. Writing that test for the single expression setter won't cost you a lot.
Now, the correct response may well be to decide that that particular bit of code shouldn't be tested after all. Sometimes this means adding a /coverage skip if: an outgrabed momerath is hard to reproduce on demand/ for a particularly tricky corner case, but more often it means that you haven't correctly separated your concerns and that particular piece of code actually belongs in a file that's not touched by the test suite.
Yeah, usually it won't cost anything to the guy who writes the test, but he is actually paid. What it costs to the business/customer/etc and whether the value added is worth of the cost is totally different matter.
I think there is clear incentive for writing unnecessary tests for certain developers - it is non-risky general work, where it is is difficult to fuck up anyway. If you are able to sell test-writing hours, what's the downside?
You can lose a business because of the heavy costs which add little value, however it is not easily arguable that a specific cost was the deal-breaker.
I doubt anyone does that. Having weird requirements like "all public methods should have at least one test" is insane. For many code bases that would be an order of magnitude more tests than a requirement of "100% test coverage"
I've had a small team nicely paid for months only to prove and document the that the product my company was to deliver won't fail in some specific scenarios, specified by the contract.
Those who don't produce the mission critical code (or believe what they produce is not on the critical path) unsurprisingly see the investment in the tests questionable. Of course, there is always a real danger of doing something "just because it is done" even if there's no real need.
Meaning if tests don't increase your confidence, but simply put in to tick a box for having a test, you aren't getting paid for that (or if you are getting paid for that, someones lost sight of what they are trying to achieve)
If Kent Beck is coding in a private bubble, he can do whatever he wants that makes code work.
But proper syntax is what gets you code that works. You aren't paid directly for it, though. Tests are not a direct path to code that works, but they can be a big help.
Code that "works" doesn't just mean it runs/compiles/passes CI/etc. It has to continuously add value. It can do this by running properly and efficiently across a wide variety of likely or infrequent conditions, as well as some exceptional scenarios. It can do this by being written clearly and not adding technical debt. It can do this by being as simple and/or as replaceable as possible. And ultimately, it can even add a final gasp of value by being easily deletable.
Which is unfortunately complete opposite of how TDD was interpreted, especially in its glory days (and in some corners, up until the present day).
Nobody needs division by zero tests if there are already guards in place so that can't happen, but it's quite helpful to have a "A goes in, B should come out" view from a client/user perspective. As long as behavior appears correct to the client and is not exploitable you're good to go.
Of course given no formal requirements, all that is possible are tests of the technical implementation. Regression errors will still be inevitable for customers and stakeholders, despite the "programmer" being able to claim his or her tests passed.
We need to find a way to stop kidding ourselves and find a way to test the right thing.
I'm not advocating testing getters/setters but not testing because "I don't write those kinds of bugs" can burn the next junior dev who might.
Testing is as much about finding bugs early as it is making your more agile in the future.
Its all a trade off that needs to be communicated to whoever is paying you.
The problem with comments like this is that they're too ambiguous. Someone who doesn't want to spend the time to write unit tests, will use this ambiguity as a mechanism for justifying their laziness.
That's one problem with the TDD mindset. If you start by looking for things to test, you might come up with unlikely scenarios or cases that don't matter much for your user.
I agree with Kent Beck mostly. I would add that tests can also be used to maintain invariants for future changes to increase maintainability. I just hope this quote isn't taken out of context.
It really depends on what your project is, what your goals for maintainability are and what programming language you use.
Two things about testing: - test to confirm your spec - if you have trouble writing tests, your design is probably flawed
a) Bad team mates. b) Future developers.
I've been burned one too many times with junior devs making cavalier changes in code they don't understand. Unit tests were THE solution for catching these changes.
Most comments seem to equate:
"regression tests"=="TDD"
... but it's really... "regression tests" is subset of "TDD"
I'm not a practitioner of TDD but I my understanding of its components are:1) the ergonomics & design of the API you're building by way of writing the tests first. In this sense, the buzzword acronym could have been EDD (Ergonomics Driven Development). Writing the usage of the API first to see how the interface feels to subsequent programmers. Arguably, a lot of incoherent/inconsistent APIs out there could have benefitted from a little TDD (e.g. func1(src, dst) doesn't match func2(dst, src))
2) a sort of specification of behavior by usage examples ... again by writing the tests first. Consider the case of programmers trying to figure out how an unfamiliar function actually works. Let's say a newbie Javascript programmer wants to know how to use .IndexOf()[1] What do many programmers do? They skip all the intro paragraphs and just hit PageDown repeatedly until they get to the section subtitled as "EXAMPLES". With TDD, instead of examples being relegated to code comments "sqrt(64) // should print 8" , it formally encodes the "should print 8" into real syntax that's understood by the automated test tools. (Test unit frameworks typically use the keyword "Expect()" as the syntax.)
3) an IDE that's "TDD aware" because it creates a quick visual feedback loop (the code that's "red" turns to "green") during initial editing. The TDD "artifacts" can also act as a "dashboard" for subsequent automated builds alerting you that something broke.
So TDD is a "workflow" and from that, you address 3 areas: (1) design (2) documentation (3) quality assurance via regression tests. With that background, the original Stackoverflow question makes more sense: how many "test cases" do I write because it looks like I can get bogged down in the test case phase?!?
[1]https://developer.mozilla.org/en-US/docs/Web/JavaScript/Refe...
That's like saying "I get paid for functioning software, not writing code."
Yes, you get paid for the output of the act, not the act itself.
Such as?
[Edit] I erased a cheap shot I took at Ken Beck. I still think the title of the article is stupid, but what the guy actually said is much more nuanced than that.
When coding on a team, I modify my strategy to carefully test code that we, collectively, tend to get wrong.
Allow me to explain:
Production use of code IS testing (manual, etc). Because it is an observation of the system state.
All of the world is testing. Every system is inherently a quantum mechanical one (ie: the observer is constantly testing the state of various systems to ascertain some level of confidence)
If you are going to test anything... then you should test the Use Cases (ie: Interactor objects). Don't have Use Case/Interactor objects that encapsulate intent? Well, you better understand it since the world, and therefore software, is all about intent.