Monkey patching is fragile, but in tests, who cares? If it breaks, you see why and you fix it. Tests don't have to be held to the same level of reliability as production code.
Monkey patching is fragile, but in tests, who cares? If it breaks, you see why and you fix it. Tests don't have to be held to the same level of reliability as production code.
I personally take issue with that statement. If your tests are not reliable and robust, you will stop trusting them. At the point, why even have tests if you cannot trust them to verify correctness of your program (or at least increase your confidence that the program is not broken)?
In my experience, taking shortcuts in tests breeds more and more shortcuts (as less experienced developers look to existing tests as reference). As the shortcuts accumulate, the tests become more brittle, less effective and harder to maintain. This is the path to having a test suite that you no longer trust.
Red -> Green -> Refactor -> Refactor tests
Also, the tests can be less reliable by generating false positives. As long as all positives are checked, and the tests are maintained to pass (by fixing the tests and production code as needed), then false positives are not going to hurt your production code's reliability.
As long as your less reliable tests are not generating false negatives, you are fine.
This seems to be the paradox in much advocacy of strategies like TDD. We start with the premise that our code is likely to contain bugs. In order to detect those bugs, we write a lot of tests, perhaps doubling the size of our code base. Now, we can do whatever we like to the production half of our code base, as long as our tests in the other half all continue to pass when we run them, because magically that testing half of our code base is completely error-free.
For example, lets say you make a mistake in your test so you are checking (result = expected_result), which is always true, instead of (result == expected_result). Now when you write your code, you run the test and it passes.
In this case your code may or may not be correct, and the test, which contains a bug, does not catch it. But the bug is not a fundamental misunderstanding of the problem, rather a simple mistake in writing the test. Following strict TDD doesn't prevent this.
Actually, following strict TDD does prevent this.
You must see a test fail before you make it pass. It goes "Red -> Green -> Refactor" not just "Green -> Refactor". Even if a valid test case passes on the first time, one should modify the test (or the code) in a known way to cause an expected failure.
Your premise (tests are only as correct as the spec of initial assumptions) is correct, but your supporting example is not.
The risk is small with simple declarative tests ("the output of f(x) == y") but with complicated systems with large state spaces, you start writing combinatorial tests to try to cover more of the space and incur a bigger risk your test code is buggy.
Few things are as frustrating as fixing a bug that should have been caught by a test and realizing the test wasn't actually working correctly. This is even more acute when working on code where speed is important (where performance is correctness, in other words) and you realize a test has been measuring the wrong things, such as asynchronous runtime showing up in the wrong place.
TDD has always seemed more like a useful tactic to be applied to certain kinds of problems than as a silver bullet.
My worry about that argument is: how do you audit your code base to make sure you really aren't relying on monkey patching in production code?
This is a concern with any abstraction technique that provides neater code by hiding some part of the behaviour, starting with simply using a programming language at a higher level than assembly.
However, in some cases, particularly more static/compiled languages, there are often tools to enforce encapsulation of various kinds. On that basis, you can systematically prove that the code you're looking at really does honour certain guarantees.
In the dynamic/interpreted languages, this can be harder, because with so much being done at run-time and so much flexibility, almost everything comes down to trust. That makes it more difficult to move from informal/convenient to formal/rigorous code as a project evolves.
Often, that may be a price worth paying. In return we can get early prototypes up and running more quickly than we could with a heavier "engineering" language. But we shouldn't underestimate the long-term implications of that trade-off.
"Never" is a big word.
Years ago, I was tasked to change the text label of a file uploader. The site was built using a home-made framework on top of WebForms. It was even built to allow components to be replaced, but since everything touched everything else -- my simple task proved to be impossible. Even one of the original authors of the framework sat at my computer trying to show me, and after a couple hours even he gave up.
Compare that to a problem I was having with a Ruby on Rails site built on top of Spree from versions ago. I wanted to set up a staging server, and (for a reason that turned out to be my fault) I couldn't get the server to stop switching to SSL -- which the staging server didn't have. After spending a little bit of time to solve it, I just monkey-patched the SSL switch to false. Deployed, it worked, and I moved on.
We're really talking about edge cases here, but it's a symbol of a higher issue. Are we to be martyrs to past issues and bugs in code, devising more abstractions just to try to protect one code from another? Or is getting the problem solved in the most efficient way possible the better solution?
I'd use the phrase "it depends on the context" instead of "never."
These problems can be very difficult to find.
The coders who have to maintain the tests care. Tests are different to production code, but the trade-off where quick-but-fragile code imposes a tax in keeping it working is the same.