Why I don't mock
blog.metaobject.com
blog.metaobject.com
Isolated unit testing is supposed to improve your designs. But when you mock everything, you end up with tests that know everything about your implementation (white box testing). The result is that refactoring breaks your tests. Your tests no longer document how to use the system; they are merely a mirror of its implementation.
Kent Beck's central point in his book is that TDD mitigates fear, allows refactoring, and gives you immediate feedback. Extensive mocking makes developers fearful to refactor (thus hurting design) and reduces the quality of feedback. Not to mention it makes your tests hard to refactor too.
Martin Fowler explains the difference between "classical" and "mockist" testing:
http://martinfowler.com/articles/mocksArentStubs.html#Classi...
An other reason could be it's a main entry point for code execution, a class that only purpose is the glue others.
In that case, having a lot of mocks is a code smell pointing that you should probably test this on the integration level rather than on the unit one (and thus, discard mocks and use "real" everything).
(but beware not to be too much self indulgent on what is a gluing class, most of the time, they are not and it's indeed a god class problem).
Personally I like to avoid mocking for cases like this, since integration tests can usually suffice.
As you said, though, it could also be the framework. If you're in C#-land, Moq is pretty terse with the syntax so the actual mocking code isn't that big.
http://blog.8thlight.com/uncle-bob/2014/05/10/WhenToMock.htm...
http://blog.8thlight.com/uncle-bob/2014/05/14/TheLittleMocke...
What am I missing?
Only if you limit yourself to testing using unit tests. Two of the strongest testing tools care not one bit about internal state.
Combinatorial testing tests all possible combinations of input pairs, and looks for the system to fail.
Fuzz testing creates a fairly large volume of noise as input, and looks for the system to fail.
Integration testing should be in your toolbox right alongside unit testing, and will exercise objects which are not appropriate for unit tests.
> inheriting a lot of external dependencies that have little or nothing to do with your test.
I've found that this kind of thinking leads to development practices which rely on unit tests & production as their sole methods of testing. If a component is part of your production system, it should also be part of your test system.
The problem is not mocking. I found it to be absolutely helpful if used right (and right often means sparsely). The problem is this "all or nothing" mentality that many programmers have. Either they mock the hell out of the codebase or they don't use it at all.
I don't use mock objects, unless it's message based business logic, often located at the topmost layer of abstraction.
In these cases, whether you're PXE booting a server or talking to 500 different APIs that all range between the latest database or cloud API, flashing firmware on a RAID controller (shudder), whatever, there's a point of diminishing returns. There's also most of the bugs that you won't be able to simulate, because they come from external factors.
The magic lies in automated integration tests.
QA automation is AMAZING stuff. Set up real servers and real test environments when you can. Test real database upgrades.
Mock testing can have it's place, but often you find yourself making internal API contracts more rigid than you want, which can limit refactoring.
You can do things like set up scenarios you know won't work, such as connecting to machines that don't exist or sending the wrong parameters.
The main thing is if the external inputs can't be made to "violate a contract", the internals aren't the problem.
Ultimately, it depends on what you are building. Simple CRUD web applications are different from systems development type applications in many cases. Many of these classes can be more easily mocked out.
Often the breaking change you'll see is an upstream API or SaaS endpoint changing in ways you didn't expect (or the difference between a Linux utility on the latest Ubuntu, an older RHEL, and Solaris), so it's often better to concentrate on testing the real thing.
And these tests are expensive, so an automated test matrix that deploys a VM fleet with lots of permutations is usually necessary.
[0] Meditations by Sean Cassidy, http://blog.seancassidy.me/meditations.html
Checking these assumptions is what integration and acceptance tests are for.
Since reading Sandi Metz's very fantastic "Practical Object-oriented design in ruby" it clear to me that mocking is great for checking that objects other the one under test receive messages sent by the methods being tested (makes tests easier to setup and write usually). If I find that I'm doing too much mocking, then its usually because the method I'm using is doing way too much. FWIW.
1. Mocks are great, we'll use them everywhere.
2. Mocks are preventing our tests from being refactored.
3. Mocks are preventing our code form being refactored.
4. Mocks are preventing progress on the platform.
5. Code is being modified and tests and mocks are being removed to keep productivity up.
6. There are no mocks and few tests.
7. Tests increase again as defects are found. No mocks required.
This is over the space of 8 years and our problem was the amount of coupling they caused.
Mocks, almost by their very nature, encourage coupling. Ask yourself the question "What tends to get mocked? " Answer: Objects which are difficult and/or heavy to instantiate, but which are outside of your direct control and cannot be changed, e.g.: database connections, network sockets, etc.
So when you want to write tests around this you are almost forced to mock them out. So you instantiate those classes (as mocks) in your test class, so that the code under test. There are now two references to that class: one in the test, and one in the code being tested.
This is coupling. QED.
Answer: every collaborator of the object under test, so that only the tested object is actually having its implementation exercised. All other objects that it needs to talk to — database connections, network sockets, or whatever — are replaced with mocks, and the test subject’s interactions with those mocks are verified, so that a) we have confidence that the tested object is interacting with its collaborators in the expected way, and b) the test itself doesn’t depend at all on the actual behaviour of those collaborators’ implementations, which behaviour should be explicitly exercised elsewhere.
That’s the opposite of coupling. I don’t understand where your answer comes from or what it means.
The issue is that you should want to make it very easy to change the API signature of internal components.
Mock objects can tend to reinforce internal coding choices when things that are not a public interface, or a service boundary, are mocked out.
If you're mocking at this level, there is a potential for refactoring of intra-class API contracts to become much harder to change.
Thus, it would be better IMHO to mock out only meaningful service boundaries, and concentrate testing at the public contracts.
You can still get very strong coverage, but you're just thinking about a different level of inputs.
Mocking and stunning in the manner you describe makes tests extremely dependent upon the current implementation. In fact, it inverts the relationship your tests should have with your development process.
Tests should allow you to refactor without breaking tests, and tests should break when APIs you call change in incompatible ways. These are the two primary long-term benefits that tests offer. Stubbing and mocking all of their internals causes tests to break when you change internal implementation (library Foo was switched to library Bar, but Foo is mocked and has a different API than Bar). It also causes tests not to break when the Foo API is incompatibly changed by another developer.
I really hope that you're not being serious here.
By this reasoning, the test suite itself causes coupling to increase. For any component C, you have at least N=2: when it's used in production code, and when it's tested. If you don't test C, you reduce N to 1. This doesn't sound like a good argument against testing, though.
I think there might be subtle differences is how we are mocking objects and how we organize our code that might make mocking have different outcomes. That's why I would like to see actual examples that can be deeply analyzed. I suspect that it's in fact the way the code is written and what the underlying platform allows that makes one more suitable than the other rather than mocks being fundamentally a poor choice.
An example of complicated mocking can be seen here:
https://github.com/openstack/nova/blob/master/nova/tests/vir...
In particular, SpawnTestCase.
Disclaimer: While I didn't write this example, I am responsible for some of the other mess
Otherwise you get nasty effects, such as time moving on while your system is processing etc.
def has_expired(now=None):
#.. implementation
instead of: def has_expired():
#.. implementation
But at some point you need to call that method somewhere else... def get_request_handler(request):
#...code
now = timezone.now()
if has_expired(now=now):
raise Error403.
And `get_request_handler` still needs to be tested, which I 'stub' the API for getting the current time in the correct timezone within the test for the request. (Taking into account the other comment), using a library called 'mock'. (https://pypi.python.org/pypi/mock)a) changed your interface in a way that invites use cases you never intended (or wanted) to support.
b) added branching code anywhere that requires now()
The FP community doesn't seem to have a problem so much with testing simply because they mostly write pure functions that are data in -> data out and that is the easiest thing in the world to test.
It occurs to me that most developers aren't writing testable OO code, and thus we have a never-ending argument about the most clever way to test (or not test) code that is not easily tested in isolation.
If we never get around to writing testable code, we aren't going to have systems that are easy to test. If we don't have easy to test systems, testing is going to suck and people won't do it.
I don't know that FP is the answer, but just asking people to write better OO systems isn't working. The problem certainly is more fundamental than simply TDD or mocking or whatever language or tool we are using.
The problem is in how we look at solving problems in code and how that relates to system maintenance over time. It's a people problem, not a tooling one.
I just know that the OO design approach isn't working super well with the people who are using it, and we now have a couple decades where design hasn't progressed significantly on the average project.
So, how to we change the human behavior that is leading to less than ideal outcomes?
In code we have interfaces for axis configuration, axis movement and so on. All implemented for different controllers. On top of that there's a layer defining specific movements. Say you want to go from point a to b while avoiding c, there's a piece of code that calculates a route (so it needs axes configurations for max velocity/range/...) and another piece that applies this route to a set of axes (so it needs axes to drive).
Apart from dragging around a bunch of motors everywhere I code, how am I supposed to properly test those two pieces without any mocks? Telling me the design is all wrong to start with is fine, but then you also have to come up with a better one of course :). Telling me that route calculation shouldn't depend directly on axis configuration but should use it's own configuration type and I should then add a conversion between the two is also no solution as it a) introduces an extra layer and b) still requires a mock to test the conversion, d'uh :)
Some exceptions are almost impossible to generate consistently in tests, or can't be simulated easily next to test that are supposed to pass when that exception state doesn't occur - think API failures, databases being yanked mid-query, etc.
Is anyone aware of a solution like this? I certainly don't have time to create it but I would definitely find it useful.
So they seem to be saying "if you write code upfront before you need it, you might not need it, therefor mocking is bad"
IMHO, that is not a valid conclusion to draw.
I thought everyone agreed that writing code upfront before you need it is always a bad idea, especially specific implementations like a database.
let's take javascript.I cant generate mocks from my code,i have to write them.I have this data access object that takes a connnection.
var dao=new Dao(connection);
dao.getById(1,function(err,res){ console.log(res) });
Tell me how do I unit test getById in isolation without mocking connection,without doing whitebox testing,without knowing what actually happens in my getById.the obvious thing here is to spy on a hypothetic connection.query, fake a db call and return a predictable result.I cant do anything without a mock.
It means that “[mocks are] a way to allow you to test code that you can't improve right now” and “[mocks are] a stepping-stone for testing and improving a large existing codebase” is a mischaracterisation of the role and benefits of mocks.
If using them like that is useful for you then that’s great, but it’s misleading to say that that’s why “they’re there”.
Steve Freeman & Nat Pryce’s book [0], for example, is an excellent illustration of what mocks are for.
The difference between a stub and a mock is about the object under test, not the collaborators that are being replaced with doubles. We use a mock when we want to test the object’s interactions with its collaborators; we use a stub when we don’t care about the interactions, only about the values that result from them.
What does that decision have to do with whether the interface is small and well-defined?
1) Use the real thing. This can be slower and a lot less reliable.
2) Capture some typical output of the 3rd party api, and make a mock that emits it.
Which technique is useful? Both are. But the second one gives you fast-running fine-grained repeatable unit tests so therefore is your first line of testing.
So you discover in an integration test (#1) that the live api has a behaviour that the mock doesn't, maybe an occasional malformed response due to an error at the other end. Great, capture it and add it to the mock. Now you can reproduce it at will and write fast tests on how you handle it.
Note that a trivial gateway-implementation, like MemoryDataStoreGateway, isn't a mock! It's a fully-functional component that works perfectly well from its consumers' perspectives. But, unlike the version that consumes a third-party component, it doesn't have nearly any failure modes that would confuse your tests. It just does what it does, simply, and lets your tests test what you're trying to test.
And note that when you stop mocking, and limit yourself to switching out fully-fledged gateway implementations, "dependency injection" stops being this huge pain with Factories and heavily-parameterized initializers. Instead, you can just make your objects discover their collaborators through a service registry. Your unit tests just register the trivial implementations in the service registry on setup().
---
† Don't get a bad taste in your mouth thinking about Spring here. A simpler, healthier example is Erlang's "global" module (http://erlang.org/doc/man/global.html).
That's all fine until you need to test handling the third-party component's failure modes.
> Note that a trivial gateway-implementation, like MemoryDataStoreGateway, isn't a mock!
Probably not. Though it depends at which point you insert the "trivial" components. If you're using a trivial component to unit test a collaborating class that isn't at the edges of the system, then that's a mock and it's helpful. It's not the only helpful technique or even the usual one, but "mocks are never helpful, always avoid them" is silly thing to say.
A mock store for one unit test is far more lightweight than a MemoryDataStoreGateway, and allows an easy way to simulate server errors, so sometimes it is a better thing to code.
> This is what "dependency injection" is really about, actually: not inserting mock-objects ... but rather allowing you to plug in ... versions of your collaborators.
It can be about both. They are not exclusive.
> Instead, you can just make your objects discover their collaborators through a service registry
You may be happy with every object reaching out to a global singleton. But I'm going to stay well away. Thanks but no thanks.
The important thing about the gateway interface is that it's an error boundary: you don't need to think about "server errors" when dealing with a DataStoreGateway; DataStoreGateways don't have server errors. The ODBC object the ODBCDataStoreGateway holds might hand it a server error, but it won't propagate that back to you. It'll just put out a plain† DataStoreGateway::InsertError or somesuch.
† The exception should probably hold a copy of the implementation-exception that triggered it, but this is just for the sake of debugging the implementation. The client should never unwrap the internal exception.
I'm not sure why you even make that distinction. More than once I've dealt with servers that occasionally time out and fail; (say once a week or so). I wish to test what I show on my UI under those conditions, and what part of the text and status codes returned from the remote server that represent (or whatever) its failure I will choose to display.
Doing so without using a mock seems pointlessly obtuse and roundabout when I can make a mock to insert whatever error response I want at any interesting point in my code.
1. Your collaborator cannot be reached.
2. Your client has a network connectivity problem.
I've never seen a client who needs to know #1. All they care about is that there was a DataStoreGateway::TransientInsertError, and that TransientInsertErrors mean retrying and notifying the UI that their work hasn't been committed yet.
#2 isn't a problem the DataStoreGateway should be responsible for, because it will affect almost everything in the system in one way or another, and you don't want to have to write that code repeatedly. It's more likely a kind of Alarm--the kind of thing that gets pushed out from wherever it is to a global alarm handler, which then does something like putting the UI into a different root view-controller state, e.g. darkening it and covering it with a modal saying "lost connection."
Your design points may be valid, but they don't really talk about mocks; why bend over backwards to avoid testing using a mock?
The difference, then, from passing them their pre-initialized collaborators directly, is that if they need to create new collaborators on their own, they can ask the service registry what the current implementation they should be instantiating is. (It's basically a change in perspective from having objects that spawn specialized, short-lived objects, to having long-lived objects that can deal with changing conditions around them. Think about, say, Docker containers using links vs. etcd discovery. It's nice to not have to bring down your webserver just because your DB moved to a new IP address.)
Also, if you were thinking about my Erlang example specifically when you said "global singleton object" -- Erlang, despite the name, doesn't have one global registry, either. Instead, Erlang's process registry is per-node -- and nodes are the boundary where you deploy a release. In fact, ignoring the pragmatic OS-threading and networking implications, an Erlang node is almost exactly a "switchboard" on which services are plugged into one another, and an Erlang release is almost exactly a declaration of the current connections on such a switchboard. Need a place where your service is registered to X instead of Y? Bring up a node where that's true.)
For network requests to an API, your option #2 is exactly what I was suggesting.
There's a lot of value in coarse-grained mocks like this for testing the system; I also find value in finer grained mocks for testing smaller components in isolation. But neither of these techniques is the only valid one.
And neither of them forces you to write things upfront that you don't need later; which is what the original article seems to imply.
Anyway, sounds like we're more or less agreeing :)
Mocking helps when you are calling into code that you can't control. E.g, when calling into an API that is broken and sometimes returns incorrect results.
I don't write very many tests anymore. I haven't mocked an object in years. I also have a lot of bugs in my code.
By enabling isolation in the test, mocking reduces the pressure to design more isolated code.
Mocking helps when you are calling into code that you can't control.
Pervasive mocking helps isolate you from the pressure to remove (deep) dependencies on code that you can't control. Redesign your code so it only interacts with such code at the edges. At those places, create APIs that are independent of the code you can't control, and adapt that code to fit your APIs.
yeah, and what if you don't have time to do this? sometimes mocks are appropriate because of schedules and other priorities. Much better to mock than to not write tests because mocking is not pure.
If your schedules and other priorities are against fixing broken code, I'd be leery of your organization's long-term health. If you're in for the long haul--and at my current gig I certainly am--then eliminating bad code (and any code as tightly coupled as what you describe) is the only sane option. Because it'll kill you eventually.
Having just written some tests (with mocks) today, I have read your comment several times and failed to extract any meaning whatsoever from it. It seems like the kind of comment that's counter-intuitive because it's just nonsense. Mocking reduces isolation by enabling isolation? Um.
> Redesign your code so it only interacts with such code at the edges.
Yes, and those edges can then be mocked when you want to test the rest of the code. Apis, interfaces etc. are a good fit to mocking. Mocking helps isolation.
TFA's argument seems to be that they ended up not needing a database, and if they'd have "used mocking" they would have started by mocking the database, therefor mocking encourages bad design. Nope, starting with things that you don't need yet encourages bad design. It doesn't matter if you're mocking those things or not.
Yes... and that process you describe can be accomplished easily using mocks or stubs (depending on what you wish to test).