> On a more serious note, the author is describing a scenario where mocks are generally not useful, ML or not: never mock the code that is under your control if you can help it.
When I'm deciding what to mock, I don't use "code under my control" as the determining factor. For me the question is, "is this the code I'm trying to test with this test?". If not, it's often a good idea to mock it even if it's your code, because you want your test failures to indicate where the error is.
Of course, you probably shouldn't be testing code that isn't yours, so that does mean you can mock code that isn't under your control.
For example, let's say you have two components, A and B, where tests of A end up triggering calls to B. In an ideal world, you have tests of B in isolation, so when writing tests of A, you don't need to test B again. If you don't mock B when you test A, then if something fails in B, you get failures in both A's tests and B's tests, and that makes it harder to figure out what's wrong. Instead, if you mock out B in A's tests, when something fails in B, you only get failures in B, and that pinpoints where the issue is.
There are caveats galore. The first caveat that comes to mind: If you're on a project with poor test coverage, B might not have any tests, and it may be worthwhile to write tests against A that don't mock B, so that you get some testing of B "for free". You can mock out B in A's tests once you have tests against B, but that may not be worth the effort. Yes, that's sort of blurring the lines between "unit test" and "integration test", but whether or not the code is tested at all is far more important than whether it's unit tested or integration tested.
[...]
LLMs are a prime candidate for mocking, because they're non-deterministic. Non-deterministic code it terrible for unit tests, because it will fail randomly, and over time that slowly normalizes ignoring test failures. You can make probablistic assertions on non-deterministic code, but there are two problems with this:
1. The more "wiggle room" you allow in your probablistic assertions, the less useful the assertion is--3rd-standard-deviation results are sometimes bugs, not just unusual results, and you don't want to ignore them, and...
2. Probablistic never equals deterministic; if you're running your tests often (and you should) then you'll get test failures on probablistic tests eventually. Broadening probablistic assertions only decreases the chance of failure, it doesn't remove it.
Probablistic tests create an impossible problem: if you tighten your assertions you struggle with meaningless failing tests and normalize ignoring test failures, while if you loosen your assertions your tests aren't really asserting anything any more. And there isn't a happy balance: if you go somewhere in the middle, you end up having both problems.
Mocking the LLM removes the non-determinism problem when the LLM is not the unit under test. That is to say, if you're testing code that calls the LLM, it may be a good idea to mock the LLM in those tests. Obviously, don't mock the LLM if your intent is to test the LLM.
Another possibility for removing non-determinism, which is probably more appropriate when you are testing the LLM, is to make it non-deterministic. You can do this by abstracting the RNG out of the LLM. Then when you want to test the LLM, you pass in an RNG that is seeded with a constant or even mock the RNG (gasp! You can do that!!?).
There are no silver bullets, and this approach has its downsides, because non-breaking changes to the LLM will likely result in changes to the now-deterministic outputs, meaning that a non-breaking change to the LLM code will still result in a bunch of breaking tests. This is why I would prefer mocking the LLM for tests that aren't intended for testing the LLM.