A significant part of writing characterisation tests can be simply staring at a code coverage report and asking “can I write a test (possibly by modifying an existing one) which hits this line/branch”. Sometimes that’s easy, sometimes that’s hard, sometimes that’s impossible (code bases, especially crapulent legacy ones, sometimes contain large sections of dead code which are impossible to reach given any input).
An LLM doesn’t have to always get it right to be useful-have it generate a whole bunch of tests, run them all, keep the ones which hit new lines/conditions, maybe even feed those results back in to see if it can iteratively improve, stop when it is no longer generating useful tests. Hopefully, that addresses most of the low-hanging fruit, and leaves the harder cases to a human.
There already exist automated test generation systems which can do some of this–for example, concolic testing-but an LLM can be viewed as just another tool in the toolbox, which may sometimes be able to generate tests which concolic testing can’t, or possibly produce the same tests quicker than concolic testing would. There is also the potential for them to interact synergistically - the LLM might produce a test which concolic testing couldn’t, but then concolic testing might then use that to discover further tests which the LLM couldn’t.