This is such a salient question. Sometimes (definitely not always) the test suites produced by LLMs are so trivial it's scary. Coverage can be an illusion for sure.
This is such a salient question. Sometimes (definitely not always) the test suites produced by LLMs are so trivial it's scary. Coverage can be an illusion for sure.
I wrote a tool called tert - I guess it's called an agent harness now - to run various test runners and log test output and coverage output to disk. FWIU stripping spaces from JSON does save tokens. It seems like feeding coverage lines-missing maps into the prompt results in better output, better LLM-authored tests.
"Refactor these tests for maintainability and coverage. Use fixtures, mocks, and parametrization"
Substance coverage - testing the actual logic, edge cases, etc. Not mere lines.
If there is prompt insufficiency, there is probably acceptance test insufficiency.
A more assuming agent could automatically develop a plan that includes presumptive acceptance tests and request feedback before spending tokens
And, if/where we need tests, we write the source so they are few, high value, and complementary. Like actual unit tests, not complex with stuff like mocks just to generate trivial coverage.