Unit Test Autogeneration: A Tale of Two Models
hackernoon.com
hackernoon.com
GitHub's Copilot, which uses OpenAI’s Generative Pre-trained Transformer model, a derivative of GPT-3, does not explicitly generate unit tests, but it can suggest code snippets for testing.
So, while Copilot can be helpful in generating some initial test cases, it is not a replacement for a comprehensive testing strategy.
Microsoft Research, the University of Pennsylvania, and the University of California, San Diego have proposed TiCoder (Test-driven Interactive Coder), which leverages user feedback to generate code based on natural language inputs consistent with user intent. It uses natural language processing and machine learning algorithms to assist developers in generating unit tests.
When a developer writes code, TiCoder asks the coder a series of questions to refine its understanding of the coder’s intent. It then provides suggestions and autocomplete options based on the code's context, syntax, and language. It generates test cases based on the code being written, suggesting assertions and testing various scenarios.
Both Copilot and TiCoder, as well as other LLM-based tools, may speed up the writing of unit tests, but they are fundamentally AI assistants to human coders who check their work, rather than productive AI-based coders in their own right. So is there a better way?
Reinforcement learning systems can be far more accurate and cost-effective than large language models because they learn by doing.
Diffblue Cover, for example, writes executable unit tests without human intervention, making it possible to automate complex, error-prone tasks at scale.
The product uses reinforcement learning to search the space of all possible test methods, write the test code automatically for each method, and select the best test among those written. The reward function for reinforcement learning is based on various criteria, including coverage of the test and aesthetics, which include coding style that look as if a human has written them. The tool creates tests for each method in an average of one second, and delivers the best test for a unit of code within one or two minutes at most.
Diffblue Cover is more similar to AlphaGo, DeepMind’s automatic system for playing the game Go, than Copilot or TiCoder. AlphaGo identifies areas of a huge search space where there are potential moves to win the game, and then uses reinforcement learning on these areas to select which move to make next. Diffblue Cover does the same with unit test methods, coming up with potential tests, evaluating them to find the best test, repeating this operation until it has built a full test suite.
If the goal is to automate the writing of 10,000 unit tests for a program no single person understands, reinforcement learning is the only real solution. Large deep-learning models just can’t compete – not least because there’s no way for humans to effectively supervise them and correct their code at that scale, and making models larger and more complicated doesn’t fix that.
While large language models like ChatGPT have wowed the world with their fluency and depth of knowledge, for precise tasks like unit testing, reinforcement learning is a more accurate and cost-effective solution.
Mathew Lodge (mathew.lodge@diffblue.com) is CEO of Diffblue, an Oxford, UK-based AI startup.