We created the first open source implementation of Meta's TestGen–LLM
codium.ai
codium.ai
I tried creating some on a personal project just using ChatGPT and it saved me a lot of toil on tests I probably wouldn’t have written. I did find I had low trust in refactoring my code, but higher than if I’d had no tests.
It seemed like a net positive for low risk cases.
Anything else with more complex logic I don't get much gain, having to read and understand the nuances of the generated logic is more taxing on my mind than just writing it myself, got tired of having to correct hallucinations most of the time.
Many people see unit tests as a chore, rather than the foundation of a codebase. If every project could get 60-80% coverage basically for free (except some super novel or intricate path, which should already have tons of unit tests), I see that as a net win for everyone.
For many of the CRUD platforms, you could even get close to 90% with a bit of hand holding.
So I see nothing wrong.
As they say: don't let perfect be the enemy of good.
Was great to have a pattern to tweak vs. blank code editor though, way faster.
So like 90% of human-written tests then?
Like most LLM generation though, it’s not a deterministic thing and like you mentioned originally it takes some verification of the output. I still think with the extra steps it saves time when applied to the right scenarios. The longer the input the higher the hallucinations count in my experience though, so I always keep the code provided in the smallest chunk possible which still has enough context.
Adding the simple unit tests ensures the current design of the internals is suddenly more "final" than before, and maybe more than intended.
I used just ChatGPT and Copilot.
For very simple test cases of simple methods/functions it's ok. Once code starts to complicate, you'll lose much more time trying to steer the LLM into writing what you need, than simply doing it yourself. In particular setting up and creating mocks seems where the most difficulty lies.
Once you create the test, you can use the LLM as a code completion tool, to test for more things once you get a complex result. Or you can ask it to extrapolate from test data to create more test cases.
LLM generated tests are the ultimate response to “Our CTO wants high test coverage but we all know it’s bullshit and provides little to no value”. They exist to tick a checkbox.
Maybe there’s a little value in using these as regression tests. Beyonce rule and all that. Kinda like double entry accounting. “Oh the code makes the same mistake in 2 places, that must mean it’s on purpose”
If you fixed it then you should have put a test on it?
Your Beyonce rule no longer means “I meant to do that” it now means “I wrote what I wrote”
But anything more complex and it is very hit or miss. I'm trying now to use GPT-4 Turbo to write some integration tests for some Go code that talks to the database and it is mostly a disaster.
It will constantly mock things that I want tested, and write useless tests that do basically nothing because either everything is mocked or the setup is not complete.
I'm settling in using it for tests for those small, pure functions, and more using it as a guide to find possible bugs / edge cases in more complex cases, then writing the tests myself and asking it in another prompt if they would cover those cases.
As most people that actually use AI heavily these days, I think the usefulness of AI for coding increases a lot if you already have a pretty good grasp of the subject and the problem space you are working on. If you already know roughly what you want and how to ask, they can be a huge time saver on the smaller and simpler things.
I like having it translate config formats too. A series of env vars to a yaml or toml or something
Specifically for generating unit regression tests the Cover-Agent tool already works quite well in the wild for some projects, especially isolated projects (as opposed to complex enterprise-level code). You can see in the few (somewhat cherry-picked) examples we posted [0] that it generates working tests that increase coverage (they were cherry-picked in the sense that these are examples we like to work with often internally at CodiumAI).
I believe that it’s possible to generate additional meaningful tests including end-to-end tests by creating a more sophisticated flow that uses prompting techniques like reflection on the code and existing tests, and generates the tests iteratively, feeding errors and failures back to the LLM to let it fix them. Just as an example. This is somewhat similar to the approach we used with AlphaCodium [1] which hit 54% on the CodeContests benchmark (DeepMind’s AlphaCode 2 hit 43% [2] with the equivalent amount of LLM calls).
If like me you think tests are important but hate writing them, please consider contributing to the open source to help make it work better for more use cases. https://github.com/Codium-ai/cover-agent
[0] https://www.youtube.com/@Codium-AI/videos [1] https://github.com/Codium-ai/AlphaCodium [2] https://storage.googleapis.com/deepmind-media/AlphaCode2/Alp...
It's hard to see value in spending resources this way right now - most notably, engineer time to review the generated tests. Improve the hit rate by an order of magnitude, and I suspect I'd feel differently.
Based on my understanding, only 1:20 passed the automated acceptance criteria (build, run, pass, increase coverage). Of those that made it through to the human review, “over 50% of the diffs submitted were accepted by developers” according to the paper
> In highly controlled cases, the ratio of generated tests to those that pass all of the steps is 1:4, and in real-world scenarios, Meta’s authors report a 1:20 ratio.
> Following the automated process, Meta had a human reviewer accept or reject tests. The authors reported an average acceptance ratio of 1:2, with a 73% acceptance rate in their best reported cases.
1:20 of real-world generated tests reach human review, or 5%. Of those, on average 1:2 are approved, or about 50%. 50% of 5% is 2.5%, or 1 in 40. Where do you see the error?
edit: Okay, I think I see it, specifically in considering engineer time vs the ~2.5% overall hit rate vs. the ~50% rate for tests reaching human review and thus requiring effort. Fair callout, thanks!
In the end I think the important part is not giving tests more trust than they deserve. If you fix the incorrect code and break a test you should give your change a closer look, but also give the test a closer look before you decide which needs to be fixed.
Is your name a reference to Gronky Scripples? https://www.youtube.com/watch?v=4KG3v365mq4