Meta's new LLM-based test generator
read.engineerscodex.com
read.engineerscodex.com
I could see it as very helpful though for an LLM to point out underspecified areas. Maybe having it propose unit tests for underspecified areas is a way to do look at that and what's happening here?
Edit: Even before LLMs were a thing, I sometimes wondered if monkeys on type writers could write my application once I've written all the tests.
People who work on legacy code bases often build what are called “characterisation tests” - tests which define how the current code base actually behaves, as opposed to how some human believes it ought to behave. They enable you to rewrite/refactor/rearchitect code while minimising the risk of introducing regressions. The problem with many legacy code bases is nobody understands how they are supposed to work, sometimes even the users believe it is supposed to work a certain way which is different from how it actually does - but the most important thing is to avoid changing behaviour except when changes are explicitly desired.
An issue with going with llms will be to validate if the behavior described are merely tolerated or if they're correct. Another will be wether something is actually tested (e.g. a code change still wouldn't break the test). Too granular output check would be an issue as well.
All in all this feels like a bad idea, but I hope to be wrong.
Characterization tests ideally need to not be written in code but defined in something resembling a configuration language - something without loops, conditionals, methods, etc. There then needs to be a strict separation of concerns kept between these definitions and code that executes them.
I wrote a testing framework (with the same name as my username) centered around this idea. Because it is YAML based, expected textual outputs can be automatically written into the test from actual inputs which saves tons of time and you can autogenerate readable stakeholder documentation that validates behavior.
It might seem unbelievable but with decent abstractions, writing tests and TDD stops being a chore and actually starts being fun - something you won't want to delegate to an LLM.
The technical aspect of the code is not actually the difficult part WRT maintenance. Someone that knows COBOL as a language can figure out what some code does and how it does it. It takes time but that is information that can be derived if you just have the code.
The main problem with COBOL is the code is often an implementation of some business or regulatory process. The COBOL maintainer retiring is taking knowledge of the code but more importantly the knowledge of the literal business logic.
The business logic and accounting/legal restraints aren't something that can necessarily be derived from the code. You can know some bit of code multiplies a value to 100 but you can't necessarily know if it supposed to do that. If the source code and documentation don't capture the why of the code the how doesn't help the future maintainer.
Often with COBOL the people that originally defined the why of the code are not just retired but dead. The first generation of maintainers may only have ever received partial knowledge of the why so even the best documenters have holes in their knowledge. They may have had exposure to the original why defines but neglected or didn't have an opportunity to document some aspects of why. The subsequent generations of maintainers are constrained by how much of the original why was documented.
Edit: pre-coffee typo
An LLM doesn’t have to always get it right to be useful-have it generate a whole bunch of tests, run them all, keep the ones which hit new lines/conditions, maybe even feed those results back in to see if it can iteratively improve, stop when it is no longer generating useful tests. Hopefully, that addresses most of the low-hanging fruit, and leaves the harder cases to a human.
There already exist automated test generation systems which can do some of this–for example, concolic testing-but an LLM can be viewed as just another tool in the toolbox, which may sometimes be able to generate tests which concolic testing can’t, or possibly produce the same tests quicker than concolic testing would. There is also the potential for them to interact synergistically - the LLM might produce a test which concolic testing couldn’t, but then concolic testing might then use that to discover further tests which the LLM couldn’t.
Characterisation tests are not supposed to be tightly coupled – they are supposed to be integration/end-to-end tests not unit tests – the point is to ensure that some business process continues to produce the same outputs given the same inputs, not that the internals of how it produces that output are unchanged. Code coverage is used as an (imperfect) measure of how complete your set of test inputs is, and as a tool to help discover new test inputs, and minimise test inputs (if two test inputs all hit the same lines/branches, maybe it is wasteful to keep both of them–although it isn't just about code coverage, e.g. extreme values such as maximums and minimums can be valuable in the test suite even if they don't actually increase coverage.)
They can take the form of unit tests if you are focusing on refactoring a specific component, and want to ensure its interactions with the rest of the application are not changed. But at some point, a larger redesign may get rid of that component entirely, at which point you can throw those unit tests away, but you'll likely keep the system-level tests
Just in case any reader hasn’t tried this, the basic idea is to make statements about code’s behavior that are weaker than a totally closed-form proof system (which also have their place) stated as “properties” than are checked up to some inherently probabilistic bound, which can be quite useful statements.
The “canonical” example is reversing a string: two applications of string reverse is generally intended to produce the input. But with 1 line of code, you can check as many weird Unicode edge cases or whatever as you have time and electricity.
I know this example seems trite, but I met this because some hard CUDA hackers doing the autodiff and kernels and shit that became PyTorch used it to tremendous effect and probably got 5x the confidence in the code for half the effort/price.
It doesn’t always work out, but when it does it’s great, and LLMs seems to be able to get a Hypothesis case sort of, closer than starting from scratch.
In my experience, at least in languages like C++ or Java, unit tests are made of tedium, so I'm absolutely not surprised that the first instinct is to use LLMs to write that for you.
Sadly, this is a rare approach, so if you cooperate with others it's hard to use it.
This system would be a godsend in the minds of engineers who think / operate that way.
I've also had managers who told me I wasn't allowed to write tests firsts as it was slower. Luckily I was able to override / ignore them as I was on loan "take it up with my boss". They're probbably thinking the same as the above engineers.
Another way to think of this is most devs hate documentation... if they had an AI that would write great docs from the code they'd love it. And these to these devs docs they don't have to write are great docs :)
Sounds like a great place to work.
You're suppose to write them together, at the same time. What's slower is spending all of your time mousing around the GUI like a caveman and then having missing tests and then trying to debug every future recurring regression forever the rest of the project's lifespan.
The only thing that's slower is perhaps trying to do TDD for the first time. After that, you're embarrassed at how much time you wasted wandering in a browser before your API tests were finished.
I feel the same way about how test code is viewed even outside of AI. A lot of the time the test code is treated as a lower priority code given to more junior engineers, which seems like the opposite of what you would want.
Getting LLMs to write tests is like getting LLMs to write my spec.
For the remaining 95% of uninteresting surfaces I'd be perfectly happy to let an AI interpolate between my key cases and write the tests that I was mostly not going to bother writing anyway.
It's nowhere as fancy as FB tool but I know it is blessed by company.
I’m guessing they want to automate tests because most engineers skimp on them. Compensating for lack of discipline.
- Software engineering, which is akin to real engineering, as it involves desigining a complex mechanism that fits a lot of real world constraints. Usually you need to develop a sophisticated mental model and exploit it for the desired results. Involved implementations and algorithms usually fall into this category. - Talking to computers, which is about describing what you need to the computer. Usually focuses on the 'what' as the 'how' is trivial. Examples include HTML/CSS, Terraform and very simple programs (like porting a business process flow from a flowchart to code). And, indeed test code.
LLMs are terrible at the former, but great at the latter.
AI-generated tests can work like compiler/sanitizer warnings. If they fail, you can audit them and decide if it was a true or false positive.
- TDD, which as you say describes the system's behavior. But it often deals with the nominal cases. It is hard to predict all that can go wrong in the initial development phase.
- tests designed to reproduce a bug. The goal of these to try very hard to make the system fail, taking inspiration with the bug's context
Maybe this LLM test generator could allow to be more proactive in the second kind?
I dread the morning after a night of getting something to work…somehow.
When you try to get the LLM to write the code, you find that it’s easier to get it to write the tests. So you do that and publish about that first.
And making those assertions shouldn't be mindless work. Documenting the failure cases should be the most interesting work of all the code you are writing. If you are finding that it isn't, then that tells you that you should be using a DSL that has already properly abstracted the failure cases for your problem space away.
* While "errors are values" makes it possible to enumerate the potential failure cases, most of the time errors are just strings. The thing you are forced to document by the lack of exception semantics is the sites at which you might be dealing with an error vs. a success value. In an IO heavy application this rounds up to "everywhere."
* Go is not really powered to do DSLs in a type-safe way, except through code generation. I would view the LLM as a type of code generator here.
* If your errors are strings, you are almost certainly doing something horribly wrong. Further, your tests would notice it is horribly wrong as making assertions on those strings would look pretty silly, so if your errors are strings that tells that your testing is horribly, horribly wrong.
* Go is not a DSL, no. It is unabashedly a systems language. Failure is the most interesting problem in systems. Again, if your failures aren’t interesting, you’re not building a system. You should not be using a systems language, you should be using a DSL.
When you have no idea what you are doing, choosing the wrong tool at every turn, an LLM might be able to help, sure.
I agree that the widespread use of code generation (LLM or otherwise) is revealing of this tension.
Do you mean Goa? Indeed, it is not a language fit for what we're talking about. It only allows description of services, not the implementation. For that you have to fall back to Go, which completely misses the whole reason for using a DSL.
> "one of your upstream calls might fail and then you return an error to your caller" is not that interesting.
Sure, and which is why you wouldn't use Go here. There is absolutely nothing in Go that is geared towards abstracting those kinds of things away. And you can't tell me that Goa tempts you. It is not a good DSL.
Most systems are pretty predictable. it("displays the user's name") isn't very novel, and is probably pretty easy for a LLM to generate.
Arguably the implementation that passes such a test is even simpler, making it a bit questionable why we have the human write that part.
From the blog:
> Most of the test cases created by Meta’s TestGen-LLM only covered an extra 2.5 lines. However, one test case covered 1326 lines! The value of that one test case is exponentially more valuable than most of the previous test cases and exponentially improves the value of TestGen-LLM. LLMs can vigorously “think outside the box” and the value of catching unexpected edge cases is very high here.
Of course "exponentially more valuable" should set off your BS detector. But to verify, from the paper:
> However, this result arose due to a single test case, which achieved 1,326 lines covered. This test case managed to ‘hit the jackpot’ in terms of unit test coverage for a single test case. Because TestGen-LLM is typically adding to the existing coverage, and seeking to cover corner cases, the typical expected number of lines of code covered per test case is much lower....The median number of lines of code added by a TestGen-LLM test in the test-a-thon was 2.5. This is a more realistic assessment of the expected additional line coverage from a single test generated by TestGen-LLM.
Nowhere do the authors mention "unexpected edge cases" or "thinking outside the box." They clearly present this 1,326 lines of coverage test as a fluke, e.g. maybe the test case checked one branch of a horrible switch statement, or perhaps it was even a fluke in how code coverage is counted. It is noteworthy that the authors do not seem to have looked into it any further, even in the "qualitative results" section.
Inaccurate editorializing really doesn't help anyone. The internet is too damn full of people pretending to understand things they pretended to read.
This article is less of a summary of a paper and rather commentary on what the results of the paper entails. After all, Hacker News is meant for discussion :)
I will say though that I do believe that I still stand by the "exponentially more valuable" portion. I think the fact that LLMs can fluke their way into "hitting a jackpot" in terms of test coverage is exactly why they're so valuable. When you have something constantly trying out different combinations, if it hits even one jackpot, like in the paper, it's extremely valuable to the team. It's a case that could have been either non-obvious or simply too tedious to write a test for manually. I think there's tremendous value in that, especially speaking as someone who has spend way too much time simply figuring out how to test something within a Big Tech codebase (F/G) when I already knew what to test.
You CAN'T have exponential growth that is not a function of some value or variable or input.
I suppose in this case you could argue you have exponential growth as a function of the discrete using-an-LLM or not-using-an-LLM, but I've never heard of exponential growth as a function of a discrete.
Often people using the term "exponential growth" in common English don't understand what it means. Sorry.
FWIW, exponential growth as a function of a discrete variable is very common (e.g. all of algorithmic complexity), but it has to be (at least modeled as) an unbounded numeric variable.
You can't have exponential growth as a function of a binary variable.
My problem is we have no clue what those lines actually were. If it was effectively dead code, then it's not surprising that it was untested, and the LLM-generated test wouldn't be valuable to the team. We have no clue what the value of the test actually was, and using a single stat like "lines of code covered" doesn't actually tell us anything. Saying the test was "exponentially more valuable" is pure speculation, and IMO not an especially well-founded one. (Sort of like saying people who write more lines of code are more productive.)
This speculation seems downright irresponsible when the paper specifically emphasizes that this result was a fluke. When the authors said "hit the jackpot" they did not mean "hit the jackpot with a valuable test", they meant "hit the jackpot with an outlier that somewhat artificially juked the stats." I truly believe if the LLM managed to write a unusually valuable test with such broad coverage they would have mentioned it in the qualitative discussion. Instead they went out of their way to dismiss the importance of the 1,326 figure.
Some of my comments within the article are more aspirational than realistic in this case, and I've made edits to reflect that.
I want to clarify that I view this LLM as a junior dev that submits PRs that pass presubmits and other verifiable, programmatic checks. A human dev then reviews the PR manually. In this case, the LLM + its processing is used to make sure that no BS is sent out of review - only potential improvements.
In no scenario should it's auto-generated code be auto-submitted into the codebase. That becomes a nightmare really fast.
One additional question - do you forsee any issues with this application where LLMs enter a non-value add "doom loop"? I can imagine a scenario where a test generation LLM gets hooked on the lower value simplistic tests, and yet management sees such a huge increase on the test metric ("100x increase in unit tests in an afternoon? Let's do it again!") that they continue to bloat the test suite to near-infinity. Now we're in a situation where all future training data is now training on complete cesspool of meaningless tests that technically add coverage, but mostly just to cover an edge case that only an LLM would create.
Not sure if that makes sense, but tl;dr - having LLMs in the loop for both code creation and code testing seems like it's a feedback loop waiting to happen, with what seems like solely negative repercussions for future LLM training data.
Rather, I prefer to view LLMs as a junior dev that submits PRs that pass presubmits and other verifiable, programmatic checks. A human dev then reviews the PR manually. In this case, the LLM + its processing is used to make sure that no BS is sent out of review - only potential improvements.
Why not treat every LLM as a dev contributing to git such that Humans, or other LLms need to gatekeep in case something like that happens? (start by treating them as Interns, rather than Professors with office hours)
Most tests are use-case based or written for checking error-handling.
Use-case tests are easy to write: you don't even need to be a programmer (in fact, it's good if use-cases are defined by a Product Owner or Tester), though of course some cases are only known to the programmer as they're the only ones who dive into the details. The programmer should come up with all use-cases the PO missed, of course, and judge whether or not they need to test those too... sometimes it's ok to not test as the cost-benefit is low. Anyway, once you have this use-case based test mentality, it's very easy to write the tests (using a proper language to do it is important! Don't use just JUnit if you're doing Java as it will be really tedious to write and you will stop midway - I know, I've been there... I highly recommend using Spock, though other frameworks to make writing test pleasurable exist).
This applies mostly for "integration tests". For unit tests, hopefully you don't find them difficult to write?! I find them quite easy to write since I know how to write testable code, which takes a while to learn but once you do, it's really easy.
If you have examples of difficult to write tests, I would be curious to see it! Perhaps we can discuss how to make them easy.
For example GUI testing is IMO still an unsolved problem. Maybe AI will help there but existing solutions are generally not worth the pain.
I work in silicon verification and the testing we do is way way way more thorough than software testing, for obvious reasons. Do you formally verify your software? Unlikely.
I can only assume you work in an easy-to-test domain on a project that doesn't have changing requirements, like... I dunno a C compiler or something.
You're dismissing my claim by basically saying I am naive. Which is not an honest argument as you know absolutely nothing about me (and I don't want to tell you more than I did here).
About changing requirements: what does that have to do with testing at all? If requirements change, you basically discard the tests for the old behaviour and start over...
I would say that the testing we do is very close to formal verification because it's close to being comprehensive - though no, we do not use methods normally classified as such. I tried to but the benefit we would get over our current approach would be negligible.
By the way, I do a lot of UI testing, and dare I say it: yes, it's easy too.
We use this sort of thing if you're curious: https://gebish.org/manual/current/#pages
Again, if you find a real example of something you find hard to test, let me know so I can evaluate it against my own situation.
To be fair most projects don't care beyond "rollback the database and maybe display a 5xx error or something". But some do. Anyway, use cases are fine. It's the failures / edge cases that cause pain.
interesting perspective - why do you think this is a bad thing?
to me, it's an opportunity to verify that the change is intended. without it, how do you know that the program does what it is supposed to do?
Thinking of it as a leaky abstraction helps me.
I try hard to separate domain logic tests from implementation specific tests.
Your code could be loosely coupled with high cohesion, but with lots of random tests like you get when code coverage is a performance metric, you have to add a lot of complexity that only relates to an implementation.
People then get used to just blindly updating the golden images. It becomes basically "the output changed, do you want to continue anyway" which is not the most useful thing. You really want it to say "the output is wrong".
I don’t want to write tests for everything. I just want to write the ones that matter.
TDD is _about_ writing tests that matter, but most people think it is about writing all unit tests first.
If you are following TDD anywhere close to the way it is described, you will only be writing tests that relate to domain functionality first.
Note how it is described here, although it is turse.
https://martinfowler.com/bliki/TestDrivenDevelopment.html
The coverage metric as a goal writing style doesn't work for TDD, sorry you were exposed to that.
You are correct that model doesn't work.
Coverage is not a goal of TDD, but in practice you will have 100% coverage by following TDD as you would never have reason to write code that isn't covered by test.
Ultimately, the purpose of coverage tools is to let you know what you might have forgotten to clean up during a refactor, to help you remove what you missed.
More importantly, why are you writing any code for things that don't matter?
Another way, if you know what the code is supposed to do, why write it down in two places?
This would be like criticizing double-entry accounting by asking "if you know what the amount is, why write it down in two places?"
We write the code down in two places because that gives us advantages that far outweigh the added effort:
- Once written, your test will catch regressions forever
- A test is often excellent documentation on what the code does
- It's now much easier to refactor the code, making it more likely that it will be refactored when needed.
We are all terrible at writing tests. We just find our own ways to do it.
But yea 9/10 times that a snapshot test fails, it's noise rather than signal.
I'm not saying that writing tests before the production code is something that should always be done.
But tests are just as much a part of the codebase as anything else, and absolutely must be written alongside the code being tested.
The most important part of the test is that it showcases intent of the developer. A test suite demonstrates the following:
* How the code should be used
* What the code does
* What the code doesn't do
* What it was written for
Then when that code is used or modified by another developer, they don't have to hunt for clues in the codebase like they're Sherlock Holmes.
If the tests aren't telling a story, you're writing tests wrong.
And until the computers gain the ability to read your mind and do a better job at understanding what you want to do, AI/LLM-based generators can't do this job for you.
Of course, if the only goal of your test suite is getting a green checkmark on a pre-commit check (and being able to show great coverage numbers), then yeah, you can double your productivity with AI.
Automatic code generators will surely help you write more bad code at lightning speed.
And if others complain that tons of boilerplate make the code bloated and hard to understand — just tell them to use AI to deal with it. Worked for you!
That really does seem to be the future of development. But not the future I'm looking forward to.
There are different types of testing, what you're describing sounds to me like testing the "core" of your code, part documentation, part validation, part stability, etc.
Other types of testing like fuzzing provide an entirely kind of value. I believe this AI- driven testing can inherit a space to target tests at the tail end of the distribution, many tests with little value. Providing extra coverage where human energy and time is lacking.
That is how I see the current state of AI tooling regardless, as a cognitive assistant.
I'd be surprised if this line of research doesn't end up being very fruitful in the coming years.
Your comment presents a way more grounded perspective on the future of LLMs in programming than the article does.
https://news.ycombinator.com/item?id=39406726
Their abstract doesn't match their actual paper contents. That's unfortunate. Their summary indicates rates in terms of test cases:
> 75% of test cases built correctly, 57% passed reliably [implying test cases by context], and 25% increased coverage [same implication] The actual report talks about test classes, where each class has one or more test cases.
> (1) 75% of test classes had at least one new test case that builds correctly.
> (2) 57% of test classes had at least one test case that builds cor- rectly and passes reliably.
> (3) 25% of test classes had at least one test case that builds cor- rectly, passes and increases line coverage compared to all other test classes that share the same build target.
Those are two very different statements. They even have a footnote acknowledging this:
> For a given attempt to extend a test class, there can be many attempts to generate a test case, so the success rate per test case is typically considerably lower than that per test class.
But then in their conclusion they misrepresent their findings again, like the abstract:
> When we use TestGen-LLM in its experimental mode (free from the confounding factors inherent in deployment), we found that the success rate per test case was 25% (See Section 3.3). However, line coverage is a stringent requirement for success. Were we to relax the requirement to require only that test cases build and pass, then the success rate rises to 57%.
Maybe the right side should be something about fuzzing and using static types. Use systemic and automated checks. I’m a little ashamed that I thought in memes so readily.
LLMs will never get any better than they are right now and haven't improved at all in 2 years. Just fancy Markov chains.
The only way they can be used to write code is by people who don't know how to code blindly commiting code to prod without any review whatsoever.
People who do know how to code couldn't possibly have a use case and it won't make them any more productive.
I'm just going to ignore all this LLM nonsense that isn't changing the world at all and you definitely should too.
Yes authoring some tests might be sped up but not necessarily maintaining them - or maintaining the code under test because you are not necessarily generating good ones. Not to mention sweating over tests usually help developers with checking the design of the code early on too; if not very testable, usually not a good design either, e.g not sufficiently abstracted component contracts which suck in a context where you need to coauthor code with others.
What some people miss is that tests are supposed to be sacrifical code, that most of which will not catch anything during their lifetime - and that is OK because it gives an automated peace of mind and saves from potential false clues when things fail. But that also means max investment into a probabilistic safeguard is not gonna pan out at all times; you will always have diminishing marginal utility as the coverage tops. Unless you're writing some high traffic part of the execution path - e.g. a standard library - touting high coverage is not gonna pay off.
Not to mention almost always an ecology of tests need be there - not just unittests but integration, system etc - to make the thing keep chugging at the end of the day. Will llm's sit at the design meetings and understand the architecture to write tests for them too? Or what they can do will be oversold at the expense of what should be done. A sense of "what is relevant" is needed while investing effort in tests - not just at write-time but also at design-time and maintain-time - which is what humans are pretty OK at, and AI tools are not.
What llms can save time with is keystrokes of an experienced developer who already has a sense of what is a good thing to test and what is not. It can also be - and has been - a hinderance with making the developers smuggle not-so-relevant things into the code.
I don't want an economy of producing keystrokes, I want an appropriately thought set of highly relevant out keystrokes, and I want the latter well separated from the former so that their objective utility - or lack thereof - can be demonstrated in time.
I showed it a TypeScript module, asked it to generate a unit test and it made a working test not only covering the happy paths but a few edge cases as well.
I’m not resonating with the downvotes here on similar comments.
ChatGPT goes above and beyond for me in many ways.
Tests seem… easy in terms of gpt capabilities.
Last week I had it write python that traversed an AST and construct a react flow graph as well as the component. I made no edits, went through a few iterations of prompt feedback, and it worked great. Many similar interesting abilities I’ve observed from gpt.
I think this is an interesting experiment but somewhat dubious. The way I see AI would best help software development is that I the programmer have a question about my or somebody else's code, which the AI then answers, sometimes with a code-proposal but not always. It should be able to answer questions like "Is there a way to simplify this code? What are some inputs that would cause an error?" etc.
AI should help us understand the code, and understand how to improve it. Not write all of it on its own because if we don't tell it what to do, it cannot know what we want it to do. Tests is a good example. What do we want it to test?
Forest for the trees if you ask me, but, to each their own.
The bad news is that often the business folks don't know what they want, either, which is how "agile" became a thing. I recognize the ship has sailed on that, but it's "cake and eat it too" to think one can have good tests and ship "PoCs that do something valuable" in 2 week increments
Everyone seems to feel one way or the other about AI code, it's a very political topic. But I would just wanna try it and see.
This is pretty interesting, because a lot of these technologies are staggeringly expensive to develop. The AI tooling I've used so far has been somewhat useful, but if it doesn't get much better it won't have been worth the cost that was paid to create it.
I'm pretty optimistic about what will be achieved but even with my optimism it's far from clear that it's actually gonna pay for itself.
Very important. Very valuable.
Imagine if `if err != nil { return err }` just stopped working tomorrow. Your tests would detect it! Outage prevented!
Only past few days though, I did find two bugs that would've been prevented if the original code was covered by a decent unit test.
So you've never made a change caused a unit test to fail? If not, how large is your codebase, and is ownership shared across multiple teams?
I caught dozens of latent or unreported bugs by writing unit tests for a 6kloc JS app which had 0% coverage before.
In my experience LLM are smart but sometimes inconsistent and over a long chat it might say things that are logically self contradictions… when you tell it that it confirms it.
It just seems like it lacks a consistent world view.
I don’t trust them yet. Maybe with even more scale they become better.
They act a little bit like young children, with a lot of domain knowledge.
Ideally, "one change" (Whatever that might be) in production code should cause exactly 1 test to fail.
How does TestGen-LLM address this problem?
what does this mean ? i would have thought they would simply use codellama. is there any research around privately finetuned code llms ? why would they be better ?
More germane to this discussion, I would guess a locally tuned model will also recognize the kinds of things they care about testing, up to and including spotting any bug fix tests that were hard won and can carry forward in any such generated tests for future code
GitHub copilot instead, as an intelligent Intellisense, is absolutely great; a real blessing and gift to coders.
Funny you mention that, since lawyers have already gotten in trouble for citing fictional cases when submitting work performed by chatgpt.
It's useful for rote things, but for anything that you depend on, you still need to give it just as much attention as if you'd done it yourself.
For something like legal work, you don’t want a raw LLM. You want an agent integrated with a legal database, in such a way that it can’t cite cases which don’t exist in the database. And it can’t generate a direct quote from a case unless that text actually occurs in the case.
You still need a human lawyer familiar with the law to pick up on subtler errors, but better technology can prevent the grosser ones. And the subtler errors (e.g. misrepresenting what a case says through partial or out of context quoting) is the kind of error human lawyers sometimes make too - like every other profession, lawyers vary greatly in their competence, and often only the grosser cases of incompetence incur sanctions
Much like ChatGPT, it can only handle recall and summation — and even then, I don’t fully trust it because it often misses key ideas.
And much like ChatGPT, it can’t do anything coherent at length without a lot of help and working around its faults.
And seems entirely unaware of similar sounding words having distinct legal meanings. Which is not good.
I wonder to what extent that’s due to current inherent limitations of the technology, and to what extent it is due to quality of implementation issues.
It is hard to say because (I presume) there is only limited public information on how it is actually implemented.
e.g. which LLM is it using? How much fine-tuning has been done? Are they using other potentially helpful techniques such as guided sampling? Or breaking down the task into parts and having multiple agents each specialised to handle one particular part?
> Much like ChatGPT
When you compare it to ChatGPT, do you mean GPT3.5 or GPT4?
I also would guess that certain areas of law (especially criminal law) may be prone to triggering “safeguards” which result in poorer performance than if those safeguards were absent. Arguing that what your client did was legal (even be it ethically unsavoury) is an essential part of a lawyer’s job
It's different than a Lawyers case, where the facts require manually cross referencing. In our case, that verification can come via directly and immediately running the code. How good the actual code is varies, but the fundamental difference remains.
Whether that's better, or useful, probably differs by situation.
in many cases it's way harder
I clearly find value in LLMs to some degree, but speeding up my coding is not yet one.. i'd love to, but i just don't understand. I can only imagine typing more in explanation for the LLM than it would take me to write it to begin with.
For instance I know I'm good at breaking up large tasks into smaller ones, but because planning and executive functioning is by far my worst energy consuming skill (adhd) I can save a lot of energy (but not time) by approaching it conversationally. I'd often use colleagues for the same thing, but then my productivity costs 2 salaries
Similarly, I have some trouble with memory and cognitive speed when I'm tired which is unfortunately often the case due to my health, I know well enough what I kinda "want" to do and I let the AI generate something that comes near what I need and I can work from there once I have the right starting point.
Just my personal experience I'd wanted to share.
I think the real problem with this is that people aren't differentiating these different types of work when they give these numbers.
1. After working for many years in Java I needed to build a service. I spent few days designing it and then a month on implementation. I used DBs and libraries I knew very well. I didn't need to access google/stackoverflow, I didn't need to look up names of (std)lib methods and their parameters and if something wasn't working it was fairly obvious what I needed to change.
2. Recently I wanted to create a simple page which fetched a bunch of stuff from some URLs and showed the results, simple stuff. But with React since that was what frontend team was using. I never used React and rarely touched web in the recent years. Most of my time was spent googling about React and how exactly CORS/SOP work in the browsers, and with polishing it took a couple of days.
I'm pretty sure that in case 1) AI wouldn't help me much. Maybe just as a more fancy code completion.
In case 2) AI would probably be a significant time save. I could just ask it to write some draft for me and then I could make few tweaks, without having to figure out React.
But somehow nobody quantifies their experience with the languages/tools when they are using AI - I'm sure there's a staggering difference between 1 month and 10 yoe.
Without getting into touchy questions of quality and talent and the whole 10x trope, there are a lot of working engineers out there that produce code a lot more slowly than others and that work on more common problems than others. My sense is copilot-like products provide the biggest boon to those people and are a lot harder for people who are more naturally prolific or who work on more esoteric things.
Would this have been difficult for me to just write? No. I wouldn't call it difficult. But it would have involved effort, which is a resource I'm happy to conserve. It's like the difference between a sandwich you make and a sandwich you ask someone to make.
Then I asked it to generate the integer ranges which correspond to each bucket. It screwed that one up, which I found out by copying the function to a scratch file and trying some representative values in the REPL. So I told it to iterate all the values between in and max and generate the ranges that way. That one worked. Net of less effort, plus I had a test for it which I could copy-paste from the REPL to the test suite.
Faster? Maybe, maybe not. But I'm rate-limited by gumption, not minutes. When it's easier to describe the function than write it, and it's simple enough that the robot won't screw the pooch, I hand it off to the LLM. It's a great addition to the toolkit.
The exact opposite will be true and the funny (sad?) part is that they will lack the skills necessary (because they got lazy and over-trusted the AI) to fix mistakes/incompatible solutions.
Maybe my work is not too typical though, I spend only a fraction of my time actually typing in code. And I do eliminate the need for boilerplate through other means (picking frameworks/libraries that are a good fit for the problem, refactoring, meta programming, scripts, suitable tool chain etc).
Is that enough to be competitive?
I’m not sure — but at scale that would be a 5% reduction in headcount for the same work, or ~$12M/yr for every 1,000 engineers.
If you can figure out how to get 2-3 hours more coding done a week, we’re talking real gains.
People are starting to rely on LLMs to do almost everything they've been hired to do for them.
Did they help people write lots of effective code faster? Yes.
Did they breed a generation of people with little intuition around how memory works? Also yes.
It has probably the same productivity boost that intellisense gave back when it came out. Which is good, but still marginal. Certainly not replacing anyone’s job.
I paste in functions / classes I want to write unit tests for. Paste in a sample unit test, and it does a solid job of writing tests for it in the same manner as in the sample.
For unit tests, you don't even need the multi-step coverage optimization in this article. You just manually inspect, adjust it etc.
But yeah, if you're just working solo on basic stuff and need to protect against off-by-one errors and accidental mutations in later refactoring, it's a great tool. You'd never write duly antagonistic tests for your own code anyway.
The Power of AI™ can figure out the true intent of the code by looking at the initial (and, potentially, buggy) implementation, and help the programmer by generating edge test cases where the code doesn't produce correct results.
The programmer will easily tell those test cases from the ones where the AI did a mistake and generated a flawed test case because the AI doesn't make such silly mistakes; clearly, it's the programmer's code that needs to be corrected.
In fact, the AI would do a better job at that, too, which clearly speeds up the development cycle.
The correct way to use the tool is to let the AI both generate the test cases and modify the code so that it would pass the tests it generates.
After all, if the AI can't figure out what you wanted to do in the first place — how can you?
Of course, there's more to it.
Whichever problems you run into can be surely attributed to writing the code in an AI-unfriendly way.
In the past, we had the adage that the code is read more times than it's written. This is still true, but we need to abandon the old habit habit of having a human reader in mind.
Just like you rearrange your furniture to make your house more accessible to the robot vacuum cleaner, you need to write the code with the AI in mind.
When you write a Google query or a prompt for ChatGPT (effectively the same thing anyway), you don't write it like you'd talk to a person.
You're going to have to write code the same way to be truly effective, and think a little bit like AI to get the most use out of it.
That might sound like a lot of work to get code that does what you want.
But of course, that's not the case.
Just use the AI for this.
The secret sauce is: the robot doesn't run the tests. It digests the code and writes tests for it.
So it doesn't know if they'll pass or not. And sure enough, some of them don't, because the code was buggy.
I've seen some awful code come out of ChatGPT, but never a bad test. Good tests are short, which makes them hard to screw up. It has a reasonable grasp on what an edge case is as well.
(edit: I didn't downvote you, but I do think your claim is over-broad)
There's a 100% likelihood that OpenAI and other LLMs providers are going to cooperate with intelligence agencies looking to deploy their backdoors to businesses across the world.
It's the perfect supply chain attack vector with complete deniability.