Generating Code Without Generating Technical Debt?
sourcery.ai
sourcery.ai
0. ask it to ask you clarifying questions, which it rarely does by default when you're prompting it to do a task; often it has a few questions, which you can efficiently update your original prompt to cover.
1. make it generate a test-suite; have it iteratively generate new tests which don't overlap with the old ones.
2. ask GPT-4 explicitly to identify improvements, edge-cases and bugfixes, and then go through the list one by one having it rewrite for each one; not infrequently, a fancy rewrite will fail the test-suite from #1, but given the failing test-case, GPT-4 can fix it.
3. once it is clean and it's either done the list or you've disapproved the suggestions, and the test-suite is passing, ask it to write a summary design doc at the beginning.
With all this, you're set up for fairly maintainable code: with the test-suite and the up-front design doc, future LLMs can handle it natively and well, and humans should be able to read it easily after #2 has finished, so you don't need to care where it came from or try to track 'taint' through all future refactorings or usage - GPT-4 can write pretty readable human-like code, it just doesn't necessarily do it the best way the first time (also like a human), so you have to apply inner-monologue ideas.
As an example, in my day job (https://speakeasyapi.dev), we sell code generation products using the OpenAPI specification to generate downstream artefacts (language SDKs, terraform providers, markdown documentation). The determinism makes it useful — API updates propagate continuously from server code, to specifications, then to the SDKs / providers / docs site. There are no breaking changes because the pipeline is deterministic and humans are in control of the API at the start. The code generation itself is just a means to an end : removing boilerplate effort and language differences by driving it from a source of truth (server api routes/types). Continuously generated, it is not debt.
We’ve put a lot of effort into trying to make an LLM agent useful in this context. However giving them control of generated code directly means it’s hard to keep the “no breaking changes”, and “consistency” restrictions that’s needed to make code generation useful.
The trick we’ve landed on to get utility out of an LLM in a code generation task, is to restrict it to manipulating a strictly typed interface document, such that it can only do non-breaking things to code (e.g. adjust comments / descriptions / examples) by making changes through this interface.
Well said @ThomasRooney.
There may be other contexts where pure LLM codegen could work well, but I haven't really encountered them personally yet.
The first version of our product (https://grit.io) was entirely LLM-powered. It was very easy to get started with, but reliability was low on enterprise-scale codebases.
Since then, we've switched to a similar approach: using LLMs to manipulate a verifiable interface, but making actual changes through deterministic code.
The first, if used, implies the code can be contextualized through another employee at the office (the one who wrote it).
The second, if used, implies you may not be able to get an answer as to why the code was written as it was, since an LLM generated it.
Unfortunately, it doesn't work with GitHub Copilot's style, which assumes really fine-grained interaction.
At the end of the day, whether you copy-pasted the code from StackOverflow, let Copilot generate it or M-X-Butterfly'd it in yourself doesn't make a difference as you still need to understand thoroughly what it does.
I mean `urlparse(url).scheme = 'https'` is so literal and straightforward that it really ought not be it's own separate function. Adding a wrapper that does nothing but add more code and rethrow the same error is just generating technical debt by writing a function that ought not be.
Arguably all code is technical debt, but so far the world isn't kind enough to bend to our will without any code.
* Why would it print a warning to stdout for invalid input?
* The function expects an URL. So if it gets sth else, it should not catch that exception on its own. The exception is right.