has anyone found a good way to improve code quality? just wondering -- LOC does seem like the wrong metric, but the code LLMs write is just too verbose
- keep prompts focused on atomic tasks.
- use expert prompting[0] when possible.
- require coding agents to verify changes.
- require coding agents to create/update unit tests with 100% coverage.
- use git to commit/revert atomic tasks manually.
- leverage planning capabilities to review instead of recover.
- consider using something like Karpathy guidelines[1].
- leverage Constraint Programming[2] concepts when
formulating prompts.
0 - https://arxiv.org/pdf/2305.14688This scales about as well as it sounds like it would.
Multi-model review does a decent job identifying things they’re outright wrong. The resulting code still doesn’t feel elegant writ large.
If you want good output, it seems that iterating on the output is inferior to providing better input inclusive of code examples. And by the time you’ve made all the decisions that go into that, something like ponytail is superfluous.
(All that said, I have ponytail installed in most harnesses.)