Lately, I've been playing around with LLMs to write code. I find
that they're great at generating small self-contained snippets.
Unfortunately, anything more than that requires a human...
I have been working on this problem quite a bit lately. I put together a writeup describing the solution that's been working well for me:https://aider.chat/docs/ctags.html
The problem I am trying to solve is that it’s difficult to use GPT-4 to modify or extend a large, complex pre-existing codebase. To modify such code, GPT needs to understand the dependencies and APIs which interconnect its subsystems. Somehow we need to provide this “code context” to GPT when we ask it to accomplish a coding task. Specifically, we need to:
1. Help GPT understand the overall codebase, so that it can decifer the meaning of code with complex dependencies and generate new code that respects and utilizes existing abstractions.
2. Convey all of this “code context” to GPT in an efficient manner that fits within the 8k-token context window.
To address these issues, I send GPT a concise map of the whole codebase. The map includes all declared variables and functions with call signatures. This "repo map" is built automatically using ctags and enables GPT to better comprehend, navigate and edit code in larger repos.
The writeup linked above goes into more detail, and provides some examples of the actual map that I send to GPT as well as examples of how well it can work.