Asking it about these things sounds like it would result in questionable, at best, responses.
(A) wait for future models that are planned to have much longer contexts
(B) fine tune a model on this specific code base, so the code base is part of the training data not the prompt
(C) Break the problem up into multiple invocations of the model. Feed each source file in separately and ask it to give a brief plain text summary of each. Then concatenate those summaries and ask it questions about it. Still probably not going to perform that well, but likely better than just giving it a large code base directly
Another issue is that, even the best of us make mistakes sometimes, but then we try the answer and see it doesn’t work (compilation error, we remembered the name of the class wrong because there is no class by that name in the source code, etc). OOTB, ChatGPT has no access to compilers/etc so it can’t validate its answers. If one gave it access to an external system for doing that, it would likely perform better.
[0] https://mobile.twitter.com/goodside/status/15988746742046187...
Of course, that doesn't tell you whether the machine understanding will be useful or not
https://platform.openai.com/docs/guides/code
I’d you’re interested in trying the very cheap models behind ChatGPT, you may want to have a look at langchain and langchain-chat for an example of how to build a chatbot that uses vectorized source code to build context-aware prompts.
I actually built a slack bot for work and daily ask it to refactor code or "write jsdocs for this function"
We're two years in but everything still feels super early given how quickly the fundamentals are improving. Would love your feedback - https://bloop.ai
I assume it would “understand” more popular open source frameworks.
For code completion for example, you can just train it with a whole bunch of code.
But to explain large code bases, you need to train it with both large codebases and explanations. As far as I know, there are no such explanations available.