We have LLMs with hundreds of thousands of tokens context windows and prompt caching that makes using them affordable. Why don’t we just stuff the whole code base in the context window?
For the time being, indexing and retrieving a good collection of 10-20 code chunks is more effective/performant in practice.
https://github.com/Storia-AI/repo2vec/blob/1864102949e720320...