Chunking 2M files a day for code search using syntax trees
docs.sweep.dev
docs.sweep.dev
I do see them generate code that fails the type checker though.
Also syntax seems a lot easier to understand for them than semantics/logic. If you've used GPT-4 it almost never makes syntax errors. Logical errors on the other hand...
It also frequently makes undefined variables and the like, however.
Looking forward to reading through your docs and repo later tonight to see how you’re addressing issues like this.
I'm building a bot that's building itself, so it doesn't have to support large legacy code bases with different languages.
But it requires parsing AST and language specific instructions. And things like metaprogramming or macros could cause some hairy confusion.
All of these factors don't hurt my use case.
Gut feeling doesn't account for much though - I'm working on an evals system to be able to quantify system performance. It won't be cheap to run.
It could easily be that your method is superior.
I'm also curious, how are you guys evaluating the performance of your models?
https://github.com/paul-gauthier/aider/blob/d5a7aac560d4584d...
Also, syntax level information are local or short sighted, it is called context-free grammar for a reason. My own observation with playing with those coding LLMs all day, is that they most likely had acquired the grammar themselves implicitly. Providing explicit regularization by enforcing grammar, is going to provide at best modest benefits, and that is dependent on good that parser is written, in many cases, it is not a given.
Cool project btw!
What is interesting to me is the sharp increase in forks, which is a good indicator that others will contribute code in the near future.
Full Disclosure: This is my tool
Edit: I guess it is just the app I am using, it looks fine on my mobile browser but odd in the app
Edit: Here's some of the papers: https://arxiv.org/abs/1911.09983 and https://aclanthology.org/2021.findings-acl.384.pdf
I’ll try to explain something I’m thinking, it comes down to a type of agglutination.
https://en.wikipedia.org/wiki/Agglutination
The ASTs need to become sequences of tokens for LLMs to work well. Also the embedding space is related, this should all be specifically optimized for code, not based on general human language .
Of course the description of the code would be english. But my point is that a subspace of the encoding should be specifically designed for an agglutination based sequential expression of ASTs.
Not sure any of that makes sense to someone with more expertise in this space than me.