AST-grep(sg) is a CLI tool for code structural search, lint, and rewriting
github.com
github.com
My tool let's you grep a regex as usual, but shows you the matches in a helpful AST aware way. It works with most popular languages, thanks to tree-sitter.
It uses the abstract syntax tree (AST) of the source code to show how the matching lines fit into the code structure. It shows relevant code from every layer of the AST, above and below the matches. It's useful when you're grepping to understand how functions, classes, variables etc are used within a non-trivial codebase.
Here's a snippet that shows grep-ast searching the django repo. Notice that it finds `ROOT_URLCONF` and then shows you the method and class that contain the matching line, including a helpful part of the docstring. If you ran this in the terminal, it would also colorize the matches.
django$ grep-ast ROOT_URLCONF
middleware/locale.py:
│from django.conf import settings
│from django.conf.urls.i18n import is_language_prefix_patterns_used
│from django.http import HttpResponseRedirect
⋮...
│class LocaleMiddleware(MiddlewareMixin):
│ """
│ Parse a request and decide what translation
│ object to install in the current thread context.
⋮...
│ def process_request(self, request):
▶ urlconf = getattr(request, "urlconf", settings.ROOT_URLCONF)
[0] https://github.com/paul-gauthier/grep-astThe command line tool is a thin wrapper around the `TreeContext` class, whose purpose is show you a set of "lines of interest" in the context of the entire AST. This all exists because my other project aider [0] uses TreeContext to display a repository map [1] so that GPT-4 can understand how the most important classes, methods, functions, etc fit into the entire code base of a git repository.
But it was easy to make a CLI interface to grep lines of interest and display them with TreeContext, and it turned out to be quite useful.
The TreeContext class is line-oriented, and is mainly interested in tracking language constructs whose scope spans multiple lines. Typically these are things like classes, methods, functions, loops, if/else constructs, etc. Given a line of interest, we look at all the multi-line scopes which contain it. For each such multi-line scope, we want to display some "header" lines to provide context.
In this example, the match for "two" is contained in the multi-line scopes of a method and a class. So we print their headers.
$ grep-ast two example.py
⋮...
│class MyClass:
│ "MyClass is great"
⋮...
│ def print2(self):
▶ print("two")
⋮...
The trick is how to determine the header for each multi-line scope? It's not ideal to just use the first line. For example, it's nice that the header for the class included the docstring as well as the bare `class MyClass:` line.For any multi-line scope, we look at all the other AST scopes which start on the same line. We take the smallest such co-occurring scope, and declare the header to be the lines that it spans. For a simple method like `def print2(self):`, that's all that gets picked up.
But a complex method like `print1()` below picks up all the lines which are part of its full function signature:
$ grep-ast one example.py
⋮...
│class MyClass:
│ "MyClass is great"
⋮...
│ def print1(
│ self,
│ prefix,
│ suffix,
│ ):
⋮...
▶ print(f"{prefix} one {suffix}")
⋮...
It's a heuristic, but it seems to work well in practice.See also parsertl-playground[1] for online edit/test grammars.
[0]https://github.com/BenHanson/gram_grep
[1]https://mingodad.github.io/parsertl-playground/playground/
I've spent a lot of time trying to find similar tools, and even list them in the README, but `AST-grep` did not come up! I was a bit confused, as I was sure such a thing must exist already. AST-grep looks much more capable and dynamic, great work, especially around the variable syntax.
One minor comment: I personally found the first Python example involving a docstring a little hard to parse (ha). It may show a variety of features, but in particular I found that it was hard to spot at a glance what had changed.
If you could use diff formatting or a screenshot with color to show the differences it would make it much easier to follow. If I get around to using it later today, I might submit a PR for that. :)
Thank you for the feedback! That sounds good, I'll add that.
Do you use tree-sitter for the AST part also?
https://github.com/go-go-golems/oak
I initially hope the queries would be more powerful, but they are really not. You can write queries and a resulting template in a yaml file. The program will scan a list of repositories for all these YAML files, and expose them as command line verbs.
Here is one to find go definitions:
https://github.com/go-go-golems/oak/blob/main/cmd/oak/querie...
This can then be run as:
oak go definitions /home/manuel/code/wesen/corporate-headquarters/geppetto/pkg/cmds/cmd.go
type GeppettoCommandDescription struct {
Name string `yaml:"name"`
Short string `yaml:"short"`
Long string `yaml:"long,omitempty"`
Flags []*parameters.ParameterDefinition `yaml:"flags,omitempty"`
Arguments []*parameters.ParameterDefinition `yaml:"arguments,omitempty"`
Layers []layers.ParameterLayer `yaml:"layers,omitempty"`
Prompt string `yaml:"prompt,omitempty"`
Messages []*geppetto_context.Message `yaml:"messages,omitempty"`
SystemPrompt string `yaml:"system-prompt,omitempty"`
}
type GeppettoCommand struct {
*glazedcmds.CommandDescription
StepSettings *settings.StepSettings
Prompt string
Messages []*geppetto_context.Message
SystemPrompt string
}
While I can use it for good effect for LLM prompting as is, I really would like to add a unification algorithm (like the one in Peter Norvig's Prolog compiler) to get better queries, and connect it to LSP as well.A lot like your project but with more of a focus on supporting data structures for incremental editing of programs. Kind of a DOM for code.
You can install the CLI utility in four different ways: https://ast-grep.github.io/guide/quick-start.html#installati...
# via Homebrew
brew install ast-grep
# via Cargo
cargo install ast-grep
# via npm
npm i @ast-grep/cli -g
# via pip
pip install ast-grep-cli
# I tested and pipx works too:
pipx install ast-grep-cli
I really like this - it means the tool is available to people with familiarity of any of those four distribution mechanisms.You can also download pre-built binaries from their releases page: https://github.com/ast-grep/ast-grep/releases/tag/0.14.2
On top of that, they offer API bindings for it in three different languages:
- Rust (not yet stable): https://docs.rs/ast-grep-core/latest/ast_grep_core/
- JavaScript/TypeScript: https://ast-grep.github.io/guide/api-usage/js-api.html
- Python: https://ast-grep.github.io/guide/api-usage/py-api.html
It's rare to see a tool/library offer this depth of language support out of the box.
The wheel just contains the two binaries (sg and ast-grep) and no Python code:
$ unzip -l ast_grep_cli-0.14.2-py3-none-macosx_10_7_x86_64.whl
Archive: ast_grep_cli-0.14.2-py3-none-macosx_10_7_x86_64.whl
Length Date Time Name
--------- ---------- ----- ----
6207 12-03-2023 07:34 ast_grep_cli-0.14.2.dist-info/METADATA
102 12-03-2023 07:34 ast_grep_cli-0.14.2.dist-info/WHEEL
1077 12-03-2023 07:34 ast_grep_cli-0.14.2.dist-info/license_files/LICENSE
1077 12-03-2023 07:34 ast_grep_cli-0.14.2.dist-info/license_files/LICENSE
32865880 12-03-2023 07:34 ast_grep_cli-0.14.2.data/scripts/sg
32865880 12-03-2023 07:34 ast_grep_cli-0.14.2.data/scripts/ast-grep
639 12-03-2023 07:34 ast_grep_cli-0.14.2.dist-info/RECORD
--------- -------
65740862 7 files
I haven't seen pip and wheels used to distribute a purely binary tool like this before.https://github.com/astral-sh/ruff/blob/main/.github/workflow...
One of the downsides of the simplicity is that rules are written in yaml. This works nicely for simple rules, but if you try to save a complex migration as a rule, you end up programming in YAML (which is very hard).
For my similar tool we decided to build a full query language for matching code, called GritQL: https://docs.grit.io/tutorials/gritql
For some reason my simple query pattern 'GetValueForKey' isn't a match. Is there a way to get the 'sg' CLI to output the AST for some line of my target file so I can see what kind of pattern I need to write?
We use gpt-4 to generate ast-grep patterns to deep-dive and verify pull-request integrity. We just rolled this feature out 3 days back and are getting excellent results!
Comments such as these are powered by AI-generate ast-grep queries: https://github.com/amorphie/contract/pull/100#discussion_r14...
Say, like performance. tree-sitter's initial parsing speed can be easily beaten by a carefully hand-crafted parser. Tree-sitter, written in C, has a similar JavaScript parsing speed as Babel, a JS-based parser. See the benchmark https://dev.to/herrington_darkholme/benchmark-typescript-par...
(4 years ago, I was more willing to put up with enterprise licensing. But in the last two years, I've seen way too many enterprise vendors try to squeeze every penny they can get from existing clients. An enterprise sales process now often means "Expect 30% annual price hikes once you're in too deep to back out." The lack of easy VC money seems to have made some enterprise vendors pretty desperate.)
There's also an open source "semgrep" project here: https://github.com/semgrep/semgrep. But this seems to be basically a vulernability scanner, going by the README.
Whereas AST-grep seems to focus heavily on things like:
1. One-off searching: "Search my tree for this pattern."
2. Refactoring: "Replace this pattern with this other pattern."
AST-grep also includes a vulnerability scanning mode like semgrep.
It's possible that semgrep also has nice support for (1) and (2), but it isn't clearly visible on their corporate landing page or the first open source README I found.
Semgrep's vulnerability scanning is much more advanced, mostly for enterprise security usage.
TLDR; I designed ast-grep to be on different tracks than semgrep.
Semgrep is for security and ast-grep is for development.
First and foremost, I have always been in awe of semgrep. Semgrep's documentation, product sites and Padioleau's podcast all gave me a lot of inspiration. Using code to find code is such a cool idea that I never need to craft an intricate regex or write a lengthy AST program. sgrep and patch from https://github.com/facebookarchive/pfff/wiki/Sgrep have helped me a lot in real large codebases.
When I used semgrep as a software engineer, instead of a security researcher, I found semgrep has not touched too much on routine development works. I can use `semgrep -e PATTERN` but the Python wrapper is not too fast compared to grep. While pattern is cool, it cannot precisely match some syntax nodes. (example, selecting generator expression in Semgrep is very hard). It also does not have API to find code programmatically.
I have also a short summary for tool comparison. https://ast-grep.github.io/advanced/tool-comparison.html
* Semgrep is security focused. It has many advanced static analysis features in its core product, such as dataflow analysis, symbolic propagation, and semantic equivalence, all of which are useful for security analysis. They are not available in ast-grep. * Semgrep’s pattern syntax also prefers matching more potentially vulnerable semantics than matching precise syntax. Semantic level information is the better level of abstraction for security model. ast-grep, on the other hand, sticks to faithfully translating users' queries syntactically. * Semgrep has a one-off search and rewrite feature, but it is not its primary focus. The CLI is a bit slow compared to other tools. ast-grep strives to be a fast CLI tool. * Semgrep has a product matrix for vulnerability detection: detecting secrets, supply chain vulnerabilities, and cross-file detection. It also has a plethora of security rules in the registry. These features will not be included in ast-grep.
[1]: https://learn.microsoft.com/en-us/azure/devops/project/searc...
I did this:
cargo install ast-grep
Then I'm searching in my code with: $ sg --pattern 'catch (...) { $$$ARGS }' --lang c++
./Server/TCPHandler.cpp
657│ catch (...)
658│ {
659│ state.io.onException();
660│ exception = std::make_unique<DB::Exception>(ErrorCodes::UNKNOWN_EXCEPTION, "Unknown exception");
661│ }
But there should be more than one result.It also does not work in Playground: https://ast-grep.github.io/playground.html#eyJtb2RlIjoiUGF0Y...
Pattern code must be valid code that tree-sitter can parse.[0]
You can use this syntax `try { $$$CODE } catch ($$$E) { $$$HANDLE }`[1]
[0] https://ast-grep.github.io/guide/pattern-syntax.html#pattern...
[1] https://ast-grep.github.io/playground.html#eyJtb2RlIjoiUGF0Y...
I find the best uses for it being answering questions like 'how is this function called' and 'what is this struct's definition':
# find struct definition
$ syns 'struct Span {}'
[./src/psi.rs:21-26]
pub struct Span {
/// Starting byte index of the span.
pub lo: usize,
/// End byte index of the span.
pub hi: usize,
}
# How is this function called? $ syns 'ast_match()'
[./src/query.rs:91] if &op.ty == op1 && self.ast_match(content1, &[\*start]).is_some() {
[./src/query.rs:125] .flat_map(move |tts| self.ast_match(tts, &[self.machine.initial]))
It's also pretty fast for most repositories, as an extreme case with the kernel source: time syns 'kmalloc()' > /dev/null
syns 'kmalloc()' > /dev/null 61,82s user 0,90s system 99% cpu 1:02,73 total
Most other repositories print all results pretty much instantly.It worked great for the use case I built it around initially but I think it would need a scripting/logic component to generalise to any conceivable refactoring.
I certainly hope some excellent AST-based CLI code search tools come to exist; hopefully this is one of them.
https://ast-grep.github.io/playground.html#eyJtb2RlIjoiQ29uZ...
I have the same problem also, haha, https://x.com/hd_nvim/status/1667059966111547392
would your hoped-for tool recognise that
1
and sin(x)^2 + cos(x)^2
are the same? (I think that identity holds, but if not you get the picture)They did mention code, and said "similarity" rather than equivalence.
But, as a trivial example, two different pieces of code can compile down to the same AST, or bytecode, or assembler.
I suspect that things like "these two functions both start with the same conditional+early return" would be more useful to -me- given the sort of things I tend to be working on. Also a 'fuzzy possible copy+paste detector' in general to help identify refactoring targets.
It also strikes me that something that was mostly 'just' a structure-aware diff so e.g. you got diffs within-if-body and similar but I'm now into vigorous hand waving because it's been ages since I've thought about this and I probably need more coffee.
I -did- do a pure maths degree many years ago but I don't generally seem to end up working on computational code
Tree Sitter also leaves a lot to be desired for C++ editing, but that's a special problem.
Was it possible to use it entirely as a CLI tool without any YAML 6 months ago?
something:
subkey: |
I can put any characters I like in here
And they "won't be messed up" by anything
Because they are part of a multi-line stringI don't know your exact use case but it sounds a little bit hard to me. I guess you want to change things like this.
```cpp
add_macro(a, b, c, d); // before
a + b + c + d; // after
```
It isn't that easy to do in pattern syntax. The pattern/replacement must support repetition and it isn't straightforward as far as I can tell from Rust's macro[0].
That said, if you need to support use cases like this, ast-grep has Python API to support programmatic usage[1].
[0] https://veykril.github.io/tlborm/decl-macros/macros-methodic...
I have been improving ast-grep's documentation since the release of the project.
It is arguably not that good/abundant for resources compared to eslint/babel. But it has improved a lot (say, example page[0] and deep dive page[1]). The doc site is also carefully crafted to make it accessible, compared to libcst or other similar projects.
I also appreciate blogs/introductions to ast-grep if the community can help! Let me know if there are issues/problems when using ast-grep!
-l ts
And an -l rs
In the examples. Those target typescript and rust. Looks like it’s built in tree-sitter, so presumably any language that supports that should workIf you leave off the language command line option it detects the language from the extension on your files.
These provide a nice frontend for writing simple rules, but I would not want to (essentially) write an entire transpiler in yaml.
For Python->JavaScript, you likely want a transpiler focused specifically on that.
Unfortunately, every such effort eventually hits serious limits in the emergent complexity for languages. There's a reason most of the SOTA techniques ML-based.