Refactoring Python with Tree-sitter and Jedi
jackevans.bearblog.dev
jackevans.bearblog.dev
cat file.py | srgn --py def --py identifiers 'database' 'db'
will replace all mentions of `database` inside identifiers inside (only!) function definitions (`def`) with `db`.An input like
import database
import pytest
@pytest.fixture()
def test_a(database):
return database
def test_b(database):
return database
database = "database"
class database:
pass
is turned into import database
import pytest
@pytest.fixture()
def test_a(db):
return db
def test_b(db):
return db
database = "database"
class database:
pass
which seems roughly like what the author is after. Mentions of "database" outside function definitions are not modified. That sort of logic I always found hard to replicate in basic GNU-like tools. If run without stdin, the above command runs recursively, in-place (careful with that one!).Note: I just wrote this, and version 0.13.2 is required for the above to work.
There’s no need to manually iterate through the tree, and use if statements to select nodes. Instead you can just write a couple of simple queries (and even use treesitters web UI to test the queries), and have the treesitter just provide all the nodes for you.
https://tree-sitter.github.io/tree-sitter/using-parsers#patt...
But the syntax for variable binding is idiosyncratic and the opposite of normal pattern languages. Writing “x” doesn’t bind the thing at the position to the variable x; instead, you have to write e.g. foo @x to bind x to the child of type foo. Insanely, some Scheme dialects use @ with the exact opposite semantics!! There’s also a bizarre # syntax for conditionals and statements.
Honestly there isn’t really an excuse for how weird they made the pattern syntax given that people have spent decades working on pattern matching for everything from XML to objects (even respecting abstraction!). I’ve slowly been souring on treesitter in general, but paraphrasing Stroustrup: there are things people complain about, and then there are things nobody uses.
It’s as much a Scheme dialect as WASM’s S-expression form is a Scheme dialect.
Treesitter’s query syntax is slightly understandable in the sense that having x match a node among siblings of type x works well for extracting values out of sibling lists. Most conventional pattern syntaxes struggle with this, e.g. how do you match the string “foo” inside of a list of strings in OCaml or Rust without leaving the match expression and resorting to a loop?
But you could imagine a syntax-rules like use of ellipses …. There’s also a more powerful pattern syntax someone worked on for implementing Scheme-like macros in non-S-expression based languages whose name escapes me right now.
I guess what I’m looking for is something that
* can carry out the kind of refactorings usually reserved for an IDE
* has a concise language for representing the refactorings so new ones can be built quite easily
* can apply the same refactoring in multiple places, with some kind of search language (addressing the task of renaming a test parameter in multiple methods)
* ideally does this across multiple programming languages
This is trivial with codegen.com. Syntax below:
# Iterate through all files in the codebase
for file in codebase.files:
# Check for functions with the pytest.fixture decorator
for function in file.functions:
if any(d.name == "fixture" for d in function.decorators):
# Rename the 'db' parameter to 'database'
db_param = function.get_parameter("db")
if db_param:
db_param.set_name("database")
# Log the modification
print(f"Modified {function.name}")
Live example: https://www.codegen.sh/codemod/4697/public/diffIs each file getting parsed individually with tree-sitter or how is the codebase object constructed?
This enables APIs such as `function.call_sites`, `symbol.usages`, `class.parent_classes`, and more!
- perl -pi -e 's/foo/bar/g' files
"-pi" means "in place edit" so it will change the files in place. If you have a purely mechanical change like he's doing here it's a very reasonable choice. If you're not as much of a cowboy as I am, you can specify a suffix and it will back the files up, so something like
perl -p -i.bak -e 's/db/database/g' py
For example then all your original '.py' files will be copied to '.py.bak' and the new renamed versions will be '.py'
For vim users (I know emacs has the same thing but I don't remember the exact invocation because it has been >20years since I used emacs as my main editor) it's worth knowing the "global" command. So you can execute a particular command only on lines that match some regex. So say you want to delete all the lines which mention cheese
:%g/cheese/d
Say you want to replace "db" with "database" but only on lines which start with "def"
:%g/^def/s/db/database/
OK cool. Now if you go 'vim *py' you can do ":argdo g/^def/s/db/database/ | update" and it will perform that global command across all the files in the arg list and save the ones which have changed.
But in this specific situation it was tricky to handle situations with things spanning over multiple lines + preventing accidental renames.
https://blog.moertel.com/posts/2013-02-18-git-second-order-d...
[1] like that's even possible in this situation
> every instance of a pytest fixture
Although it's probably good enough for 99% of the use cases, and any extra accidental renames could be reverted when you look at the diff.
Maybe it could be covered with a multi line regex using `\_.`
If you're looking for a higher level interface, GritQL[0] is built on top of tree-sitter and could handle the same refactor with this query:
language python
`def $_($_): $_` as $func where $func <: contains `database` => `db`
[0] https://github.com/getgrit/gritql pattern: |-
def $FUNC(..., database, ...):
$...BODY
fix: |-
def $FUNC(..., db, ...):
$...BODYOr libCST: https://github.com/Instagram/LibCST docs: https://libcst.readthedocs.io/en/latest/ :
> LibCST parses Python 3.0 -> 3.12 source code as a CST tree that keeps all formatting details (comments, whitespaces, parentheses, etc). It’s useful for building automated refactoring (codemod) applications and linters.
libcst_transformer.py: https://gist.github.com/sangwoo-joh/26e9007ebc2de256b0b3deed... :
> example code for renaming variables using libcst [w/ Visitors and Transformers]
Refactoring because it doesn't pass formal verification: https://deal.readthedocs.io/basic/verification.html#backgrou... :
> 2021. deal-solver. We released a tool that converts Python code (including deal contracts) into Z3 theorems that can be formally verified
> Pymode can rename everything: classes, functions, modules, packages, methods, variables and keyword arguments.
> Keymap for rename method/function/class/variables under cursor
let g:pymode_rope_rename_bind = '<C-c>rr
python-rope/ropevim also has mappings for refactorings like renaming a variable: https://github.com/python-rope/ropevim#keybinding : C-c r r :RopeRename
C-c f find occurrences
https://github.com/python-rope/ropevim#finding-occurrencesTheir README now recommends pylsp-rope:
> If you are using ropevim, consider using pylsp-rope in Vim
python-rope/pylsp-rope: https://github.com/python-rope/pylsp-rope :
> Finding Occurrences: The find occurrences command (C-c f by default) can be used to find the occurrences of a python name. If unsure option is yes, it will also show unsure occurrences; unsure occurrences are indicated with a ? mark in the end. Note that ropevim uses the quickfix feature of vim for marking occurrence locations. [...]
> Rename: When Rename is triggered, rename the symbol under the cursor. If the symbol under the cursor points to a module/package, it will move that module/package files
SpaceVim > Available Layers > lang#python > LSP key Bindings: https://spacevim.org/layers/lang/python/#lsp-key-bindings :
SPC l e rename symbol
Vscode Python variable renaming:Vscode tips and tricks > Multi cursor selection: https://code.visualstudio.com/docs/getstarted/tips-and-trick... :
> You can add additional cursors to all occurrences of the current selection with Ctrl+Shift+L. [And then rename the occurrences in the local file]
https://code.visualstudio.com/docs/editor/refactoring#_renam... :
> Rename symbol: Renaming is a common operation related to refactoring source code, and VS Code has a separate Rename Symbol command (F2). Some languages support renaming a symbol across files. Press F2, type the new desired name, and press Enter. All instances of the symbol across all files will be renamed
Because pytest requires a preprocessing step, renaming fixtures is tough, and also for jupyter notebooks %%ipytest is necessary to call functions that start with test_ and upgrade assert keywords to expressions; e.g `assert a == b, error_expr` is preprocessed into `assertEqual(a,b, error_expr)` with an AssertionError message even for comparisons of large lists and strings.
Here are two approaches in Emacs.
https://emacs.stackexchange.com/a/69571
https://rigsomelight.com/2010/02/14/emacs-interactively-find...
there's tsmod at least https://github.com/WolkSoftware/tsmod
i've heard of fastmod, codemod but never used them.
I believe the idea is that those identifiers are semantically related: that fixture decorator inspects the formal parameter names so that it can pass the appropriate arguments to each test when the tests are run. A sufficiently smart IDE and/or language server would thus know that these identifiers are related, and performing a rename on one instance would thus rename all of the others.
And maybe you were being facetious, but an IDE is an “Integrated Development Environment”.
Edit: Yep. Took all of 60 seconds to find what I’m looking for, as I type this from my phone while sitting in my throne room: https://docs.pytest.org/en/6.2.x/fixture.html
See the “Fixtures can request other fixtures” section, which describes the scenario from TFA.
And this post describes the PyCharm support for refactoring fixtures: https://www.jetbrains.com/guide/pytest/tutorials/visual_pyte...
>rename every instance of a pytest fixture from database -> db
Every instance of a fixture, not every instance of all fixtures of the same name.
https://www.jetbrains.com/guide/pytest/tutorials/visual_pyte...