Show HN: Python Source Code Refactoring Toolkit via AST
github.com
github.com
Any suggestion on an AST ‘framework’ that would help me parse code easier? Language specific or generic, even if it only sort-of fits (I don’t even know what I want).
Automated Code analysers exist, but I want something manual.
I found it last week. It may be able to detect a potential issue in multiple languages, but not sure if it supports refactoring.
It's mostly a pet project and I just wanted to share since it could maybe at least inspire something.
I've been using tree-sitter for syntax highlighting, simple refactors, and custom text objects in NeoVim recently (and love it, thanks NeoVim devs!). Have a play with the syntax tree it generates https://tree-sitter.github.io/tree-sitter/playground (it can generate a tree for most languages, not just the ones in this playground).
It still needs a lot of handholding, but at the same time I get to use gems like -
... = TemplateTranslator({
'isinstance($a, str)': "typeof($a) == 'string'",
'isinstance($a, bool)': "typeof($a) == 'boolean'",
'len($a)': '$a.length',
'str($a)': '$a.toString()',
'getattr($a, $b)': '$a and $a[$b]',
'getattr($a, $b, $c)': '($a and $a[$b]) or $c',
'hasattr($a, $b)': '$b in $a',
'$a.join($b)': '$b.join($a)',
... })The main concern with using the ast module from stdlib is that you lose all formatting/whitespace/comments in your syntax tree, so writing any changes back to the original sources requires doing a lot more extra work to preserve the original formatting, put back comments, etc. This is entire point of lib2to3/fissix and LibCST, allowing large scale CST manipulation while preserving all of those comments and formatting. We do recognize the limitations of lib2to3/fissix, though, so there have been some backburner plans to move Bowler onto LibCST, as well as building a PEG-based parser for LibCST specifically to enable support for 3.10 and future syntax/grammar changes. But of course, this is very difficult to give any ETA or target for release.
The real start point for this project was to find / replace all type(<literal>)'s in CPython codebase with type(type(<literal>)) (e.g type('') would become type(str)) which is very light weight transformation, and I was able to write a script which did it without having any major problems about style on over 2000 files. Here it is for the reference: https://github.com/isidentical/refactor/blob/master/examples...
Also one thing to note here is that; in the last couple of years, thanks to black (and yapf), the adoptance of code formatters have really increased which is very nice for custom refactoring tools like refactor since the end-code would be refactored anyways so that means if you convert a multi line call, or a list to a single line version then the formatter you use probably reformat that segment anyways.
But thanks for authoring Bowler! It is a very cool project.
Things like "find any function that uses kwargs" or "find any class that defines a method named `bar`" can be easily expressed with lib2to3's matching grammar, and no other CST that I'm aware of (that isn't itself based on lib2to3) has equivalent functionality. This is something we wanted to add to LibCST, but haven't had the time to focus on given other priorities. Meanwhile, we used LibCST to write a safer alternative to isort: https://usort.readthedocs.io
From https://news.ycombinator.com/item?id=24511280 :
> Additional lists of static analysis, dynamic analysis, SAST, DAST, and other source code analysis tools […]
For my use-cases, not much changes between the two - 1. Replace opening and closing braces with colon. 2. Remove semi colons 3. Replace function with def 4. Remove type definitions 5. Replace array manipulation functions with python equivalents. 6. Remove tokens like const, let and var
For my use-cases, not much changes between the two -
1. Replace opening and closing braces with colon.
2. Remove semi colons
3. Replace function with def
4. Remove type definitions
5. Replace array manipulation functions with python equivalents.
6. Remove tokens like const, let and var
However, I think it will be much easier than translating the other way around :)
CPython is so unbelievably slow when compared to languages like PHP and Javascript.
When using PyPy, I keep encountering road bumps because many tools are still expecting that one uses CPython.
Dear world - please accelerate the conversion to PyPy!
In contrast, a tracing JIT can dynamically elide most of this work without changing the semantics.
I wish their readme gave some kind of estimate of how the speedups might compare to Cython and PyPy. But it's currently alpha, so there may be big differences by the time it's ready for production use.
Similar approach to the Sorbet Compiler for Ruby (albeit targeting C instead of llvm... I wonder how that might impact optimizability).
https://cffi.readthedocs.io/en/latest/
I do agree that a standardized Python should have an object model that is a clean public contract, but given that Python is CPython and there is no standardization, that the library runtime is married to the core also makes this incredibly difficult politically.