Topiary: A code formatting engine leveraging Tree-sitter
tweag.io
tweag.io
It's fascinating seeing these tools that facilitate building better programming language experiences. I've called them "Tooling for Tooling", basically tools that make it easier to create tools like formatters, linters, etc.
We share the same curiosity for the effectiveness of our approach! Right now, we want to make Topiary great for languages with less complex formatting rules, and "good enough" for languages that are a bit more complex. Where, on a one-off basis, you don't feel the need to get the dedicated formatter.
We don't yet have any ambitions to compete directly with Prettier and rustfmt among others.
Having said that, we are quite proud of how the OCaml rules turned out, and even had some great results with the Rust rules.
As we explore more, and expand the complexity of our tree-sitter scopes, who knows what kind of things we might be able to format!
It's all very exciting!
I keep wondering, there a reason everything isn’t just based off treesitter these days? If I were tasked with writing tsserver, my instinct would be to layer it over treesitter. Does anyone know if there are practical reasons that doesn’t happen, or is it just legacy?
The neovim world is slowly converging on using it for syntax. This project for formatting. I personally use it in neovim for things like highlighting and formatting sql within strings in python.
I can only see more editors just giving up on this and going all in of treesitter.
It's not only neovim, btw, the next emacs is also gonna have it included.
There SO MANY perfect tools for language implementation and ast manipulation..! Starting with ML family languages that are very good at this.
People have been looking for universal approach to parsing for so long... maybe there is one, maybe not, but treesitter was never meant to be one.
And it's great for what it does!
If you want to create a parser for a toy language that produces an AST or a single error, then sure, that's trivial. But if you want a parser that does good error recovery, produces a high fidelity CST, and reuses memory in an efficient manner (red-green trees ideally), that's a lot of work. And that's table stakes for good programming language tooling. We're not in the era of emacs plugins that do regex syntax highlighting and call it a day. If there was a framework that could accomplish this, and function as a parser for the compiler (which is not so crazy, since most modern compilers are also the engines for tooling, i.e. language servers)
I agree that tree-sitter was never meant to be a universal solution, but I think it's easily could be with some adjustments. And because of the existing infrastructure, because of the existing parsers, I think that it's reasonable to consider pushing tree-sitter in that direction instead of creating yet another parsing framework.
As somebody who came up with a couple of quick modes and parsers for Emacs and in Emacs Lisp I can say that for people like myself it's a blessing. I sincerely hate how there are numerous implementations of everything in dozens of editors out there, but nobody benefits from each others work in a reasonable way... Treesitter's universal community-centric approach kind of resonates with the stronger side of OSS: suddenly all of these little steps individuals do contribute to the ecosystem as a whole.
Now, admittedly, all I need is an axe. I know I need an axe, treesit gives it to me and this makes me a happy little contributor.
So let's say somebody comes up with a factory of a tool. All inside: properly incremental, smart error handling, tree editing, transformations and stuff. Something tells me it would much harder to contribute a simplified barely working grammar for that thing. And this kind of kills the point of emacsy-sh moonlight hacker tool.
It would be useful, sure, but would it work in practise?
Where I disagree is that IMO, tree-sitter already is very close to this ideal model. It has incremental parsing. It has great tree querying. Where it needs help is an AST facade over the raw syntax tree, which is very much feasible. rust-sitter[1] does it for instance. Tree-editing and tree construction is also very much doable. I don't think it'd have an impact on grammar construction at all. As for error recovery, I think it could function as a reparsing feature where you can drop down to a manual parser (or even a secondary grammar) that is more tolerant. Or an error recovery function that can be written in any language. tree-sitter already has the ability to use a manual lexer written in native code, so this is not such a stretch.
My only issue with it is that it really only does half the job. You get a CST of sorts, but if you want to do anything with it you pretty much have to hand write another parser for that node tree.
In contrast parser combinator libraries like Nom and Chumsky give you "the final output".
My research into programming language parsing started with a very specific problem: I like folding code, and I like disabling ("commenting out") code to test behavior. Well, but (with rare exceptions: Xcode and nowadays some languages in VSCode) "commenting out" code breaks folding. I never got around to really solving it, but the learning involved (including about tree sitter) was very cool.
Tree-sitter deals with errors better than most parser generators but if you just lex and separate into chunks then you can much more flexibly format broken code.
aaa(
bbb,
ccc(
ddd,
eee,
),
fff,
)I agree with you in that there are many languages where skipping parsing altogether could still result in a good formatter, and I would love to see a Topiary-like project attempt it.
I don't feel confident in saying that that holds for most languages however, worrying that it can lead to a lot of ambiguity in languages with more complex formatting conventions.
Regardless, the eventual goal of Topiary is to be able to format the widest possible spectrum of languages, and so limiting ourselves to just lexing didn't seem like the right choice at the time.
Like you mention, this does mean we give up being able to format broken code. In fact, we currently even ensure that TS is able to parse the entire input before formatting. This is a shame, but ultimately what we decided was the best approach for Topiary to achieve its goal.
We are not sure right now because Topiary is still very much an experiment.
Having said that, we are constantly surprised what we can do with Topiary. So with a dedicated Python developer willing to draft a set of rules, it might just be possible!