Atom understands your code better than ever before
blog.github.com
blog.github.com
Admittedly, Tree-sitter has the opportunity to centralize some of the work in incremental parsing that LSP pushes to individual language tooling, but arguably in the modern world of watch-based programming most language tooling is going to feel a need to build it regardless.
The maintenance headache, though, is that Tree-sitter has its own custom grammar DSL, which likely won't look anything like or share anything with the native grammars of the languages it is parsing, and which ultimately has to settle for a lowest-common denominator approach to grammar parsing. Even if it is more sophisticated than so-called "crude" RegEx-based syntax highlighters, it can still at best be a separately maintained emulation of a language's grammar versus the LSP encouraging a language to provide its own intelligence presumably from the very same parser it uses to get work done.
[1] https://docs.microsoft.com/en-us/visualstudio/extensibility/...
LSP is probably the best way to provide the fixed set of classic IDE features that LSP does - inline diagnostics, autocomplete, go-to definition, etc.
We think that for many other features, Tree-sitter is a much cleaner solution than pushing the logic out into the language server. For example, if you want accurate syntax highlighting that's lightweight and updates immediately (as opposed to in a delayed fashion like in most IDEs), you need incremental parsing. Incremental parsing is a highly specialized problem, so writing an incremental parser in the form of a Tree-sitter grammar is much simpler than modifying a language's compiler toolchain in order to make it incremental, and bundling that whole toolchain into an app like Atom.
The same goes for other features besides syntax highlighting. The syntax tree is now available via a uniform, in-process API in Atom, and is always up-to-date, so you can script the editor to manipulate your code intelligently. I'm not totally sure what kinds of things we'll end up building on top of that, but I think it's a different set of things than LSP will end up providing.
- VS Code implements syntax-aware code folding using the LSP. This means that folding is super flexible but it is also makes computing folds fairly expensive. And every time the document changes, the language server has to re-compute and update the folds. In almost all cases, language servers is just generating folds based on the document's syntax anyways.
Tree-sitter is interesting because it lets folding and other syntax based language features be calculated accurately and quickly on the client, freeing up the language server to do more interesting things (or just go to sleep for a moment). Document outlines are similar; VS Code uses the LSP for this but in many cases the same syntax derived outline could be generated by tree-sitter.
- The LSP probably isn't well suited to general syntax highlighting due to computation cost and communication chattiness concerns, but the LSP may eventually support semantic syntax highlighting [1]. This could, for example, allow an editor to color all singletons hotpink, which requires a semantic understanding of the code. Semantic highlighting would augment the base highlighting provided by tree-sitter or by a TextMate grammar.
I'm the developer of VS Code's JavaScript/TypeScript and Markdown support, and am interested in tree-sitter if only in the hope the it will free us from TextMate grammars. If you want to see just how far regular expressions can be pushed, just go browsing through some of these bad boys; TypeScript's is a classic [2].
Keep up the great work Max!
[1]: https://github.com/Microsoft/language-server-protocol/issues...
[2]: https://github.com/Microsoft/TypeScript-TmLanguage/blob/16c5...
Developing a language server is however a more complicated task - even if there would be a template that does most of the boilerplate. Therefore I would continue to welcome in-process highlighting mechanisms like textmate grammars and tree-sitter as a baseline. They could be augment by LSP features whenever a plugin author feels it's necessary.
Regarding tree-sitter itself: I've read the documentation, and it looks super-interesting. I've developed textmate grammars before and wasn't really satisfied with them, tree-sitter looks like it can provide a lot better results.
I would love to see those getting into VsCode and other editors too. Then good grammars can again be shared between editors! Maybe getting it into VsCode is now easier, since both are now somehow Microsoft editors? :)
I'm a layperson, so please forgive the noob question:
I've written some LL(k) grammars (using ANTLR). I really wanted to have incremental parsing, for a two-way structured editor I envisioned for my UI DSL. But I could never wrap my head around either LR or Wagner et al's paper, much less how Wagner's algo could apply to LL grammars.
Is there hope for us ANTLR people to get something like tree-sitter? (I ask tparr every few years, to see if someone smarter than me has tried. No joy. Yet.)
My hope is to make parser development with Tree-sitter easy enough that you wouldn't have to pave a new path for your structured editor (unless you wanted to) - you could just create a Tree-sitter grammar for your DSL and implement your custom editing logic as an Atom package.
I have tried to make Tree-sitter more approachable than, say, Bison in a several ways. Usually, once a language's basic structure is in place, the process of adding a new feature is pretty declarative and easy. But it's still true that in the course of developing a grammar from scratch, some understanding of LR is important.
To be fair, incremental compilation is way simpler than making the support of incremental X where X is part of the toolchain. Also, although the Tree-sitter approach taken by the Rust plugin for IntelliJ seems like "the right way forward", I think in the long run having the compiler do that for you is a way-way easier-to-manage approach.
Hands down the best code completion and awareness of any plugin on any language I’ve seen, ever. It would, without fail, speedily and correctly pull up the actual methods attached to a Pandas DataFrame, along with their documentation. Worked equally well with any class I wrote, or any library I imported.
Semantic highlighting is awesome, and I miss it in every language I don’t have it in. Legitimately worth every dollar I spent on the licence.
I'd like to eventually use VSCode, but so far they just don't have the same refactoring capabilities.
You can manually set the language pretty easily though via the command palette or by clicking the language in the bottom-right.