Question: do you have any plans to integrate tree-sitter with the language-server project(s)? If fast, accurate parsing of any programming language is now easily implemented in language servers via tree-sitter, it seems to make sense for LSP to expand its protocol to include syntax highlighting as well.
I haven't specifically pursued integration with LSP, but Tree-sitter has been used to build a couple of language servers which work with both VSCode and Atom (and probably other editors):
* Bash - https://github.com/mads-hartmann/bash-language-server * Ruby - https://github.com/rubyide/vscode-ruby/tree/master/server
Yes! The state of parsing is such a sad thing, discouraging so much other progress, that it's wonderful to have it improving. It's appreciated.
EDIT: And my thanks for all your guidance this evening. Good night.
I noticed [1] doesn't have a version field. "No version: means version:1" is ok, but I then wondered.
[1] https://github.com/tree-sitter/tree-sitter-json/blob/master/...
At one time, I had the thought that folks who disliked JavaScript might want to write their own grammar DSL in another language, which could generate compatible JSON to be consumed by the core Tree-sitter compiler library. Now, I think this flexibility is pretty unnecessary.
I've also thought that some applications might want to inspect the grammar JSON at runtime in order to do some kind of meta-programming. I haven't really thought that through though.
In either case, versioning would probably be a good idea. The actual parser is already versioned so that we can error out if you try to use a parser with an incompatible version of the runtime. We could just add that same version number into the grammar.
One possible approach might be "part of spec" but "optional" and "not actively it use"? So present in schema and tests, to somewhat reduce the chance of getting locked into not having it by hypothetical others failing to support it. And easily available, so absence doesn't discourage any hypothetical interest. But without spending much time on it, absent non-hypothetical need. Maybe.
> that same version number
Hmm. Parser-runtime versioning would seem only lightly coupled with grammar-parser versioning? Unless there's a grammar-runtime coupling to express?
EDIT: Ok, I was thinking of "grammar" in too general a sense, decoupled from choice of parser tech, rather than the usual "bison grammar tightly coupled to bison parser". Failure to RTFM. Grammar's `conflict:`, `inline:`, `word:`, `externals:` are all somewhat runtime coupled. Though perhaps looser than parser-runtime? It can be nice to flexibly move work between a compiler and its runtime, without affecting the external api.
Is usually what I end up doing. But there's already the npm package versioning to capture api versions. So one question is whether grammar json is likely to be moved around in space or time in ways for which that's insufficient. One story could be moving generated grammars among different programming language tree-sitting implementations, which can be out of sync, and where there's no out-of-band information on how to deal with it. Another story could be grammar json being captured in repo, as with tree-sitter-json, while avoiding "be careful updating the version of tree-sitter you use". Generalizing that, any other cases where you don't want to generate the json at runtime, and thus it gets stored longer-term.
Can `externals` call back into the parser? That is, is the parser reentrant?
Can `externals` be a zero-length assertions? That is, is there a progress requirement for success? And thus serve as arbitrary-computation parse-rejecting "semantic actions".
Do `externals` have access to GLR parse state?
You cannot call back into the parser; it’s just for tokenization.
You can produce zero length tokens, and we do use this feature a lot.
You dont have access to the parse state directly, but the external scanner is passed an array of Boolean values that indicates which external tokens are expected in the current state.
Thanks again.