Let's write a treesitter major mode for Emacs
masteringemacs.org
masteringemacs.org
While Emacs 29.1 comes with "treesitter" built-in, you still need to manually build and install any treesitter language plugin implementing the actual language specific parser. This can be fiddly and frustrating doing it yourself.
I had a quick success with using this convenience script: https://github.com/casouri/tree-sitter-module/. It provides fully-automated builds for the most popular languages (including typescript, c and c++).
This is how it works for "typescript":
1. Clone the repository: https://github.com/casouri/tree-sitter-module/
2. Install "build-essentials" (providing a c/c++ compiler if you're on Linux).
3. run "./build typescript" from within the repo
4. Copy the resulting shared library from "dist/libtree-sitter-typescript.so" into your "~/.emacs.d/tree-sitter/".
5. Open a random typescript file and try "M-x typescript-ts-mode" which should not give you any error but instead nice syntax highlighting.
You might find there is a treesitter plugin for your language available and it is even supported by "tree-sitter-module" but there is still no major mode, yet. Happened to me for Perl 5.
https://www.masteringemacs.org/article/how-to-get-started-tr...
What I read from this is that tree-sitter isn't considered quite ready by the Emacs maintainers, perhaps because of the restricted number of actual treesitter modes, or maybe because the treesitter support itself is not quite considered there yet?
sudo snap install emacs
It works w/o problems as it is using snapcraft's "classic" runtime. It comes with both native compilation and tresitter support.Edit: I can confirm it works nicely for "javascript". Cool!
(setq treesit-language-source-alist
'((typescript "https://github.com/tree-sitter/tree-sitter-typescript" "master" "typescript/src")
(tsx "https://github.com/tree-sitter/tree-sitter-typescript" "master" "tsx/src")))
(mapc #'treesit-install-language-grammar (mapcar #'car treesit-language-source-alist))https://www.masteringemacs.org/article/how-to-get-started-tr...
Part of this is surely that I don't know wtf I'm doing, but it seemed like there was not an underlying data structure held in memory that you could conveniently query / manipulate, but rather, most of the existing org functionality built some kind of structure each time you did an operation.
Would appreciate any pointers, code examples, tutorials that show how to effectively navigate / manipulate an org structure and have it reflected in the buffer, if there is such a thing.
But even with this I found it pretty awful.
https://github.com/spegoraro/org-alert/blob/master/org-alert...
I am very far from being knowledgeable about programming on the Emacs platform, but I am trying to learn. I grabbed the name M-x-AI.com a while back with the goal of integrating other people’s Emacs packages with some of my own hacks into a better AI dev work environment and writing a short book on it. I have been using Emacs since, I think, 1982. There are so many good new packages for integrating CoPilot, GPT-4, etc., as well as major Emacs platform improvements that are too many to list.
If so, which modes and packages do you use?
https://github.com/emacs-jupyter/jupyter#org-mode-source-blo...
I guess the reality just isn't so rosy.
Even if all it was supposed to do is syntax highlighting, different languages are usually highlighted differently and presumably the person maintianing the bindings would need to be a language expert with strong opinions. The options are they either maintain their own mode or a module in `treesitter-mode` which is basically 2 ways of describing the same situation.
Although to address your other point it is quite funny that after 40 years and 29 versions of experimenting in OS design, GNU Emacs is starting to implement syntax highlighting properly for its system text editor.
We're in a transition period while everything is rewritten to use tree-sitter. In a few years all the default major modes in Emacs are probably going to be tree-sitter based - unless the maintainers believe them to be better than tree-sitter.
Not really. You setup hooks, usually based by fileextension, and emacs execute the hooks. Emacs itself has no understanding of languages. And this will not work well when there is no hook executed.
> Modes are mixable.
You can load multiple minor-modes, but there is always just one major-mode per buffer, and languages are usually major-modes. But ok, treesitter can be handled as a minor. But it still needs explicit support. You cannot give treesitter control over a range of text, let if figure out the language automatically and let it make it's thing. This needs explicit support.
So, the line of the joke goes, it must not really be a text editor because it is not very good at it and has a wild array of other capabilities. Like "M-x dunnet". Must be an OS.
Of course the reality is that Emacs just copies good ideas from other people. A pretty normal Emacs setup uses Vim keybindings, LSP and now tree-sitter to get a good experience programming.
Too bad other people don't copy good ideas from Emacs. Why other text editors/processors have no equivalent of view-lossage is a mystery to me. Also, where-is, describe-key, insert-char, transpose-chars and lots of other very useful stuff...
Currently language servers support semantic highlighting which is close, but as far as I can tell, it seems meant to be supplemental to in-editor local highlighting rather than a full replacement.
(The code being valid s-sexps doesn’t even guarantee it’s syntactically valid after macro expansion, but even if it does this is a bad idea.)
Though handling opening and closing a string and still ending up in an illegal state would need some more logic.. Maybe not worth it.
If you continue down this line of thinking a few more steps, eventually you produce tree-sitter.
Sounds like a feature - don't need to scan for tiny squiggly red lines, or examine some tiny font output in a status bar somewhere.
I'd love it.
Your comment implies a language server taps into the language's compiler/interpreter. That is a popular misconception. Almost all LSP severs don't actually use the compiler backend of the language they are servicing. All the LSP server implementations I've seen at least just use a static parsing approach (or even worse i.e. just a tokenizer) similar to what Treesitter does. It is just not limited to a single file.
That's why I still think Treesitter has the potential to not only improve syntax highlighting but also to simplify language servers or even -- in the long run with some extensions of the the language modes -- replace it altogether.
But it is a misconception which framed the discussion, not a misconception of the answerer. The claim was that LSP was
> performed by 100% accurate parsers instead of approximations
A good overview here: https://matklad.github.io/2022/04/25/why-lsp.html
As somebody who wrote a tree-sitter grammar recently I can confirm that the library definitely has its share of... Interesting choices. But there's nothing else like it as most parser generators are non-incremental, don't do generalized parsing and don't provide decent error recovery.
Curious to hear more.
1. NodeJS-based utilities. The tooling around js doesn't lean itself well to integration with with the dev's environment.
2. The sdk itself is a mix of Rust, javascript and C. I'd say that this is at least one language too many - makes it harder to contribute. Ideally such fundamental projects should be as simple and homogeneous as possible.
3. I keep wondering if generating languages other than C should be supported.
The core idea of the library is brilliant either way so it sort of compensates all of the above.
> Fourth, Microsoft itself doesn’t try to take advantage of M + N. There’s no universal LSP implementation in VS Code. Instead, each language is required to have a dedicated plugin with physically independent implementations of LSP.
There's no (official) LSP implementation for Typescript either. Instead of using LSP, Microsoft maintains tsserver which uses a custom protocol for better integration.
Take-home message: don't try to write a universal tool that solves everything, as it will be lowest common denominator. Create building blocks such as Tree-sitter, which make specific (M×N) integrations easier and more powerful.
Also LSP was never going to completely eliminate the need for language-specific plugins. But it does significantly reduce the amount of work needed to build each one, providing a consistent baseline.
I didn't know tree-sitter-highlight exists though. It seems to provide M+N highlighting and includes a CLI tool. https://tree-sitter.github.io/tree-sitter/syntax-highlightin...
Similarly, tree-sitter tags seems to provide M+N code indexing: https://tree-sitter.github.io/tree-sitter/code-navigation-sy...
In editors this always seems extremely esoteric comparatively: I've tried doing it in a few.
I'm sure brilliant people find it easy, but I'm merely average on a good day.
I haven't tried extending any of these modern electron based editors, can anyone speak to that?
Now it's much, much easier providing there's a TreeSitter parser for your language.
I don't know of anything else that bridges the gap like this.
https://www.masteringemacs.org/article/tree-sitter-complicat...
And also why using LSP to furnish your editor with highlight markers is an inelegant solution for many languages.
But fast enough re-parsing of fragments and recovery from errors is a much more complex problem, that often doesn't have a single correct answer, and it's also a much newer problem in as much as syntax-highlighting is much newer feature, being preceded largely by "offline" pretty-printers with very different constraints.
The extent to which modern compilers try to parse past errors still varies greatly, with a whole lot not even trying to.
But just any recovering parser also does not mean the problem is solved. E.g. you've typed "foo". Now you type "(". It'd be very annoying if your editor now re-colors everything as an error, so you typically want some error recovery. But how soon? Do you assume the tokens immediate afterwards are par of what was a valid expression until you typed "foo", or are they a valid part of an argument list? And where do they end? Do you just delay re-parsing until the user has typed more? Or left the line? Sometimes that can help, sometimes it will just make things worse.
Parsing methods that work fine if you assume you can "reset" the parse at many different points which tend to constrain the area considered an error and so reducing the size of a typical re-parse will fail badly if you want stricter re-parsing that frequently may trigger reparsing most of a file, for example.
A lot of this is subjective, and picking the "right" way of handling it largely comes down to unpacking humans unstated preferences, and trying to reconcile competing and possibly contradictory preferences.
https://www.masteringemacs.org/article/combobulate-structure...
There is also evil-textobj-tree-sitter for tree-sitter based text objects for Evil mode:
Lisps essentially have no syntax, that's why it is trivial to manipulate a lisp program structurally.
I wonder, though, if a "semi-structural" editor with a keypress to "complete what is otherwise an error" would work well. E.g. I write "begin" and so imbalances an expression, and the editor looked at what I typed that changed the parse from successful to an error and puts up options to insert "rescue ... end" or just "end". Maybe coupled with a warning on save if the file does not parse.
Just an aside. I wouldn't associate syntax highlighting with new. It's almost as old as text editors. Accurate syntax highlighting in real-time visual text editors is over 40 years old by now and was becoming common by the late 1980s.
GNU Emacs gained syntax highlighting in 1989 and it was considered late to the party. Many programming editors and IDEs had syntax highlighting by then.
https://steve-yegge.blogspot.com/2008/03/js2-mode-new-javasc...
I guess comp. sci. people studying languages have been more interested in syntactically valid programs than the opposite.
I see some people say it's possible and use both together but I thought for the most part language servers offer the same set of features, and probably better? My current mental model for how to use them together is that the majority of the languages I quickly read I set up treesitter for speed. For languages I read extensively or write I set up a language server.
lsp (and lsp-mode) are mostly concerned with IDE functionality- go-to definition, show references, displaying project errors in real time without explicitly building, etc.
tree-sitter builds a syntax tree of your source code; its applications are things like syntax highlighting and structural navigation of your code.
there is some overlap in functionality, lsp has somewhat supported mechanisms for syntax highlighting iirc, but they are fairly orthogonal overall
so yes, it makes sense to use them together
https://microsoft.github.io/language-server-protocol/specifi...
Obviously it depends on the LSP implementation but it looks expressive enough to replace treesitter?
I was tremendously sad to see that the Typescript Language Server wasn't owned by Microsoft <https://microsoft.github.io/language-server-protocol/impleme...>, since if there was any sanity in the world a spec bump would travel with a reference implementation showing how they envision such a thing being used
But, I found that the Typescript Language Server that they did list does indeed have a semantic-tokens module in it, although it's much shorter than I would have expected from reading that section in the spec: https://github.com/typescript-language-server/typescript-lan...
> Transforms the semantic token spans given by the ts-server into lsp compatible spans
It stands to reason: a language server often does way more than just incremental parsing of the source code into a concrete syntax tree. By limiting itself to syntax, tree-sitter can be much faster.
But well, I still love emacs :-)
I have not embraced some syntax highlighting but with a very subtle color scheme. I tried LSP for a moment but the performance was such that it was a net negative.
I allocate a week over the holiday for "tooling refactors" and I'll probably kick the tires of emacs 29. The big refactor is building a new rig.