Emacs: Feature/tree-sitter merged into master
lists.gnu.org
lists.gnu.org
This talk: https://www.thestrangeloop.com/2018/tree-sitter---a-new-pars... by the author is quite instructive as well.
That doesn't really bother me. The line of code you're currently typing is generally so invalid that there's not much point trying to color it.
But what does bother me is the code below where you're typing changing its color as the IDE tries to make your partial line fit.
Slow parsing is why emacs used regexps for syntax highlighting instead of parsers in the first place.
Granted, maybe the projects weren’t using tree-sitter correctly. But regex parsing is surprisingly practical, so despite being really amazing at its goals, tree-sitter may not have a definitive advantage
The new implementation has been authored by Yuan Fu in close collaboration with the core Emacs maintainers and the rest of the community. It has been an ongoing effort for many, many months.
This is great news, and means that also core Emacs language-binding provided as part of Emacs itself will now be able to make use of tree-sitter based parsers as well, something which wouldn't have been happening if they would have to depend on a third-party package to get those bindings.
I've been somewhat involved in the process, although not a major player, but needless to say I'm very excited about these news and can't wait to see what sort of improvements this enables across the line once people start using it.
It also includes some new language-modes which has never been part of Emacs before (like typescript).
I'd love to see C# on the list, but that might depend on me having the time to land production-grade major-mode, so that might end up happening later rather than sooner.
Anyway, from what I understand what has been merged so far should all be available as part of Emacs 29 once released.
https://www.masteringemacs.org/article/tree-sitter-complicat...
Thanks, your articles and your book are the best guides into the world of Emacs.
Do you have any tips or guides for using treesitter for syntax highlighting/structural editing and eglot/LSP-mode for everything else?
If you don't have tree-sitter your syntax highlighting will be done by the regex based font-lock-mode. I don't think eglot/lsp-mode make that slower, and I believe tree-sitter should speed it up (and make it more correct) without affecting them. I haven't tried it yet, though.
I don't think I'd want that. Syntax highlighting and indentation are things I want instant feedback from.
That affects the answer to the question. I assume you'd need to persuade lsp-mode not to do this and leave it to tree-sitter, but I don't know how to do that.
https://emacs-lsp.github.io/lsp-mode/page/settings/semantic-...
What lang do you use?
As a user who hadn't kept up with development news until recently, I'd always mentally sorted Emacs into the same taxonomy as stuff like `find`: old, powerful, with a clunky interface and a stodgy resistance to updating how it does things (though not without reason).
I'm increasingly feeling like that's an unfair classification on my part--I'm genuinely super excited to see where Emacs is in 5 years.
Both neovim and Emacs are being improved at breakneck pace, and it is quite incredible for such an old piece of software with, dare I say, a quirky contribution model. The maintainers are working really hard on keeping it current and competitive.
I've been using Emacs primarily for org-mode/roam/babel for a few years now. I'm very glad for its existence, I really think I've become a more effective DevOps person because of it.
Instead it’s backwards, now we have hard-to-use concurrency primitives and still shitty UIs.
My current experience with Emacs concurrency is mostly negative - occasionally, an async-heavy package (like e.g. Magit-style UI for Docker) will break, and I find it hard to figure out why. Futures-heavy code I've seen tends to keep critical data local (lexically let-bound), which is the opposite of what you want in a malleable system like Emacs. For example, I'd like to have a way to list all unresolved futures everywhere in Emacs, the way I can with e.g. external processes. But it seems that at least the async library used (aio, IIRC) is not designed for that.
I think you could get this done by advising promise creation/resolution functions, aio-promise and aio-resolve. The async/await macros are wrappers around generators-over-promises in this library.
But yes, in general Emacs concurrency sucks. The least bad option I found was using promises' implementation (chuntaro/emacs-promise) that uses `cl-defgeneric` for `promise-then` and (obviously) moving as much processing to a subprocess as possible. The former allows you to make any type "thenable" by implementing the method for it, which is nice for bundling the state around async operations. cl-defstructs are nice for the purpose.
Right now we can run subprocesses without blocking anything with "make-process", but interacting with the process is pretty clunky, and you have to use the process sentinel and filter to perform callbacks when the process changes state or exits. There is quite a lot of boilerplate to setup for all of this and the control flow is pretty confusing IMO.
A nice "async/await" style interface to these things could really go a long way I think!
There is one more, possibly gigantic, thing though: Better handling of very long strings. I know the data structures for strings have various tradeoffs, but properly abstracted, it should be possible to even give a choice, no? So users could choose the data structure, based on their use cases. But I know little about the internals and maybe that is all too low level to be something a user could choose from the user interface or configuration.
I hope string data structure is properly abstracted from, so that it is exchangable for another data structure, but I have my doubts. Would like to be surprised here and anyone credibly telling me, that string data structure in Emacs has an abstraction barrier around it, and is actually exchangable, by implementing basic string functions like "get nth character" or "get substring" in terms of another data structure.
If it is not properly abstracted from, then of course it could be a nightmare to change the data structure.
> Emacs is now capable of editing files with very long lines.
> The display of long lines has been optimized, and Emacs should no
> longer choke when a buffer on display contains long lines.
> ...
If all languages changed their reference parsers to tree-sitter, this would be moot, but that seems unlikely. Language parsers are often optimized beyond what is possible in a general purpose parser generator like tree-sitter and/or have ambiguities that cannot be resolved with the tree-sitter dsl.
What feels perhaps likely in the future is that a standard parse tree api emerges, analogous to lsp, and then language parsers could emit trees traversable by this api. Maybe it's just the tree-sitter c api with an alternate front end? Hard to say, but I suspect either something better than (but likely at least partially inspired by) tree-sitter will emerge or we will get stuck in a local minimum with tooling based on slightly incorrect language parsers.
Doubtful, last time I tried tree-sitter would parse invalid inputs without even tagging any errors in the parse tree. For example, it would silently accept extra tokens, or keywords in the place of identifiers. Replacing the built-in lexer and then validating the parse tree for correctness would be close to writing the grammar twice.
And accepting partially correct inputs within the compiler toolchain isn't too hard, so I don't really see the advantage of agreeing on tree-sitter and not just on a parse tree representation that editors can then query, as you then suggested. If the big deal is having it execute client-side or being sandboxed, I feel that's orthogonal to parsing algorithms.
Author of DiffLens (https://marketplace.visualstudio.com/items?itemName=DiffLens...) here. A uniform API for traversing a parse tree for all languages would be amazing for DiffLens! However, I fear languages are different enough that this ideal may never be reached :) Or maybe there would be a core set of APIs and extensions for the idiosyncrasies of each language. For DiffLens though, we try to use the language's official parser/compiler if it exposes an AST
https://github.com/wilfred/difftastic
https://github.com/afnanenayet/diffsitter
(I have not tried either personally.)
Emacs modes can use those parsers on buffer contents, e.g. for syntax colouring/highlighting, finding matching delimiters (e.g. moving the cursor over an `if`, and having all the corresponding clauses (e.g. else/elif/fi) highlighted), for contextual editing (e.g. escaping " when inside a string), etc.
This can be remarkably tricky to get right; e.g. consider languages which can splice expressions inside strings (which can themselves contain strings, containing spliced expressions, etc.)
Using tree-sitter should make this easier and more robust (i.e. less time spent implementing parsers; more time spent implementing features!). I think it would also allow grammars to be re-used across different tools, which should improve support for obscure/niche languages.
Think highlighting for html/JS/CSS in a single file or fully featured highlighting inside markdown code snippets.
https://www.masteringemacs.org/article/tree-sitter-complicat...
In general it's for functionality that needs to understand syntax but doesn't need a full compilation-level understanding of the code, that can benefit from much faster responses than an LSP server can provide and that people may want working out of the box for many languages without having to install and configure language servers, generate compilation DBs, etc.
All the basics of major modes that up until now have been implemented using bespoke algorithms, rewritten ad-hoc for each language.
That said, tree-sitter should make it possible to create paredit-like implementations for languages not LISP and other stuff like that, which IMO could turn out to be really neat.
As a change, this is quite significant, but not directly aimed at end-users.
lsp can do some of these tings as well, but sending the entire buffer over to the lsp server every time you want to update the buffer is an expensive operation. tree-sitter does it locally.
tree-sitter is an Emacs binding for Tree-sitter, an incremental parsing system.
It aims to be the foundation for a new breed of Emacs packages that understand code structurally.
And anyway, emacs is hardly a "text editor". It's a text manipulation engine, likely the only one.
There's an `async` call for system functions in elisp.