Speeding up VSCode extensions in 2022
jason-williams.co.uk
jason-williams.co.uk
I do this with an extension I develop, and it works very well. Microsoft even provides a library to make this easy, and sample code[2]. In my case the language server is written in Rust, and uses tree-sitter. This combination feels like a super power. (The first version of my extension was written in clojurescript, using the instaparse parser. It quickly became apparent that it was way too slow.)
A note on my experience with tree-sitter: It's awesome in every way except one. It's so fast that so far I haven't even bothered with incremental parsing: I can do full parsing on every keystroke; and I know that if I ever need a speed boost, I can add incremental parsing. The API is sane and easy to use. The query API is powerful. But the main weakness of tree-sitter is that the error messages are nearly content-free. The most common error looks like this: "(error)". In my case I can deal with that, but I can imagine that for many purposes that's not sufficient. There's been an issue open for 3 years about improving the error messages[3] but I haven't seen a ton of progress (unless I missed it?).
1. https://microsoft.github.io/language-server-protocol/specifi... 2. https://github.com/microsoft/vscode-extension-samples/tree/m... 3. https://github.com/tree-sitter/tree-sitter/issues/255
One more: the Rube-Goldberg build setup. You need rust, npm, node and a C compiler as well as Docker or Emscripten. Tree-sitter is mostly written in Rust and there are Rust bindings, but you can't use them to for the wasm compilation target because it requires linking a C generated grammar which breaks things. I hope they'll move on to something that's possible to build and use with pure Rust.
> If you really need to do some CPU intensive work, it’s now possible to offload some of the workload to a language server. This allows you to implement the bulk of your extension in another language (for instance, writing Rust code and compiling it down to WASM).
The section above also mentions tree-sitter.
The WASM bindings also perform very well (better in some cases), so it can be used anywhere WASM can. Which, it seems to me, means it’s a very good candidate to replace Babel without depending on language-specific tooling like SWC/Rome/Bun.
I think this is a much better approach that baking tree sitter into VS Code and continuing to use TextMate grammars (or tree sitter specific grammars) to add syntax highlighting. That way it can be applied to any editor that supports LSP integration. Really the world would be a better place for IDEs and text editors if we could just get LSP more standardized, and VS Code's dominance is a decent place to start with it. I'm tired of relying on editor and IDE authors for language support.
It's not really addressed. The semantic tokens API is intended for semantic highlighting:
> Semantic tokenization allows language servers to provide additional token information based on the language server's knowledge on how to resolve symbols in the context of a project.
Abusing the semantic tokens API for syntactic highlighting is slow, unnecessarily complex (why do I need to implement a language server just to do syntactic highlighting?), and only a partial solution (still need a TM grammar, still don't get correct code folding, etc.).
Am I right in thinking that the Semantic Tokens part of the spec (currently) falls short of saying "this is for all your syntax-highlighting needs, please go ahead and implement full-blown syntax highlighters using it"
I definitely agree with you that the more we can share between different editors/IDEs the better.
they're implementing both, with tree sitter being 'dumb' version of LSP syntax highlighting: https://github.com/microsoft/vscode-anycode
Can someone explain the paragraph below -- I thought "this is an invocation of a function named bar" was what we mean by semantic information:
> All features are based on parse trees and there is no semantic information - that means there is no guarantee for correctness. Parse trees allow to identify declarations and usages, like "these lines define a function named foo" or "this is an invocation of a function named bar"
- they’re importing query files, with a .scm extension, in TypeScript
- they’re using esbuild to handle those imports
It’s surprising to me that the Microsoft, who created and maintains TypeScript, is using an alternative TypeScript compiler… in part to query semantic information from TypeScript.
>The fact that we now have these complex grammars that end up producing beautiful tokens is more of a testament to the amazing computing power available to us than to the design of the [TextMate] grammar semantics.
Amazing computing power available to them indeed.
- The "main process" which manages the windows (renderer processes)
- The renderer process" contains the UI thread for each window, the renderer process can have its own worker threads
- The extension host loads extensions in proc, extensions are free to create their own threads/processes. The separate process for extensions protects extensions from freezing the renderer
- Various other processes that live off either the main process or the "shared process", such as the pty host for terminals which enables the terminal reconnection feature when reloading a window (also file watcher, search process)
We've been shuffling where processes are launched from recently but the actual processes and their purpose probably won't change. You can view a process tree via Help > Open Process Explorer to help understand this better.
EDIT: Formatting
VS Code is not single-threaded in general. It's event-loop is.