A source of knowledge for Language Server Protocol implementations
langserver.org
langserver.org
Instead, the standard requires running half a dozen processes with an API that doesn't provide anything except the exact data required for a number of handpicked use-cases.
Considering that even modern IDEs (again, see IntelliJ for examples) can do more, I don't see how this is a good way forward.
I think you may be underestimating the feasibility of what you propose.
This whole LSP thing is a mindbogglingly bad idea, brought to you by the same kinds of thought processes that created the disaster that is today’s WWW.
I guess the reason is that the original use-case was VS Code & TypeScript (?) and developing a lowest common denominator API for all clients would mean C, so they would have had to program to a C API even though both the client (VS Code?) and the server (Node?) are running JavaScript.
But then maybe the answer should be to provide better C invoke wrappers for high-level languages, not to use HTTP instead of C.
How to support intellisense for a language one time and have it work on many editors?
The other side of that is that we're not in the 90s with single core processors and there's been a lot of hardware & software optimization over time. Most people can run something like VSCode and a language server and still have multiple cores left over — and since the events which trigger language server interactions are generally made by a human with a keyboard it's not like you need to be in the microsecond range to keep up.
A PC full of programs built with these assumptions will probably grind to a halt all the time despite having very high spec hardware.
I mean, taking your argument seriously would mean everything should be hand-tuned assembly. Obviously we figured out that other factors like time to write, security, portability, flexibility, etc. matter as well and engineering is all about finding acceptable balances between them. Microsoft has been writing developer tools since the 1970s and in the absence of actual evidence I’m going to assume they made a well-reasoned decision.
While you're right that productivity versus performance is a trade-off, and an editor is not necessarily a high performance application, its not clear to me whether future optimizations would reduce the gap, as much as optimizing compilers did vis-a-vis C and assembly.
In any case, that aside, the core guarantee of software stability with LSP remains to be seen.
I don't follow this conclusion: haven't we already seen it with the way language servers crash and are just restarted without other side effects?
> Do you have profiler data showing that the microseconds needed to pass a message between a process is a significant limiting factor
What? I think you didn't get my point. Let me try again.
You can look at a single operation and say "oh, that's nothing, it's so cheap, it only takes a millisecond". Even though there's a way to do the same thing that takes much less time.
So this kind of measurement gives you a rational to do things the "wrong" way or shall we saw the "slow" way because you deem it insignificant.
Now imagine that everything the computer is built that way.
Layers upon layers of abstractions.
Each layer made thousands of decisions with the same mindset.
The mindset of sacrificing performance because "well it's easier for me this way".
And it's exactly because of this mindset.
Now you have a super computer that's doing busy work all the time. You think every program on your machine would start instantly because the hardware is so advanced, but nothing acts this way. Everything is still slow.
This is not really fear mongering, this is basically the state of software today. _Most_ software runs very slow, without actually doing that much.
The technologies that win are the ones that account for that.
Furthermore, many languages ship with a runtime which does not play nice with other runtimes. Try running a JVM, .NET VM, Golang runtime, and Erlang BEAM VM in the same process and see what happens. Better yet, try debugging it.
You are simulating a condition that will almost never happen unless your machine itself is dying in which case you have bigger problems than auto-completion to worry about.
Sure, could do some hyper complex setup with multiple VMs etc... or you could just run one command to temporarily increase latency on localhost.
If you know some better way of doing this, feel free to tell me.
Setting up a new interface is pretty easy on most unix OS's and then you can increase the latency on just that interface without impacting the interface most software on your machine is built to expect to be blazing fast. And you can mess with that interface to your hearts content knowing the only thing affecting it is the applications you are running on it.
I'd claim that for most languages you could define a formal grammar as an EBNF that would provide some basic utility.
Such a grammar would probably be overly permissive (it doesn't know anything about references, types or host environments) but it would provide you a baseline for syntax checking and autocompletion and could generate an AST that more language-specific rules could evaluate.
In particular, I know that using a BNF is not useful for shell, having ported the POSIX grammar to ANTLR.
(Some programming languages are context-sensitive, some programming languages allow ambiguity in the main grammar and have precedence rules or "tie-breakers" to deal with those situations.)
"Universal grammar engines" and "Universal ASTs" are wondrous dreams had by many academics and like the old using RegEx in the wrong place adage: now you have N * M more problems to solve.
I think the OP is looking for something more like this. This paper is from the same author as Nix, so he has some credibility. But I think this paper is not well written, and I'm not sure about the underlying ideas either. (The output of these pure declarative parsers is more complicated to consume as far as I remember.)
Still, the paper does shows how much more there is to consider than "BNF". Thinking "BNF" will solve the problem is a naive view of languages, and as you point out, is very similar to the problem with not understanding what languages "regexes" can express.
http://eelcovisser.org/post/135/pure-and-declarative-syntax-...
Mainstream parser generators pose restrictions on syntax definitions that follow from their implementation algorithm. They hamper evolution, maintainability, and compositionality of syntax definitions.
The point is to move the use-cases to the server/protocol to avoid the "(m languages) x (n IDEs)" problem. The protocol is at https://github.com/Microsoft/language-server-protocol with contribution guidelines so it should "allow more use-cases in the future".
> Instead, the standard requires running half a dozen processes with an API that doesn't provide anything except the exact data required for a number of handpicked use-cases.
Why half a dozen? Two seem enough (IDE + language server). It's also the minimum given that IDEs are using different languages/runtimes and that lot of languages have most of their tooling written in the targeted language.
If you only work in a single language, yes. However, it's quite common that the same project includes multiple different languages, linked or nested into each other. E.g., a web project might have HTML, CSS and (possibly nested) JS for the frontend (Replace with Coffeescript, React, SASS etc as needed) and PHP, Java, or yet another configuration of JS as the backend - not including build scripts and config files. So if you switch between them reasonably often, you'd have language server processes for each of them running in the background.
Alternatively you could have more coarse-grained servers that handle multiple languages (e.g. HTML+JS+CSS, node.js+npm, java+maven etc) but that would seem to make adoption even harder.
Meaning if your language's compiler and tooling are self hosted, you can write the plug-in in that language and leverage the very tooling used at build/runtime.
In other words, you can choose the right tool for the job. One that already exists and is tested, stable, and full featured. Or, you could rewrite everything in C or something that speaks it, run them all in the same process, meaning one bad plug-in brings the whole thing down.
And if you have to reimplement, that ups the chance that your version will behave subtly different than the real thing.
(This is assuming that reducing the number of processes improves performance; otherwise, a process per language might be simpler.)
This is a text-book case for refactoring.
I think this is assuming too much about how the language server is going to work. For example, does it parse everything greedily, or does it know how to parse stuff on demand? If the AST is in-memory in the editor, that might not work too well with giant projects that need lazy parsing. A language server might also want to do really fancy things like bust its file cache on INotify events.
https://github.com/google/kythe
EDIT: Looking that the graphs, the project is obviously not "dead", but it seems it didn't get any traction outside Google.
https://public.gitsense.com/insight/github?r=google/kythe#b%...
I think a lot of infrastructure would be simplified if each language (or interpreter/compiler) would provide a flag or tool to convert raw syntax into some sort of s-expression (or JSON or XML if you prefer). This wouldn't be a "full AST", but just the "skeleton", enough for us to do tree traversals/transformations instead of having to lex, parse, etc.
It already exists and is called Language Server Protocol.
Is there a specific reason you think it would be slow? Translating between text formats is pretty quick, and there's no need to do anything else that compilers normally do (e.g. building an AST; type-checking; resolving imports; optimising; generating code; etc.)
> At the very least you have to keep the compiler process running and you need a protocol to communicate with the compiler.
That may have a slight speed benefit, but for something as fast as parsing I think it would be over-engineering. Just pipe through stdio.
> It already exists and is called Language Server Protocol.
LSP is certainly interesting, although it's very new and seems rather limited.
I'm mostly curious why so many languages, after choosing (perfectly reasonably) to avoid an s-expression syntax, end up "throwing the baby out with the bathwater" by providing no "backend" syntax for tooling. Some tools represent code as e.g. XML for their own purposes, but I've never seen a tool-agnostic format, or a language "endorse" one of these as an interchange format. Instead, I've seen a huge amount of effort wasted on writing and maintaining a whole bunch of bespoke parsers and pretty-printers.
What it's awesome for is new editors and languages. If editors implement this interface they immediately have decent support for a decent number of languages, and if a language implements this it immediately has basic support in all editors supporting LSP.
Microsoft originally developed this for their Visual Studio Code and Typescript integration, both of which faced the typical challenges new editors and languages have with support.
Isn't most of the machinery to extract that information from source text present in the standard library? https://golang.org/pkg/go/
(I made a list of useful, stably working ones about half a year ago when hacking on a custom VScode plugin, many but not all of them linters, haven't looked into latest developments: https://github.com/metaleap/go-util/blob/master/dev/go/devgo... )
Just that maybe none have fully implemented LSP. I guess even a from-scratch Golang LSP implementation might do well to just utilize these tools for underlying lang intel --- rather than these tools themselves bloating up by implementing some maybe-just-transient-maybe-the-future current-day MS-backed protocol.
Code completion has a much higher bar for acceptable UI, though. I trust a Language Server Client with tons of users/contributors to deliver non-idiosyncratic UI sooner than MVP UI integrations of specific tools.
A good example of this is "Tern" for JavaScript. It (the backend indexing/completion engine) is absolutely killer, but the tightly coupled editor plugin/package just for Tern is decidedly "meh". I suspect because hardly anyone is developing, or indeed using, that specific integration.
---
Regardless of the overall merit of the Language Server initiative, the "matrix vs column" problem they articulate is bang on the money in terms of delivering quality (non-idiosyncratic) UI.
Apt for the client-side. But as for "there's no Go language server"---
It's still more sensible for a say Golang LSP implementation (the server side) to rely on the existing `gocode` tool and related tooling as underlying to deliver auto-completion, than reinvent the wheel here. These have already been battle-tested and encountered+fixed god-knows-how-many quirks that can occur when attempting to furnish a most-useful-for-current-cursor-context code completion.
Which kinda was my point. Anyone is free to wire up a say Golang LSP implementation without needing to redo all that the exising tooling already does well. If nobody felt like doing such an LSP server to-date, it's probably mostly because it's a young protocol still looking to find adoption beyond VS --- the ole' chicken'n'egg =)
https://discuss.kotlinlang.org/t/any-plan-for-supporting-lan...
Java seems usefully far along however...
The root cause seems to be the community. Few Kotlin developers use anything but an IDE so there's little motivation to bring tooling to other editors like LSP would