Chroma – A General purpose syntax highlighter in Go
github.com
github.com
Sourcegraph also wrote a server using syntect which provides an API for highlighting, which they use to power their new server, so you can use it from any language (at a cost): https://about.sourcegraph.com/blog/announcing-sourcegraph-2 https://github.com/sourcegraph/syntect_server
Tread lightly, for one more comment about Rust will summon the Rust Evangelism Strike Force.
The more modern grammars are written with a regex style that makes heavier use of the stack machine and so only uses regexes that can be turned into a DFA. There's ongoing work to use a layer on top of Rust's super-fast DFA-based regex engine to accelerate these grammars https://github.com/trishume/syntect/pull/34.
The problem with using non-regex based grammars is that you have to write them yourself. Syntect is something like 4000 lines of code but the grammars it uses total around 35k lines, and that's just the included ones, not the full ecosystem of online grammars. Basically unless you only want to support a small set of languages, a non-regex-based highlighting library is fairly infeasibly for a single hobbyist.
Will it really be faster? It didn't seem so from that GitHub thread.
As a data point for anyone curious, I'm using Syntect myself in a toy project. With Oniguruma (C NFA regexes), it highlights 200 lines of Rust in 40 ms on a 10 W TDP Celeron, which is all right, but a bit slower than I expected.
It's definitely possible to get better performance out of the underlying model, since Sublime does, but they have a custom DFA-based engine that can test regexes in parallel with captures, which Rust's regex engine can't.
Anyway, great job!
was this really necessary?
I think it's interesting to compare syntax highlighting approaches. If syntect was literally the same thing, but only usable in Rust, I wouldn't have commented. But syntect uses a different approach that's better for some use cases, and as demonstrated by Sourcegraph, is useable from a Go program (albeit with a cost), so is a plausible alternative to consider. The tradeoff is of course that it isn't directly in Go, and so may be slower, and also it supports fewer languages out of the box (although you could probably exceed Pygments with all online tmLanguage and sublime-syntax files).
What is the "richer highlighting and semantic information" to which you refer?
A static binary will be so much easier to maintain.
Repeat that mentality enough times, and I can't wait for the next heart-bleed to come out.
Distro-maintainers can't exactly be ecstatic about people using Go for more and more software.
https://gist.github.com/kyleconroy/a2741b9e6cf45beb3515d81ee...
Note that lexers.Analyse() will almost always fail at the moment, as I've only written support for a couple of languages.
Kudos! So often we see a translation of tool or library to another language, but no way to leverage existing data / code bases. Nice!
https://mobile.twitter.com/GoHugoIO/status/91158254224715366...
The improvement is due to two factors, with the first being by far the biggest factor:
1. Hugo no longer has to call an external `pygmentize` tool for every highlight. This removes the overhead of the fork/exec, as well as the (not-insignificant) overhead of the Python interpreter starting up.
2. Go is generally a faster language than Python.
The caveat with 2 is that Python can spend large amounts of time in C, eg. doing regex matching.