Microlight.js, a code highlighting library
asvd.github.io
asvd.github.io
I've looked through some textmate-style syntax highlighting packages (used in sublime and in github's atom and probably others), and most of them need big (or somewhat big) sets of alternations for a bunch of keywords, and more often than not they are just set up as a list of full keywords with no thought to order or size.
Combining them into something like the below should theoretically be faster while also taking up less space (which is important in web libraries), and I feel like it wouldn't even be all that difficult.
de(bugger|cimal|clare|f(ault|er)?|init|l(egate|ete)?)
Is there something out there which can do this, and would it even be worth it or is this something best left to the JIT/optimizer of the regex engine?https://en.wikipedia.org/wiki/Trie#Algorithms
I'm not so sure it would take up much less space though, if you take gzip compression into account. See for example here:
https://github.com/google/closure-compiler/wiki/FAQ#closure-...
As for the size aspect, it might not make much of a difference in the gzipped filesize but there are other benefits to smaller raw source as well (but honestly I don't have any idea if it would be at all worth anything, I'm assuming no).
Example : "{|||Wow! |Hey! |Looks {nice|cool}! }{{I like|I love|Like|Love} {this|your}|{Very nice|Nice|Cool}} {photo|foto|picture|pic|image|img}{.|!|!!| :)| ;)| :D| <3}"
[1]: https://github.com/noprompt/frak
edit: typo
I might take a look at this later and see if incorporating this or something like it has any kind of meaningful impact.
TLDR; frak transforms collections of strings into regular expressions for matching those strings. It is available as a command line utility and for the browser as a JavaScript library.
An 2009 article on the V8 blog about a regex optimization they did: http://blog.chromium.org/2009/02/irregexp-google-chromes-new...
A general purpose regex optimizer can't make assumptions like you don't care about the ordering of sub-groups, but a tool like this can.
If you are just looking for the fastest way to to match any 1 of these 40 keywords, this could make a fast regex that can minimize backtracking.
Like I said, I have no idea if it would make any real difference, but if it did incorporating something like that into a build process for syntax highlighting packages could improve performance for a bunch of people.
It’s a command in the “bundle development” bundle.
1. Use of regular expressions. Hard to get effective solution with that. Yet that famous parsing HTML by regex answer: http://stackoverflow.com/questions/1732348/regex-match-open-...
2. It modifies the DOM. Not desirable in most cases and heavy: each DOM element takes 0.2k..1.0k in memory. Yet DOM handling in browsers is O(N) complex (N - number of DOM elements).
Just in case, in Sciter (http://sciter.com) I've added an option to style character runs so without DOM modification: Selection.applyMark(runStart, runEnd, name) and special ::mark(name) pseudo-element in CSS to style those runs.
Illustrations: http://sciter.com/tokenizer-mark-syntax-colorizer/
[1]: http://pygments.org/
I don't like waste for the sake of waste.
Let's say you used the smallest markup possible for each highlighted token. Something like <i class="x">[token here]</i>. You'd only be able to serve up at most 120 server-highlighted tokens before this JS library becomes a smaller payload, and that's without even defining the CSS styling or considering tokens longer than 1 character.
120 isn't even enough tokens to highlight the tiny code snippet at the top of this demo page. The same tiny snippet highlighted with Pygments comes out to 5819 bytes of markup alone (no styling) – already more than 2.5 times the size of this whole library. Plus you can highlight any number and size of code snippets while just serving the library once... which one is wasteful again? :)
It's not going to revolutionise syntax highlighting, but I think it has plenty use cases.
Almost un readable on a fairly quick phone...
There are definitely unacceptable performance issues with the syntax highlighter javascript.
It looks great, it's a cool project, but it is not suitable for use anywhere.
It actually can use RegExp's, but they're only being used to locate the language primitive types words (e.g. char, wchar_t, etc..).
Most of the library is actually written as a hand-implemented state machine rather than relying on regular expressions.
That's eerily beautiful.
And GPU compositing, 'cause damn.
I'd check myself, but I'm away from a PC for a while.
http://asvd.github.io/microlight/haskel.png
well, since the lib is general, it's built upon compromises. But I am open for suggestions concerning updating the logic for some particular cases
Hardly for "any programming language".
There are few cases where having wchar_t in any code snippet would not indicate some sort of keyword/type annotation
library size is extremely compact
2.2k, seriously, can you imagine!
I wonder if I can do better than this, since it seems mostly a matter of a few regular expressions and then DOM manipulation?EDIT
Reviewing the source code [0] the state-machine approach, when properly implemented can beat* [1] an equivalent RE performance-wise.
It still may be interesting to see if the code size could substantially reduced this way.
[0] https://github.com/asvd/microlight/blob/master/microlight.js
* [1] Disclaimer: In my experience, and admittedly not using javascript. I have recently confirmed minimal hand-implemented state machines generally beating Regexp's in Golang.
What we are actually looking at here is the fact that some words are highlighted in any programming languages despite the fact that the language itself is unknown.
Usually the user has to manually select which programming language is being used and THEN the words get highlighted.