Nectar: A Native JavaScript Compiler Inspired by Crystal Lang and Nim Lang
blog.seraum.com
blog.seraum.com
I'd be extremely (and pleasantly!) surprised if this turns out to be real. It is very common for people to compile a subset of the language and achieve great speedups, only to discover it doesn't scale to the full generality of the language. That would be my guess for what's happening here.
Best of luck to the Nectar team, looking forward to seeing what you've got. If this works as you say it does, you'll have built something incredible!
- they're not really specified anywhere
- they use `eval` and similar dynamic constructs
- they allow you to mutate your data in weird ways (in particular, you can change variables dynamically by knowing the name of the variable, aka "variable variables").
I looked at how to solve these problems in PHP by embedding the PHP interpreter into the compiler as well as into the compiled program. Then I looked at how to do alias analysis in PHP - alias analysis is a type of static analysis to figure out which names in the program point to the same value. This had been done before for a subset of PHP but not for the full generality of the language, in particular dealing with variable variables.
PHP Specification: https://github.com/php/php-langspec/blob/master/spec/00-spec...
JavaScript Specification: http://www.ecma-international.org/publications/files/ECMA-ST...
PHP has an unofficial spec now but didn't then. Here's what I wrote about it when it came out: https://circleci.com/blog/critiquing-facebooks-new-php-spec/
This doesn't make sense
> - they allow you to mutate your data in weird ways (in particular, you can change variables dynamically by knowing the name of the variable, aka "variable variables").
There are a few ways to naively do this. The first and easiest I'd say is just to make a jump table of all the variables in scope. Take a hash of the string and access that. The hardest way is to optimize out this access. Check to see where all of the callers are and inline the function to make the string hard-baked in. After that you can simply replace $some_string() with string_contents(); and get 0 overhead.
> - they use `eval` and similar dynamic constructs
This is harder but eval is largely discouraged and mostly unused (I hope). It can still be supported by linking in the compiler at runtime and calling the compiler on the string passed. It's slow but would allow it to work. Alternatively you could also do the inlining optimization trick.
Have a read of my thesis. I researched whether eval is used much (answer, yes). I looked at the challenge of inlining (you have to know what function is called - non-trivial). The naive jump-table thing you describe is basically impossible (or at least, it makes optimizing around it impossible). If you don't know the string value, all bets on all variables are off. You need to keep the value out of registers, you can't reorder operations, you're basically at square one.
Getting even remotely comparable performance in an ahead-of-time compiler—without spending eons waiting for complex whole program analysis—is, as far as I know, an unsolved problem for real-world programs. I'd like to see what "really, really fast" means for this compiler.
so essentially it's likely "JavaScript-like compiler" rather than "JavaScript compiler"... similar to how Crystal is Ruby-like but not really Ruby.
However, apart from that, I don't see what can't be done: you can simply look at what functions get called, with what parameters, and generate code for each fundamental type (doubles, strings, …) used. It's like making C++-style templates for all function parameters.
The flip side is awful compilation times on large projects, and (I expect) poor results on maths for which JS VMs detect small ints.
I actually wanted to do something like this as a follow-up from my experiments with JS type inference, but I lacked time…
A simple example would be something like `x = (isUserOnMars() ? "foo" : 1)`. The type of `x` depends on if the user is on Mars. In practice they never will be, but a compiler can't tell that and so must consider that case and make code appropriately generic. A JIT however can see that it always appears to be false, and optimise around `x` tending to be a number, with bailout if it's wrong (which it likely never will be). Then the use of `x` may propagate a long way through the program in all sorts of places, extending the effect.
Profile-guided optimization seeks to do this for compiled languages like C++, but you also have a boatload more data to do it with. And, y'know. You're profiling the application. You're running it.
My first guess for the chopping block would be the functionality of franken-objects like Function.arguments (especially properties like arguments.callee).
Moreover, you can carry out heavy static analysis that can often narrow down the possible types for a and b. The main problem with this is that static analysis that is good enough will run very slow, hence it not an option for the web-browser.
But then many people pooh-pooh'd me in a similar way when I said I could make Ruby 10x faster which was my project for the last few years, so best of luck to you!
JITs have been around for 30+ years and they're umm OK. As in not magical.
JIT performance benefits are highly dependent on the language.
They do. Just like people write for PyPy for performance.
The only challenge to Java's performance is from languages without dynamic GC and with detailed static type information: C, C++, Rust. Even then, Java is usually "only" half the speed.
And AOT is coming for fast start-up, not good steady-state performance.
Indeed the performance benefits of JITing have been difficult to measure: https://arxiv.org/abs/1602.00602
This isn't true. At run-time you only need to pay the cost of dispatching once per (type of) input variable.
C++-style templates are popular enough that this concept should be readily understood by now: You pay with heavy compile times (IIRC Crystal takes almost a GB of ram to compile its compiler).
> Benchmarks should be provided before making performance claims!
Yes, absolutely. And reproducible.
That also means you can't inline (potentially) polymorphic calls. Inlining is the great big hammer when it comes to optimization.
Neat for sure, but not a lot of stuff you can "usefully" do in asm.js in terms of end-user-facing "RIA", you have numbers, byte arrays, function calls, numeric operators, primops keywords. (Hell of a way to go type-safe: eliminate all types but numbers and byte-arrays! ;) No JS-native/runtime-own strings, no DOM, no built-in objects (window etc), hard to impractical to try to do other record/object handling. Wrap all interaction with stuff you don't get in asm.js in exports and "foreign"s. (OK fair game for code generation but still what a mess and constant switches between the asmjs context and the normal interpreted context aren't free either --- the bulk of "RIA" logic seems to be string/list ops, DOM/window interop, basic cheap control flow.) All the funky emscripten stuff that works in asm.js just also-compiles-and-then-uses the original memory/byte-based logic for these (records, lists, strings, objects etc) of the source's native runtime/compiler/libs.
As for wasm, seems to aim to mostly follow in the above footsteps "at first" (read, the next half decade at best, realistically...)
(But yeah, asm.js is of course still neat for overwhelmingly-numerical high load computations (game-playing / media-editing stuff etc comes to mind) as well as showcasing/proof-of-concepting both older C / OpenGL games and newer 3D engine demos..)
Nim, likewise, is static.
In my experience writing real world server programs, Node.js is faster than Go in many cases. Part of this is a lack of Go library maturity, but it's also due to the fact that V8 optimizes code at runtime based on how it's being used. This project seems to come from a naive, "interpreted slow, compiled fast" view of the world, leading to false assumptions like Go is going to be automatically faster than Node.js because Go is AOT compiled and Node is JIT compiled.
Maybe it's an idea to make this a native Typescript compiler as it will profit much more from being compiled than pure JavaScript will?
This way it would probably still be possible to compile pure JavaScript but providing type information would allow the compiler to apply many more optimisations.
If I'm not mistaken V8 already does something similar to this with WASM (and I think asm.js).
https://github.com/andrei-markeev/ts2c
Very early, but the author seems to be actively working on it.
If you want to take advantage of Go's speed alongside JS, check out https://github.com/ry/v8worker
Indeed, this blog post is very vague and offers no solid evidence that this project even exists. I was very sceptical of it from the beginning and I'm surprised it was voted up so much. I guess generating hype is easy when JavaScript is involved.
If anyone knows how and can make this happen I'd like to see this on some godbolt-esq site.
http://www.2ality.com/2015/02/soundscript.html
https://github.com/v8/v8/wiki/Experiments%20with%20Strengthe...
Looking forward to more details.
How does it compare to EncloseJS and nexe (which grab V8's unoptimized compiled machine code)?
What does it target (ES3 CP, ES 5, ES 6)?
Does it require/support async server-side apps?
Does it have http, http/2 or other IP protocols?
Can it compile regexpes to native code like V8?