Speedy Web Compiler – A Rust port of Babel and Closure Compiler
github.com
github.com
I actually like JS as a programming language for the web, I just don't like scripting languages in general for command line tools.
JS is a universal binary format that tends to just work from my experience. You do occasionally run into issues if you're on an old node version, but other than that it has been pretty reliable for me. But it may just be a difference in familiarity.
Of course, wasm will need threading first for this project.
So, is npm getting a wasm build infrastructure to support this?
I get most of my Rust binaries through cargo when on Ubuntu, except in a few cases when I really want rg (still not available in Ubuntu 18.04) but don’t want to waste 25% storage on the Rust tool chain.
What do you think is the essential difference between a 'script' and a 'program'?
I never said "program". I contrasted a script vs. a "binary" by which I meant a compiled executable. What I meant to highlight was the difference between the JS toolchain which is delivered as an artifact which needs an installed JS interpreter to evaluate. I've had issues my node/npm version to what the script expects, as well as managing different expectations by different scripts. In contrast, executables which are compiled for your architecture in advance and can be executed on their own without further dependencies are much simpler to use. Just download and run. Yes, there's static vs dynamic linking, but I mean to highlight how nice the (typical) Go workflow is which tends to statically include everything, moreso than rust, but rust isn't far behind.
It's a murky distinction, because it's very common to preprocess programs before distribution to the extent that the original code is hardly discernable; for example, using Uglify or transpiling ES2017 to ES3. And you can package programs in both Python and Node such that they actually don't depend on your system having a Node or Python interpreter installed. And I doubt people would refer to C programs as scripts, even if they're distributed in a BSD ports- or Gentoo-style packaging system that defaults to distributing source code rather than precompiled binaries.
But I think a decent rule of thumb is that if you can run the original source code file(s) directly by adding the right shebang to the top, it's a script.
the library implements how to do the lower level -- which things you can do, and takes care of doing it
and the script deals with which things you want to do and when
some examples:
I made some opencv bindings for JS, and the library changes did the heavy image processing algorthims, and the JS script chose which ones to use with what parameters
for some video game or another, there was a library for what an NPC can do, and scripts for how they act within the world, including things that are more like movie scripts detailing whow says what adn where they walk
It is a hardware interpreter.
You could bake a JS interpreter in silico. There are actual hardawre implementations of the JVM. You can also run machine language in emulators.
CPUs are just hardware interpreters for their respective machine language.
I guess you could argue that an interpreter is specifically a piece of software, in which case my point wouldn't stand.
https://github.com/alangpierce/sucrase
In my case, it's still running as JS, but rearchitected to solve the more straightforward problem where you don't need to compile to ES5. I've thought about rewriting it in Rust to see how much faster it gets, though currently I'm trying to get it running in WebAssembly via AssemblyScript ( https://github.com/AssemblyScript/assemblyscript ), which has been promising.
I'm curious about the about the Babel performance comparison. Benchmarking JS is tricky because, from my observations, the performance improves by a factor of ~20 if you give it 5-10 seconds for the JIT to fully optimize everything. https://github.com/alangpierce/sucrase/issues/216 . Rust and WebAssembly both have a significant advantage in that sense when running on small datasets.
sucrase: 900.720ms
swc: 4514.190ms
babel: 16126.752ms
Sucrase had the jsx and imports transforms enabled, swc had all transforms enabled (they're not configurable), and Babel had @babel/preset-env enabled. So I'm a little skeptical of swc's 20x speedup claim, but certainly depends on a lot of details.Feedback/questions always welcome, of course! Feel free to file an issue or chat in gitter. Fortunately it's a pretty well-scoped project that's mostly done at this point.
https://github.com/alangpierce/sucrase/issues/216
Most recent prototype branch (doesn't include some patches to AssemblyScript to get it working):
https://github.com/alangpierce/sucrase/commits/assemblyscrip...
I've also chatted about it a little on the AssemblyScript Slack.
Babel's parser alone has quite a formidable suite of fixtures (a few thousand as I recall): https://github.com/babel/babel/tree/master/packages/babel-pa...
The nice thing about implementing a JS-JS compiler is that there is already a huge number of tests out there, like these, that you can use to TDD your compiler's edge-cases.
SWC's CONTRIBUTING mentions this:
> Include tests that cover all non-trivial code. The existing tests in test/ provide templates on how to test swc's behavior in a sandbox-environment. The internal crate testing provides a vast amount of helpers to minimize boilerplate. See [testing/lib.rs] for an introduction to writing tests.
But I cannot find a `/test` directory, and it does not appear in the .gitignore either.
EDIT: Ah, the tests are in `/tests.rs` with JS embedded inside rust code. This makes it a bit harder to compare directly with babel's suite, but claims to be lifted from test262, which babel also based many of their tests on (albeit recategorized). hzoo's comment here has details: https://news.ycombinator.com/item?id=18746905
That's not to say that babel isn't too slow or that this project isn't correct, just that it would scare me to try to work with it. Does babel have a regression suite? Does this pass it?
[1] current project is very slow w/ 250kloc (2.5mm kloc w/ dependencies). even seen it be slow on my last project which was only 40kloc.
Hot reloading _helps_, but it's definitely an issue that javascript tooling spends seconds on tasks other languages' tooling spends tens or hundreds of milliseconds on. (The issue isn't just limited to tsc either; just starting the build system can take a decent amount of time with webpack, while make or cmake starts in milliseconds.)
In terms of perf, I'm sure rust will be much faster than JS. And this project's main goal is to be a faster version than Babel. I would suggest the perf test would be different though because it's testing code I would consider not to be representative of what Babel actually runs on. Example with the transform test: https://github.com/swc-project/swc/blob/222bdc191fcab7714319... it's a very small file with minimal transforms even necessary at all. And the parser test is for all minified js files, none of which are written in ES6+ https://github.com/swc-project/swc/blob/master/ecmascript/pa.... Currently it's not really testing ES6+, just how at a baseline it's much faster to process a ES5 file, so it's possible that it might be even faster.
edit: I will note that Babel does a lot more than what this project is doing currently (plugins, other proposals, ts/flow, sourcemaps, etc) but this is more about what could be so if people are willing to more on this effort in the long run (with funding/support/people) it could be super useful! I'm curious about how that will work when our project has such a struggle with contributors and Babel itself is written in JS. (I work on Babel :D)
Unless we're talking about "understanding neural network internals", which is a big open problem, there just really aren't "black magic" problems inherent to particular subjects. It's almost always poor spec definition or poor code design, if there are any hand-waivey bits to be found.
If anything, compiler design (along with other subjects like text editor development) is a very deeply studied, very well understood, very well documented field. It should be easier than any other fields less central to computer science and software development, just on account of not being so thoroughly studied.
EDIT: fixed autocorrect problems
Compilers are also definitely not well documented. I think they're rather notoriously undocumented with much of our knowledge of the subject passed on orally and existing only in the knowledge of people working in the field.
For example a couple of years ago I went looking for details of a very common compiler technique, called safepoints, used in many major industrial compilers today, and found that there was essentially no written information on them at all. That's pretty extraordinary.
1. Safepoints are kind of off-topic, because they're part of the interplay between a complex optimizing compiler and the garbage collector/runtime. Babel and its ilk are source-source compilers targeting a language that doesn't have manual memory management. To my mind, that makes things much simpler.
2. How long ago were you searching for safepoint info? I've collected several articles and talks relating to safepoints. I don't know of anything that discusses safepoints in a way that you could implement them from scratch just based on the discussion, but they're tied to a complex part of the runtime, so any example/tutorial code would probably be parochial to a particular VM.
The best discussion of them at the moment is this one - and as you'll notice they're essentially having to do analysis of what is already out there because it just wasn't written down anywhere at the time. They don't have much of a 'related work' section or a bibliography because there wasn't much to refer to.
https://dl.acm.org/citation.cfm?id=2754169.2754187
And implementing them from scratch is what I wanted to know how to do. For a lot of compiler topics the only way to learn is to sit down with an existing expert and physically ask them. Compared to other CS fields that’s really bad.
There is some truth to this, but as someone with a PhD in compilers, let me say that this area is very accessible to people without PhDs. The first time I met @dannybee, I asked him where he went to grad school. Turns out he went to law school. Didn't stop him writing the GCC alias analysis framework though.
One of my focus areas during my degree was system programming, including compilers and language design.
Already in the late 90s, there were several good books about compilers.
There are tons of introductory texts, a few intermediate texts, a couple of very narrow advanced texts, and then it just runs out. A great deal of our knowledge is undocumented.
But there are classes of bugs that the language eliminates. Thread unsafe dataraces (not important in JS), unhandled exceptions, with no GC it helps reduce memory usage and make it more consistent, the type system allows you to build complex state machines easily and correctly, the type system itself can allow you to express states that should never occur as compilation failures, the compiler protects against concurrent updates in iterators and other situations.
The list goes on, but these are the things that make it such a powerful language and allow me to focus on logic and not runtime bugs. Oh, and it’s really fast which is a nice bonus, generally as fast as C.
I just know that even if the compiler of the language I am writing is very smart and can catch a lot of my silly mistakes, it would behoove me to think that I could write "bug free" code as a result. I would still write my code defensively, even in rust.
I would argue it is much faster to get to a final, 'correct' working version because you spend less time debugging. Time is also drastically cut down the more familiar you are with the language, but still, I'd love to see some data on dev time between languages with robust type systems vs dynamic.
Mostly I would say fewer bugs; the bugs that are avoided vary in severity. Many, probably most, are the severe but obvious bugs where you would have hit the first time you tried running the code in another language, but which are instead now caught by the compiler—that’s why “time to first run” is not a reasonable metric for comparing writing code in such languages, but “time to first meaningfully successful run”! But it’s also common for Rust to catch and prevent far more insidious bugs; especially, in my experience, in two areas: firstly, with respect to ownership—things like passing data structures around and improperly mutating them where you should have cloned it instead, which is structurally impossible in Rust (disregarding things like Rc<RwLock<T>>, which make it possible, but make the intent very much clearer); and secondly, where enums (tagged unions) can be used, due to the greater and more precise expressivity.
I have much greater confidence that my first pass at a piece of code will work correctly in Rust, once it compiles, than I do in JavaScript—and I’ve been working in JavaScript a lot more than Rust in the last couple of years, yet it’s still true.
Expressive type system, strong and static - JS is weak and dynamic. This is not limited to Rust - but if you're aiming for correct code pretty much anything after C/C++ and maybe PHP is better than JS and it's million and one way to write unexpected bugs and 10.000 ways to do the same basic thing (with an NPM package no less) of which each dev/library chooses it's own path for cool factor.
Who would even argue that JS is better for writing correct code over a language that's obsessive about security and correctness. Last time I checked Rust wouldn't even upcast ints FFS, meanwhile in JS you can't even be sure what "this" refers to in any given context and 0 == '0' == [0].
If you change Rust to ML, the phrase would be way more natural.
Other bugs such as the goto fail bug are eliminated by dead code spotting and consistent formatting.
You can write bugs in Rust but they tend to be business logic bugs, or incorrect explicit decisions.
[1] Naturally excluding the 18 hours of productivity that will have been killed getting said unnecessary javascript->javascript build system set up and running in the first place.
this being said, i do wish this project (and any others like it) great success. we definitely do need something faster when compiling js.
(this tool reminds me of a similar project for a small client. they required a fast js compiler. the project was a failure because no-one wanted to debug non-js code when the build failed)
This not only makes it easy to keep up to the latest of ECMA's standards and extensions like JSX, Flow, TypeScript – it also makes it surprisingly easy to implement your own syntax extensions to JS (eg; https://wcjohnson.github.io/lightscript/, which I have worked on).
Curious how this project is thinking about plugins and staying up-to-date sustainably.
Does it really? Last time I looked into this, the "parser plugins" were all fake: single-line modules which simply switched on hidden config flags in Babel to enable functionality already hard-coded into the core parser. Left a bad taste in my mouth that the developers weren't up-front about the true (non-)extensibility of the system.
This is the kind of thing Rust should be excellent at.
I would actually think this (compilers) is a thing where Rust doesn’t have that big of an advantage. Compilers are very special. They mostly deal with some kind of tree transformation. And as soon as those trees use owned nodes (E.g. in the form of Box of Rc) things are not that far from what a speedy managed language would do. Compilers are also non realtime at all. It doesn’t really matter if there’s a GC pause or not, and whether individual allocations are quick or slow. The only thing that matters is the overall throughput/performance.
Now this statement shouldn’t mean that Rust is bad at doing compilers (enough projects tell otherwise). Only that in other domains (E.g. realtime and low level code, graphics, audio, etc) it’s advantages might be even bigger.
I think Babel 7 had some perf improvements (including tweaks contributed by Benedikt Meurer from the v8 team), but I don't think they've ever had time to focus on just improving perf.
It's true, speed is not really a concern on my mind at the moment - rather that we have enough people even willing to contribute back at all
Is that somewhat similar to D's std.parallelism module? A simple example of its use here:
https://jugad2.blogspot.com/2016/12/simple-parallel-processi...
There's still going to be a finite number of cores on a machine. I can't see gaining parallelism being that big a boost (as opposed to gaining concurrency(if not already there))
Also a quick check with google trends does seem to support that feeling.
https://trends.google.com/trends/explore?cat=31&date=all&q=S...
Swift is probably used a lot less than Babel but it yields a lot more spikes...
One complication is the fact that Babel results are often cached, so with a reliable cache, Babel performance barely matters at all. However, every caching approach I've seen has reliability issues in practice. Certainly a tool is more pleasant to use without a cache, since "maybe I have a bad cache" is never a worry when running into strange behavior.
It also depends a lot on your scale, how you compile your code, and what operations you want to be fast. When switching to Sucrase (my own Babel alternative project, mentioned elsewhere) on a large codebase, webpack startup went from 40 seconds to 20 seconds, though incremental builds weren't any different. Running tests (in node with a require hook) became much faster, from minutes to seconds when uncached and about a 2x speedup when Babel pulled from cache. In both of these cases, the remaining time is now in other things: webpack processing the files, node running the imports, etc. So there's a lot more besides Babel, but Babel still is a non-trivial component in the cases I've run into.
On that note, they don't say what plugins/presets were used in the Babel comparison. Was it the same set of features that are supported in swc?
I would recommend using the benchmark at https://github.com/v8/web-tooling-benchmark in the future if it's currently not possible since currently it only seems to be testing ES5 (jquery.min.js) for the parser and a short ES6 snippet for the transforms.