Show HN: Sucrase, 20x faster smaller-scoped alternative to Babel
github.com
github.com
We had a custom parser in our build toolchain in Perl that needed ~40min for a normal build. I rewrote it in F# (similar performance to C#) and got it down to 3min. Because I thought it was fun I ported it to Rust, and it now runs in 15 secs (mostly because of memory safe, zero-copy string handling and generally better control of memory allocation).
Probably in C# it could have been a bit faster than in F# if you use some of the optimization capabilities. But I doubt you could get close to Rust. Even reading one of the files in .NET takes as long as Rust takes for everything.
In C++ it would definitely be possible to reach Rust performance, it's just way harder to keep your memory intact without the borrow-checking compiler.
Now I don't know how Perl and JS compare, but I'd guess they're in a similar ballpark for performance.
I still doubt JS/V8 is faster than one of the VM languages (.NET or JVM), they don't need to be interpreted and have much less dynamic stuff so they can optimize better.
If you rewrote something like Closure Compiler in C, I doubt you could get more than a 2x-4x speedup. The bigger fruit lay in minimizing the amount of work you do each time through the optimization loop.
The problem with most hobby projects that write transpilers is, people take a simple app like TodoMVC and say "look how amazing the edit/refresh is with this transpiler", but design decisions made early can come back to bite you much later when you're trying to handle a codebase with hundreds of thousands of lines.
Transpiled higher level languages tend to encourage people to create many layers of abstraction and re-use a ton of existing libraries, which tends to bloat the inputs to the compiler.
Really good point. Closure Compiler, at least historically, rejected many valid optimizations because they aren't common enough to justify growing the rate of number of passes.
> If you rewrote something like Closure Compiler in C, I doubt you could get more than a 2x-4x speedup.
Closure Compiler is written in Java. I don't know exactly where the state of the art of JIT is right now, but I would bet C wouldn't be that much faster.
More importantly, this is a compiler targeted towards one-time compilations to permanently reduce large JavaScript payloads per millions of downloads, and not a compiler that is required during development. As such, blunting its effect to save a few seconds is pretty meaningless, so I doubt the maintainers ever considered "less optimization".
That said, it does allow for "dangerous" but more aggressive optimizations that require assurances from the JavaScript or you'll break the code. In that way, Closure offers user-specified levels of optimization.
EDIT: A secondary and less-obvious effect is that using a smaller number of total optimizations produces more internally-consistent code, as opposed to producing unusual and internally-unique constructions for rare optimizations. Internal consistency is great for the next step after compilation: run-length compression.
Yes, but still, some projects are orders of magnitude larger than other projects. Also, some users might be willing to wait an hour, others only a minute.
There are, essentially, an infinite number of optimizations you could make to Closure, though probably several thousand are reasonable. Every marginal optimization needs to run though the entire AST and many of them require prior optimizations to be re-run. As 'cromwellian pointed out, the number of passes is the dominating factor in speed. At some point, it's no longer worth it.
For Google production code, we typically let things run long, because if you shave off say, 30k from Gmail * 1 billion active users, you've just saved a lot of bandwidth.
I'd be interested in seeing what a large typescript project would look like run through https://github.com/angular/tsickle.
https://github.com/google/closure-compiler
https://github.com/BuckleScript/bucklescript
https://github.com/fastpack/fastpack
Seems like on the average it does offer a boost in performance.
And there is some aditional work providing javascript parsers for rust (which you could build tools like babel on top of): https://github.com/dherman/esprit
But I believe the topic here was about runtime performance of using a language to compile JS, not about the build speed of working in that language itself. In which case you’ll still get some wins writing a JS toolchain in BuckleScript (compiled to JS), just from the JiT-friendliness of the BuckleScript JS output.
But realistically, you’d be compiling to native OCaml through the same codebase. We did see a 10-25x perf jump from converting a part of a Babel pipeline to native OCaml. I mean, these languages are basically designed over decades with AST manipulation in mind, so that’s not surprising.
I certainly thought about writing this in Rust or C++, and still plan on exploring that. Still, it's nice to stay in the JS world, e.g. easy integration with webpack and all of the other tools.
Edit: https://github.com/alangpierce/sucrase/blob/master/integrati...
Annoyingly, Webpack still spends a lot of time parsing the JS files with its own parser, so you cut down on the transpile time but still need to do a relatively slow parse of all files. I've thought about making Webpack (or similar) use a fast parser like the one Sucrase uses, although it certainly seems like it would be a lot of work.
As in, how much does this actually buy you in a real world scenario?
Some specific numbers I just measured for my codebase at work:
Babel/TypeScript, cold cache: 66 seconds
Babel/TypeScript, warm cache: 55 seconds
Sucrase, cold cache: 52 seconds
Sucrase, warm cache: 49 seconds
This is on a codebase of about 400,000 lines of code (some JS compiled with Babel, some TypeScript compiled with tsc). Transpilation is parallelized using happypack (and I'm running it on a 4-core machine), but webpack parsing/processing is all single-threaded, so it ends up taking more time. I've done a little prototyping on using a Sucrase-like approach to speed up webpack (most of it is indeed parse time), but it would certainly be a project to get it working in practice.With a warm cache, it's not running sucrase/babel/typescript at all. I think the running time difference there is because Babel and TypeScript are emitting ES5, which is a little more code for webpack to parse. Probably a more correct comparison would be to configure Babel and TypeScript to target newer JS, but these should all be seen as rough numbers anyway.
Technically there are a few interesting differences under the hood, such as Sucrase not wasting time generating a full-blown AST (which Bublé needs to do in order to handle things like block scoping).
Anyway, I won't waffle on any further as I'm sure Alan can explain the differences better — I just wanted to chime in to say that I'm excited about Sucrase. I've written a [Rollup plugin][1] for it, and I'm planning to use it with my TypeScript projects.
Architecturally, they both skip a bunch of work that Babel does (AST transformations and AST formatting), but Sucrase is a bit more ambitious. Bublé does a normal parse step and then does replacements using magic-string. The original hope was that Sucrase would be able to skip parsing and just tokenize the JS and do all of the necessary transforms from the tokens. It turns out that that's basically impossible, but Sucrase still tries to find a middle ground where it does something reminiscent of parsing but doesn't need to produce an AST at the end. It's still able to resolve variable scopes (which is needed for the TypeScript and imports transforms) by providing an array of scopes in addition to an array of tokens, and it provides enough useful context in the tokens to put the pieces together if necessary.
I am guessing some loading time would be improved if code converted from ES2105 to ES5 is not required
It would be great to have a Browserify plugin (I'm not asking for anything, just speaking my thoughts).
Regarding Browserify support, PRs are welcome. :-) https://github.com/alangpierce/sucrase/tree/master/integrati...
My specific use case so far has been running Mocha tests on a large codebase, so it's replacing babel-register (which uses a cache) with sucrase/register (which doesn't). For that use case, it seems to significantly speed up test startup time with a cold cache and somewhat speed up startup time with a warm cache. Just loading and saving the babel-register cache takes a fair amount of time, at least in Babel 6.
It's good to keep caching in mind when thinking about this stuff. In many cases, Sucrase won't be faster than loading results from cache. Sucrase helps with initial startup time and avoids the cache fragility issues that I've seen with lots of caching systems. And, of course, you could cache Sucrase results and get something at least as fast as cached Babel.