Why Is Esbuild Fast?
esbuild.github.io
esbuild.github.io
It's annoying to me how much time I had to spend outside of school before I realized that this is the real recipe for fast code. In particular, even code that is nominally O(n^2) by standard measures is still a dis-economy of scale: the bigger the problem the slower it gets. In school, there's a premium given for finding a polynomial-time solution, but in practice every job where I've been tasked with making something fast, only linear time complexity was good enough.
In practice these linear algorithms tend to also be completely data parallel and have great cache locality. Not only do these properties make the code fast the second you harness them, but then tend to be exceedingly simple architectures to maintain over time.
For some other reading I really like on software performance in the real world:
But of course we remember from our CS school that some things simply cannot be done in linear time. I’m not sure what happens when your employer tells you to sort a big list and insists that only linear time complexity is good enough.
I only said it took linear time, I didn’t say it was fast...
Boy am I glad I didn't take CS
This means quicksort and the like end up faster in practice for any "sort this list" operations. Radix/bucket sorting can be SUPER awesome when what you need is a single level of bucketing at all rather than exact order.
No problem, just create a file of the necessary size on disk if you don't have enough RAM available.
Today's comment was brought to you by a large number and the letters big-O and c... where c, a large number, is the hidden constant of proportionality waiting to bite you every time you mention big-O. :-)
If we're reading from disk, then we DEFINITELY don't want radix, as then it's almost guaranteed that all time will be dominated by loading pages rather than cpu cache locality being leveraged by linear time algorithms - which is what this entire discussion is about.
Of course, it was only half a joke. For some pages, we really did want data quickly, so we were willing to accept stale data or somewhat inaccurate data, it that helped. For other pages/sections, accuracy was more important so caching wasn't possible or invalidation was critical.
But product design is always about tradeoffs between what features are desirable, what features are possible, what features are easy to build, and how features interact. Features that are possible but not easy to build can sometimes be replaced with features that are easy to build and close enough.
Presumably you weren't tasked with making something fast, were linear time complexity was overkill. In other words, quadratic complexity may have been fine in many cases, but you weren't tasked to make those fast.
I'm not saying you claimed otherwise, but somebody might misread you as saying "in practice, only linear time complexity is ever good enough".
It's also worth noting that if your size is bounded a simpler quadratic algorithm might be more maintainable overtime than a linear one (if it's easier to grok). Say max(n) = 25, it doesn't really matter if it's linear (say 4N = 100) or quadratic (N^2 = 625). In practice they're both fast enough (even though one is 6x faster) and not worth optimizing but if it can grown unbounded that's when the BigO runtimes start to matter. For something like ESBuild your input size grows larger and larger as your codebase increases and that's when you start to see slowness creep in.
[0]: https://github.com/evanw/esbuild/blob/master/internal/js_par...
Still a great achievement to produce a bundler and minifier as a one-person team, but not sure we should use the amount of code line changes as a measure for productivity.
Relevant Macintosh folklore: https://www.folklore.org/StoryView.py?story=Negative_2000_Li...
I've written parsers/lexers/code generators by hand for much smaller toy languages with no thought about performance at all and those were huge undertakings for me. I was just trying to convey my impressedness with some numbers that we can relate to somewhat.
But yeah, he has a "net" loc contribution of 284-148 = 136k lines of code. Assuming the code base is resonable (which I don't doubt), it's a pretty big project which he has built, which I'm impressed by.
> Everything in esbuild is written from scratch.
> There are a lot of performance benefits with writing everything yourself instead of using 3rd-party libraries. You can have performance in mind from the beginning, you can make sure everything uses consistent data structures to avoid expensive conversions, and you can make wide architectural changes whenever necessary. The drawback is of course that it's a lot of work.
> For example, many bundlers use the official TypeScript compiler as a parser. But it was built to serve the goals of the TypeScript compiler team and they do not have performance as a top priority. Their code makes pretty heavy use of megamorphic object shapes and unnecessary dynamic property accesses (both well-known JavaScript speed bumps). And the TypeScript parser appears to still run the type checker even when type checking is disabled. None of these are an issue with esbuild's custom TypeScript parser.
1k lines per day sounds pretty reasonable when writing your own TypeScript parser, among other components
He's obviously changing things to be more optimal as he goes, not transcribing.
Having built a compiler before, this seems normal, especially since the parser is written manually (as opposed to generated). Parsers are always a lot of lines of code.
Not all of them are the most complex, and in fact a lot will just be defining all the different nodes of your AST and branching on them, which, as blocks do, tend to inflate code a lot. I wouldn't expect a JavaScript parser to be any more compact though, or the development (especially in a codebase with few contributors) to be any less "thrashy".
I know what you mean, but I also hate using lines of code per day to measure productiveness, and I hate it when managers in corporations look at number of lines added in PRs to measure how productive a dev is.
I was going to say "start of coding" but some tasks are just so boring that I may procrastinate on starting it. That also indicates an element of difficulty (in giving a shit), but not what we're trying to measure here.
I discovered the Mikado Method quite on my own as a consequence of trying to tackle increasingly complex cases of 'large-scale' or 'top-down' refactoring before accepting that as the oxyest of morons. I was sold when I realized that when you finally discover the crux of the problem, the 'Rome' that all roads lead to, between half and 80% of the code you just wrote isn't strictly necessary. Throwing it all out wholesale is how you avoid code hoarding. Yes, once in a while you go out to the proverbial trash can to pull something back out that you actually did need, but the tabula rasa aspect is profoundly useful psychologically, especially for that 'tax' bit I mentioned above.
There is always some function in the middle of the call tree, that if you add or change the meaning of an argument then the feature or bug fix becomes a logical conclusion of that new semantic. Either the code above it or below it hardly needs to change, decoupling the feature from all but a handful of functions/concerns.
A bad programmer will happily plow through adding a new argument to fifteen method calls and then propagating that change to a hundred call sites. And that code will be buggy as hell because their coworkers have already written them off. If it's too much code for others to review, you can be sure there are bugs in it. And if you didn't listen the last ten times people told you to stop writing so goddamned much code, you aren't going to listen this time, either. So have at it, sport. We'll just make sure you get the blame for the bugs, until someone in management wakes up. If they don't, then they see that the sections of code people "won't touch" but you will are getting bigger, and mistake your ownership of this code for a sign of prowess instead of a sign that you are inspiring apathy in others.
https://talks.golang.org/2012/goforc.slide#10
I find it similar to the fact that, while I am thankful that the operator's manual for my toaster was written at a 4th grade reading level, I find it unlikely that anyone has that task in mind when they say they want to be a writer when they grow up.
Some serious work there.
While that is impressive, I think this is also an indication of a problem: the grammar is becoming unwieldy. For example, even for someone as prolific as Evan, he must decide between feature depth (e.g. bundle splitting) and feature breadth (e.g. implementing the equivalent to @babel/preset-flow, which is used by both flow and hegel[0] type systems)[1]
Esbuild supporting Typescript and JSX is undoubtedly a byproduct of these grammar extensions having become popular. But supporting extra grammar extensions does add to complexity, sometimes in non-trivial ways. In Typescript, for example, you can import types using runtime import syntax, making it ambiguous whether running side-effects from a library is intentional or not. This gets problematic once you consider treeshaking and not-really-standard things like conditional resolution via package.json's `browser` field.
It gets even more fun when you realize that module resolution isn't even specified, meaning that as far as a bundler is concerned, something like `import 'lodash'` means completely different things for a browser vs a project installed via npm vs one installed via yarn v2...
Ugh, why does it have to be different between npm and yarn v2?
I'm an occasional JavaScripter and yetserday wrote an automation with Puppeteer.
The first line I wrote was
import puppeteer from 'puppeteer'
only for Node to fail on that. I thought that syntaxt was supported, but I had to change it to const puppeteer = require('puppeteer');
I've no idea if it's a problem with my Node version (12) or what, but it was a real "Huh?" moment.It's because Node was invented before EcmaScript modules existed. CommonJS worked well and supporting both was awkward, so the transition took a while.
- https://nodejs.org/api/esm.html
- https://nodejs.medium.com/node-js-version-14-available-now-8...
The easiest way is via the "type": "module" flag in the package.json file for your project.
Like the dumb broken Maps and Sets?
CommonJS was easy to implement anywhere and dead simple. We only need static analyzers and tree-shaking optimizations because the JS ecosystem went wild on dependencies. I honestly think we were better off 10 years ago.
> NodeJS never attempted to be compatible with the Web
It goes the other way around? Node caved in and support modules now, what has been done in tc-39 for web modules to be compatible with node (considering it predates all of it)?
There is no “JS community vs Node community” btw.
Mainly just because yarn v2 went off book[0]. npm vs yarn v1 is a more sane wold.
[0] A perspective on Yarn v2 from the creator of Yarn v1: https://twitter.com/sebmck/status/1300664946645069830
SWC, another competitor project written in rust is also handled exclusively by a single person. Half a million lines of code and they have even built a type checker for typescript which is not included.
Also one of the things I love about Figma is their obsession on performance. It feels such a joy to use. Runs circles around any Adobe product. It’s written in a browser. Figma and vscode are my inspiration for browser based tools.
My guess is an alignment between purpose, fun, and skillset. I would say (not for the sake of bragging but for qualifying the legitimacy of my position) - I probably pull similar hours across CTO-ing, investing, and learning biochemistry. Then again, I'm commenting on HN and Evan isn't, so maybe not. :P
1. Regarding bundle size: Optimization level used (SIMPLE optimizations yield output that is at least expected from a minifier these days)
2. Regarding compilation speed:
* Whether the JavaScript-only compiler was used over the JVM compiler: besides being slower, a number of optimizations are not available to the former.
* In many conditions a pre-heated JVM (or Nailgun), or using the NPM-distributed native compiler binary yields faster compilation.
* Which compiler flags have been tuned; e.g. such that the parsing of browser externs is bypassed during compilation for Node.js-targeted bundles.
Relevant discussion: https://github.com/evanw/esbuild/issues/425
This project is a prototypical example of the power of open source. An incredible contribution to the JavaScript ecosystem.
You also don't need a transpiler like Babel if you drop support for IE11. Almost all ES6 features– such as classes, proxies, arrow functions, and async are supported by modern browsers and their older versions. As long as you avoid features like static, private, and ??=, you get instant compile times.
It's not 2015. You don't need a transpiler to write modern JavaScript. As of writing, the latest version of Safari is 14 and the latest version of Chrome is 88. These features fully work in Safari 13 and Chrome 70, which have less than 1% of marketshare.
In development, use an index.html that loads it directly from the filesystem, using a single-pass annotation remover such as https://github.com/alangpierce/sucrase A single-pass transpiler doesn't build an AST, instead just looping over everything once, which will surely make it faster than esbuild.
In production, use a bundler like snowpack.
And saying it quite a lot in this thread.
1: https://www.typescriptlang.org/docs/handbook/intro-to-js-ts....
2. How do you handle the performance impact of your browser doing multiple sequantial requests because it doesn't know your whole dependecy tree?
See my comment above.
> 2. How do you handle the performance impact of your browser doing multiple sequantial requests because it doesn't know your whole dependecy tree?
Put all your dependencies in the HTML file, so it fetches them all in parallel. So instead of just loading the root application file and forcing the browser to resolve dependencies:
<script type="module" src="/scripts/app.js" defer></script>
Load the dependencies like this: <script type="module" src="/scripts/lib/gui.js" defer></script>
<script type="module" src="/scripts/lib/guielements.js" defer></script>
<script type="module" src="/scripts/lib/network.js" defer></script>
<script type="module" src="/scripts/lib/input.js" defer></script>
<script type="module" src="/scripts/lib/rendering.js" defer></script>
<script type="module" src="/scripts/lib/sockets.js" defer></script>
<script type="module" src="/scripts/lib/world.js" defer></script>
<script type="module" src="/scripts/lib/ui/classupgrades.js" defer></script>
<script type="module" src="/scripts/lib/ui/toasts.js" defer></script>
<script type="module" src="/scripts/lib/netstore/minimap.js" defer></script>
<script type="module" src="/scripts/data/configuration.js" defer></script>
<script type="module" src="/scripts/app.js" defer></script>
You can see a demonstration of this technique on my production site, a complex JavaScript SPA that loads everything (including dynamic content) in less than one second: http://vnav.ioFor a small performance boost, you can load smaller files first via <link rel=preload>.
I'm on Ubuntu 20.04 using the latest Firefox, if that helps.
EDIT: Seems to load fine in Chromium.
Also, there are tools that only bundle npm deps, like Snowpack or Vite. I think those would be perfect for your use case.
> Put all your dependencies in the HTML file, so it fetches them all in parallel.
> ...
> You can see a demonstration of this technique on my production site, a complex JavaScript SPA that loads everything (including dynamic content) in less than one second
But it doesn't work like this in practice! There is a limit to the number of concurrent connections a browser will make, Chrome for example is <=10 IIRC. You can see this in the waterfall:
`rendering.js`, the last `.js` in the queue, stalled for 300ms.
Additionally, each round trip is dependent on:
- The user's latency
- Your server's response time
So with each connection there is overhead. You would be much better served concatenating these files.
`util.js` took 510ms, of which 288ms was spent stalled (i.e. waiting for a connection) after which spent 220ms waiting (time-to-first-byte).
Furthermore, Lighthouse gives your page a performance score of 60, which isn't great. Key metrics:
- 1.7 seconds to paint (which is okay)
- 7.7 seconds to interactive (which is terrible)
Finally, why on earth are you redirecting to HTTP after serving an HTTPS response? This makes your page load even slower:
if (window.location.protocol === "https:") {
window.location.href = 'http:' + window.location.href.slice().slice(6);
}In HTTP/2, which our site will switch to soon, the concurrent limit is 100. The world's changed, change with it, and throw away those bundlers.
Either way, you can compare to our direct competitors: https://diep.io https://arras.io
> Finally, why on earth are you redirecting to HTTP after serving an HTTPS response?
This isn't really related to the argument at hand. My app is currently on HTTP for reasons downstream.
I suggest using a single-pass transpiler like Surcase[0] that doesn't translate to IE11-compatible syntax, and just loops through the string once, avoiding generating an AST– making it much faster– removing TypeScript annotations and desugaring JSX.
If you additionally need to support IE11, you can use a development build with Surcase and a production build with a bundler like Snowpack. C/C++ developers have been doing things like this for ages: compiling files as objects during development, and compiling them into one big binary for production.
When I do, I use a script tag, and put it before my ESM modules:
<script src="/scripts/data/contractor.js"></script>
Use them via global variables.I generally avoid third-party JavaScript modules for performance and maintainability reasons. When I do use them, I'll frequently vendorize them (put them in my source tree) and make ESM modules for them.
Also your website is serving dozens of cascading files instead of a single minified one.
Feels like 2010 to me.
Bundlers aren't the "wheels" here. The real "wheels" here are ESM modules, which have been available in every modern browser for a while. Bundlers are just pre-wheel, prehistoric, stopgap technology from the 2010s that we should be moving away from.
And there's no problem serving multiple files in 2020 because we have HTTP2 multiplexing now. In fact, it's probably more efficient than using bundling, because you can cache much better. Minification is also virtually unnecessary with Brotli, and not minifying has the added benefit of making the debugging experience much better.
And bundlers haven't been "perfected". Not even close. Webpack is without a doubt the worst piece of technology in my stack right now, together with Babel. Those two are terrible by themselves, but they also manage to "infect" other things elsewhere in my stack: for example, ESLint and Jest need Webpack/Babel plugins to work.
The wheels I’m talking about are the actual modules that exist on npm used in production by millions daily. Everyone here can write some code that does something common like, say, left-pad, and then they screw it in some obvious way that was fixed in 2014 by the third user of that library.
Sure I agree that having no build is nice, but user performance is not comparable and you’re limited to your own code.
In practice compression is not a big issue if you use something with a pre-defined dictionary specialized for HTML/CSS/JS, which is Brotli. The advantages of bundling are not exactly too large. Sure you might not be deduplicating some identifiers, but the lion's share of code in practice is normal javascript/CSS/html.
Also, in the long term, caching and the removal of code added by bundling more than makes up for any losses from compression.
Also, with Multiplexing there's no problem with download multiple files sequentially, because you don't need TCP handshakes and HTTPS negotiation, which were the reason for bundling in 2010. In fact, multiplexing multiple files might be faster in some cases, because you're able to parallelise the downloads and the parsing, rather than having to download a big bundle before parsing it.
There are of course caveats to this, but neither bundling or separate assets are silver bullets.
I love using plain ESM and I prefer to pull in stuff that only relies on that. But when I need graphing on a single page I'm not going to try to vendor Vega and I'm not going to pull it in on all pages. When I need xls parsing I don't want to vendor xls.js and maintain the diff myself.
So I have to use something like snowpack to make it work with my ESM system and dynamic imports.
If you don't have to require heavy libraries for certain pages then that's great and I encourage to use pure ESM without build systems for that, but you should also recognize that not all use-cases are ready for it.
Currently I use snowpack to handle dependencies but still run all my own code unbundled and unminified.
Add a <script src=> tag.
Access it by object destructuring a global variable. You can access ES5 stuff from ES6, just let ES5 run first.
That position isn't without good reason. Refer, for example, to the linked post, explaining the engineering decisions behind esbuild. ("It's Golang, not JS" is only part of the answer to "Why is esbuild fast?" The implications seem to be lost on many of esbuild's users, given the nature of esbuild itself and the job it's supposed to do.)
I mean there's a lot of garbage out there as with any ecosystem but saying "we all agree the JS ecosystem sucks" reeks of immaturity.
I love that TypeScript and esbuild both have 0 dependencies. For the first time I can have a build pipeline without transient dependencies.
I really dislike this line of thought: Babel isn't just a polyfill tool. It's an incredible way to enhance your code. Further, the future of javascript changes isn't static and shouldn't be.
[1]: https://github.com/evanw/esbuild/blob/master/docs/architectu...
It's good that this tool is shaking up the JS development ecosystem. It was getting stagnant (babel) and overly complicated (webpack). We can probably expect more tooling to migrate to golang and rust. I think the esbuild author accomplished his goal.
Its two biggest shortcomings are:
1. It can't compile const enums (TypeScript feature)
2. It can't compile down to ES5 (with some exceptions)
So I still need to run TypeScript + esbuild side-by-side. TypeScript compiles to ES5 and esbuild bundles the ES5 files.
Anyway, it is a massive improvement over TypeScript + webpack.
https://github.com/evanw/esbuild/issues/253#issuecomment-773...
They're also very helpful for debugging. Using await something in the console is way more convenient than something.then().
[1] https://developer.mozilla.org/en-US/docs/Glossary/IIFE
[2] https://github.com/tc39/proposal-top-level-await/blob/HEAD/R...
And good job on finding those top level await spec bugs.
For tsc it's a single line in package.json (+tsconfig with outdir: "dist").
For esbuild I am running the following Node script (saved as .mjs):
import esbuild from 'esbuild';
let args = process.argv.slice(2);
await esbuild.build({
watch: args.includes('-w'),
entryPoints: [
'dist/entry1.js',
'dist/entry2.js',
'dist/entry3.js',
],
bundle: true,
target: 'es5',
outdir: 'bundles/',
});Little by little I take large third party libraries and rewrite them on my own to drop 5MB includes. I'm sure that is part of it. Devextreme and Syncfusion are libraries I would strongly recommend avoiding.
As for Angular CLI which relies heavily on web pack, it has years of grinding tweaking bug fixing work getting all the right behaviors for all the cases. It's going to take a lot to re-create that in a non-breaking way using something different under the hood.
I first learned of Esbuild early 2020 when I complained on Twitter that JS build tools are awfully slow [1]. Most my (mainly Vue) projects use Webpack and starting the dev server takes typically anywhere between 5-50 seconds.
With Vite, cold devserver start in a mid-sized project on my 2015 MBP is ~2-5 seconds, but vast majority of the starts happen almost instantly as the deps are cached. Hot module reloading and most changes is also very fast and feel instantaneous (where in Webpack projects they would take anywhere between 1-20 seconds). This has HUGE affect in developing experience as you can just keep writing and almost never need to wait for any changes. (The one exception being TailWind/PostCSS builds that take few seconds every time the config or main import changes.)
I absolutely love Esbuild and Vite. Been using them daily for few months now, and I'm currently in the process of converting all my old vue-cli based projects to Vite using a project template I made that has TypeScript, Tailwind, and e2e tests w/ GitLab and GitHub CI configured [2]. If you work with React or Vue and Webpack and haven't yet tested out Vite or any of the other Esbuild based build tools, definitely do check them out!
[1]: https://twitter.com/uninen/status/1230673711910576130 [2]: https://github.com/Uninen/vite-ts-tailwind-starter
1. You can support IE11 without Babel (I got rid of it many years ago already): Typescript has transpilation, and you can add polyfills
2. As much as I like JSDoc-augmented vanilla JS, it doesn't catch nearly as much as real TS
3. Webapps are made of more than just scripts, loaders are even the big reason why Webpack got popular initially. Luckily, Webpack 5 can automatically recognize assets referenced in "new URL()" calls, so you don't need types for images or shaders anymore.
I'd like to be able to transpile my code but keep the same directory structure in the output. For example, I have an index.js file that exports all my components and I'd like to just create a "build" folder output with the same exact file structure, but with all the files exported from index.js having been transpiled.
Does anyone know if that is supported?
Personally, I wrote a crude JS script using the esbuild transform API + node-watch to compile .ts files and copy files/folders during development. This is all very recent, so I can't ensure that it's bulletproof quite yet.
[0] https://swc.rs/
For my site I basically only do minifying for js, and bundling+minifying for css, and it all gets done in 0.2 seconds. Right now the longest time in my deploy process is actually just doing a git fetch (with depth=1).
Not necessarily so: https://v8.dev/blog/code-caching
That would be an interesting comparison of performance and features.
You also have the editor finding issues for you, which is where I want that caught.
[1] https://github.com/TypeStrong/fork-ts-checker-webpack-plugin...