Node Modules at War: Why CommonJS and ES Modules Can’t Get Along
redfin.engineering
redfin.engineering
I’ll conclude with three guidelines for library authors to follow:
Provide a CJS version of your library
Provide a thin ESM wrapper for your CJS
Add an exports map to your package.json
</quote>
This seems like the same fallacy of python’s 2to3, and why they should have written 3to2 instead. If ESM is the new way, library authors should instead have a CJS shim to prevent fossilization of cruft.
The (comparatively) few Python 2 projects that were alive by the release of Python 3 and still alive by the EOL of Python 2 could probably have been upgraded manually, rather than with a tool.
On the other hand, people not wanting to write new code in Python 3 because it would be incompatible with the Python 2 ecosystem – that was a real problem. That problem could have been prevented with a 3to2 tool, which would compile the Python 3 code to Python 2, allowing it to be used with the existing ecosystem.
If only people would have written new code in Python 3 right away, we would have been in a far better position by the time Python 2 would have been EOL'd, purely due to natural turnaround of code bases.
Porting Python 2 code to Python 3 was never the problem. Writing new code in Python 3 was.
You're, frankly, high as a kite. As the author of one such port on a large codebase, porting Python 2 code to Python 3 was absolutely a problem.
The issue with 2to3 is that it assumed you'd do the conversion as a one-shot and straight flip over with entirely separate codebases (or dropping Python 2 entirely) from there on, which is completely insane.
What the vast majority of transitioning projects needed was a way to migrate progressively through cross-version codebases:
* libraries weren't going to drop older Python versions, the very few which attempted that (I'm aware of dateutil at least) were extremely painful to users and IIRC ended up rolling back to a more progressive codebase
* programs being distributed were not going to drop Python 2 until EOF at least
* large codebases simply could not afford to perform a completely untested and uncheckable one-shot transition
That's why the community created cross-version support tools like six and friends, and lobbied hard for the reintroduction of cross-version features into Python 3 e.g. reintroduction of `callable` (3.2), reintroduction of the "u" prefix (3.3), reintroduction of % on bytes (3.5) and I'm sure others.
> Again, there are exceptions, but the vast majority of websites or data processing scripts or whatnot are significantly re-written or decommissioned in – pulling a number out of my arse – five years or so. […] Python 3 because it would be incompatible with the Python 2 ecosystem
What do you think "the python 2 ecosystem is" except for a ton of projects which needed to be ported from Python 2 to Python 3 exactly? "the vast majority of websites or data processing scripts" is not an ecosystem, tooling and libraries are.
Do you somehow think numpy and lxml and django just scrapped their entire Python 2 codebase and rewrote the entire thing from scratch? Of course not.
> That problem could have been prevented with a 3to2 tool, which would compile the Python 3 code to Python 2, allowing it to be used with the existing ecosystem.
3to2 would not have worked any better better than 2to3. The syntactic incompatibilities between the version were not the hard part, the semantic ones were, and 3to2 would not have succeeded any better than 2to3 there because it simply couldn't.
All 3to2 would have given you is a broken Python 2 codebase from a Python 3 codebase you could not even run.
https://news.ycombinator.com/newsguidelines.html
"Be kind. Don't be snarky. "
So for over a decade, you'd be writing Python 3 code but building a Python 2 project. After a decade, one would hopefully have managed to port everything over to Python 3, and can then drop the "compile to Python 2" step.
A 3to2 tool would solve exactly the problems you're bringing up, with no intoxication needed!
----
I don't see why 3to2 would not have worked. It's a compiler. We have written compilers before. We can write compilers. Compilers work. Both languages are Turing complete and about the same in expressiveness, so clearly compiling one to the other is something we can do.
What we cannot (reasonably) do is produce nice, human-looking target code from a compiler. This is what 2to3 attempted to do, and what a 3to2 would never need to do.
Who cares if the Python 2 code it produced looked a little wonky, as long as it was correct and roughly recogniseable as being built from the original Python 3 code? The Python 3 code is the one humans write and edit. The Python 2 code is just compiler output to make it compatible with all of the Python 2 ecosystem until one is ready to go fully Python 3.
----
This is also not something completely foreign and unheard of. This is exactly how I've been porting legacy JavaScript applications to a more modern ECMAScript approach: write new code in the modern way, compile it down to legacy code compatible with the existing legacy code, and then run that.
Slowly but surely, more and more things are being written in the modern fashion, and eventually, when everything has caught up, the compatibility conversion to the legacy code can be dropped.
It doesn't. Not that its job is easy, but its job is to take constructs which are brand new and translate them to an older version of the language.
Of the same language.
It doesn't have to deal with the language or the APIs literally working differently when used the same way.
The Python version of Babel would be 3to3 compiler.
No. You'd have to migrate everything to Python 3 at once, and then "maintain" two different and incompatible codebases, one generated from the other.
> A 3to2 tool would solve exactly the problems you're bringing up, with no intoxication needed!
It would half solve (at best) one of them.
> I don't see why 3to2 would not have worked. It's a compiler.
Unless your 3to2 would literally reimplement Python 3 in Python 2, it would be a translator / bridge. Which can't work due to the semantics difference between Python 2 and Python 3: not every change translates mechanically and reliably. If they did, 2to3 would have worked.
> Who cares if the Python 2 code it produced looked a little wonky,
The maintainer who has to debug, probably.
> as long as it was correct
Which it wouldn't be.
> This is also not something completely foreign and unheard of. This is exactly how I've been porting legacy JavaScript applications to a more modern ECMAScript approach: write new code in the modern way, compile it down to legacy code compatible with the existing legacy code, and then run that.
The "modern way" of javascript is additive, it's about adding new features to the language, some of which can be implemented in terms of the old one (and which you can thus "backport" through a compiler or extending the existing APIs).
You don't run into pieces of code which are syntactically identical but semantically divergent. Which you absolutely do in P2 v P3.
Absolutely not. Well, you could, but the better approach would be to write individual modules in Python 3, and then compile those down to Python 2 at which point they can talk to the existing Python 2 code freely.
> Unless your 3to2 would literally reimplement Python 3 in Python 2
That's exactly what I'm suggesting. Take the existing Python 3 compiler and swap out the backend to generate Python 2 code instead of bytecode.
This is stuff people do every day for other languages. It is very far from rocket science.
the real issue is 2to3 didn't work so well, or exposed the exact issues why python 3's breaking changes were needed (80% sloppy bytes/str/unicode handling).
The end result is something that can be Py3 native whenever I want. 2to3 is the right tool for the job. There is no way to map all Py3 features back into Py2 so 3to2 would be a broken mess that only supported a subset of the language.
Oh but there is. If you can map them to bytecode, you can map them to Python 2, which is a fully capable language. It won't be a one-to-one map, but that's also not required.
Sure the generated code is a horrific mess to look at but even that's a solved problem with sourcemaps if you need to debug it.
This is incredibly insightful. I'm frankly a bit stunned. I would never have thought of this (and clearly the Python people also didn't) but it is absolutely true. It's also a very transferable strategy for any time compatibility is to be broken.
Is this an original thought of yours or do you have references to further reading?
OP wrote, note that it’s easy to write an ESM wrapper for CJS libraries, but it’s not possible to write a CJS wrapper for ESM libraries.
A library author can provide cjs bindings to their library for as long as cjs remains relevant with the assumption that they won't use top level await.
You can - and people widely do - compile a CJS version of your library from ESM sources. You'll accept that there may be multiple copies of your library in the program but it does work. But in any case there's no reason to actually author in CJS, everything including the ESM shim can be generated from ESM.
Last time I checked they seemed to lean into using rollup to take care of the translation under the covers.
From what I have seen seems pretty typical to simply export the named properties using a fresh module.exports = {...} and then use default as the main export?
CJS has a single export which might be an object. And objects have named properties/fields.
On the other hand, ESM CAN have unnamed "default" export and CAN have arbitrary number of named exports.
And this is a constant source of confusion and frustration for me when working with so many npm modules lately...
To nitpick on the nitpick: ESM only has named exports. It just treats the export named "default" differently in some scenarios and has special syntax for it.
So, roughly:
* CJS: Exports a single value which may or may not be an object. When assigning the exported value to local variables, destructuring can be used to access individual exports properties.
* ESM: Exports a namespace of named bindings. When importing the bindings to local aliases, the name of the alias can be changed using `as`. There's special syntax for importing the binding with the name "default".
It's somewhat unfortunate that the syntax _looks_ so close without being all that close.
I find it fascinating watching our understanding of async computation mature over the years.
[1] http://journal.stuffwithstuff.com/2015/02/01/what-color-is-y...
I cannot find the others though.
>async computation mature over the years
A good thing, really, but would be much better if the experience from few decades ago wasn't ignored by cool kids who mature.
Those benefits being that the async/await syntax and semantics for working with Promises more closely resembles the imperative code it replaces, whereas the join calculus is its own [sometimes easy, admittedly] learning curve with its own syntax and semantics. Most of the complaints about Promises are the usual complaints about monadic wrapping types that the types become "viral" and "color" all associated code. But that just means that the failure cases are more easily picked up by a static type checker than semantic failures in learning the join calculus.
function cpsfoo(cb) {...}
function foo() {
var thread = this_thread()
cpsfoo(x => thread.resume(x))
return thread.yield()
}
Error and immediate cb invokation handling omitted for clarity. Similar wrapper or wrapper-generator (uncps(f), unpromise(f)) may be done for other primitives.... but this is exactly the reason why callback hell existed. Error handling down the callback branches were a total mess, especially with networked state transfers.
This is an example for something that might come up in Javascript, using async/await:
const a = await fetchA();
const b = await fetchB(a);
const stuff = await fetchStuff(a, b);
const transformedStuff = await Promise.all(stuff.map(someAsyncTransform));
Error handling is just try/catch.Note that you generally have no control over most of these function being asynchronous. Threads or coroutines do nothing to help you write straightforward code when all your libraries are based on fine-grained callbacks/promises. Writing out that chain of dependencies in CPS would be a nightmare.
Node is built on V8, which is built for browsers. Browsers have a single threaded execution model that also hides an event loop that runs the entire tab.
CommonJS was a mistake that the JS world is going to live with for years more.
It's time we all published standard JS modules and only standard JS modules to npm.
(Author here.)
Node 14 supports ESM, but Node 12 only supports ESM behind the --experimental-modules flag, and its error handling for incorrectly importing CJS is not good.
In 2020, I think it's a bit early to go ESM only, especially since it's straightforward for CJS libraries to support both CJS and ESM clients. (I document how to do this in the article.)
IE had many non-standard APIs that were much more convenient to use than what the standard was providing (innerHTML, document.all)
And rather than moving to the standard, more and more code was written targeting IE which also had a conveniently large user-base.
ESM is the standard. CJS is the solution proprietary to node.js.
https://www.sitepoint.com/deno-module-system-a-beginners-gui...
Note that - if you can get an esm out of your cjs - and it doesn't use any nodejs apis unsupported by deno, then you might be able to use it:
> (...) Several CDNs can convert npm/CommonJS packages to ES2015 module URLs, including: Skypack.dev, jspm.org, unpkg.com (add a ?module querystring to a URL)
Ed: submitted as: https://news.ycombinator.com/item?id=24069037 since I couldn't find an earlier submission, and maybe it'll generate some interesting discussion on deno / modern js vs node.
In particular I wonder if it'll in practice be easier to share (more) code between backend and front-end.
Always has been.
import x from "library";
import {x} from "library";
import * as x from "library";
import {x as y} from "library";
Just to compare Java uses this: import java.lang.Double;
import java.lang.*;
No renaming, not a gazillion options during in- and export, but one statement. If you want to rename something, assign it to a variable.I don't understand why ES-module authors made it that complicated, when they had the chance to introduce a proper module system.
Why did they not just copy Java's system and be done with it.
I think they did it like that for compatibility with CommonJS.
Your last sentence is like a C dev asking a Java dev why they don’t just compile to x64 and be done with it. It comes across as flippant and will likely attract downvotes.
1. Import the default export as the name `x`.
2. Import named export `x` as the name `x`.
3. Import all of the named exports under an object named `x`.
4. Import named export `x` as the name `y`.
All of those the space of namespaces is global. You can't have 2 modules named math.quaterinions in any of those languages. You can in JavaScript. To put it another way, in order to prevent name clashes in C++/C#/Java you have to know the name of every other project on the planet. If someone uses the same namespace then if you ever need to use those 2 projects together one of them will have to be re-factored.
That problem doesn't exist in modern JavaScript. There are no namespaces. There is just floating references to code.
namespace foo {
class Bar ...
}
In another file declare namespace foo {
class Bar ...
}
Now use both foo.Bar from file 1 and foo.Bar from file 2 in the same project. Aliases don't help with thisDifferent assemblies with different assembly names but same namespace and class? Just use the fully qualified type name (and alias it for a shorter name).
Single assembly/project and namespace with the same class in multiple files? If you're just spreading out the code then use partial classes, otherwise why would you define the same type twice?
C# (and other languages) that use virtual namespaces can do everything that file-based namespaces can do, but also support many more scenarios.
There is one gazillion things that are much more a problem than "namespace conflicts" in packaging and package management.
Including security & code authenticity, parallel multi-version handling, sensibility to side effects, speed, auditing and reproducibility.
And in these domain the JavaScript packaging system generally range from being 'catastrophic' to 'absolutely terrible'.
If I have to take an example package management system for the world I would certainly quote cargo, Nix, Guix or Spack but certainly not npm which is quite close to what you should really not do for a package manager.
2. Get the exported value x
3. Get all exported values and call the object x
4. Get the exported value x and call it y
They're all different because they allow you to do different things, and fairly readable.
Edit: Just realised dimgl said the same 45min ago, my bad!
[1] https://developer.mozilla.org/en-US/docs/Web/JavaScript/Refe...
- no recursive destructuring
- `import {y} from 'x'` is not the same as `import x from 'x'; const {y} = x`
But I think your second point is more to do with the fact `import` by default imports `default` instead of the whole thing.
If you use import * as x from 'x'; then `const {y} = x;` works.
I personally will consider it is `import from` being different from a typical object instead of {x} here being different from typical destructuring assignment (by that I mean they still follow the same pattern; not necessarily say they're the same thing.)
import org.joda.time.Instant as JodaInstant;
The rest of the weirdness is to make it look like JS destructuring
That class doing that transformation will need to refer to both the Joda and java.time class.
https://www.npmjs.com/package/esm
I highly recommend it. Really let’s you unify on ESM modules without much thought
> "await":false - A boolean for top-level await in modules without ESM exports. (Node 10+)
So, yeah, seems to rely on "most ESM modules don't use top-level await". Seems like a useful option to have, thanks!
I guess it's all irrelevant when you don't use Node, aren't publishing libraries, and are compiling all JavaScript into bundles for deployment and never read the bundled code.
You pick one of these when you first create your app which is typically based on whatever framework or template you are starting from, and you just always stick with that.
So much of the complexity of webpack, parcel, rollup, pika, snowpack, et al is dealing with the weird mix of CJS and ESM spread across npm at this point.
One of the main differences between the bundlers at this point is their opinion of CJS and ESM handling. You may not notice it as a bundler user, but you probably see a lot of indirect results in things like bundle size and ease of use. So it is not entirely irrelevant to frontend work, but you just may not realize that why you prefer one bundler over another today is heavily influenced by their approach to the CJS / ESM split. (A good bundler you shouldn't have to realize that behavior is what contributes so much to the bundler's "feel" and your gut-reaction/instinct to working with it.) Learning the differences can be useful if you need to better pick your next bundler based on technical underpinnings, though admittedly that sort of "comparison shopping" is probably a rare need.
// cjs.js
module.exports = (a, b) => a + b
exports.default = 'default'
exports.nice = 69
// esm.js
import * as mod from './cjs'
// what is `mod` now?
In babel, `mod` will be `{ default: [Function], nice: 69 }`, and the string 'default' will be discarded because ESM always exports a `{ [key: string]: any }` while CJS can export whatevey they want (`any`).When converting TypeScript that using `import default xxx` to JS using `tsc`, it would convert it to something like `exports.default = xxx`.
To import it, you have to use `const XXX = require('./mo.js').default` instead of just `const XXX = require('./mo.js')` in plain ES6 JS.
This is fine but Babel used to [1] be able to convert `const XXX = require('./mo.js')` to valid code, which made lots of people wrote it (incorrectly) that way (I was very baffled why people wrote like that since it does not work in plain ES6).
Using the default prop doesn't work in tsc and breaks all type information because tsc just doesn't know of any default prop.
A TS file which compiles to an es6 ("esnext") module cannot import cjs. Either node breaks or TS' type-checking fails. Former if you omit default when importing and latter if you introduce .default in the TS codebase.
It's a bit complicated, what I just wrote about is reflected in the chart in step 4-6 here https://github.com/microsoft/TypeScript/issues/18442#issueco...
Thank you.
And dynamic imports are still complicated sometimes
There are high chances you are relying on some non-spec compliant behaviors but you don't know about it. For example, ESM exports have immutable bindings, which means, mocking exports for unit testing is impossible. If it works for you, it's because you transpile the code from ESM to CJS (either via `babel`, `rewire` or similar tool).
Unsure what this means. The only reason this happens is because an unfinished copy is provided to the dependency that is causing a circular reference.
I wonder if deno (maybe with a tsc reimplementation written in rust) will take its place for new projects. Personally, I've already suffered enough with python 2/3 and decided, after 9 years writing node, to use uniquely rust for new backend development.
Frontend development is quite a slow process as well, but all the faster alternatives imply knowing a different language - which impacts developers availability (given historically frontend developers are mostly JS).
My first instinct these days is how that funky new bit of Lego fits with all the other bits of Lego.
When you start there, it's becomes easy to understand just how incumbent C++ and javascript are, even with all their necessary spells and arcane witchcraft.
PHP and Perl are probably more recent examples comparable to JS here that provided mostly only low level plumbing (include/require) support and nearly wound up with multiple incompatible module loading patterns.
Even C/C++ the "module" system started as a text processing hack in a macro pre-processor that used to be a separate tool (though I don't think any modern C/C++ toolchain uses a separate macro pre-processor, as there are tons of optimization layers now and some of the macro pre-processor knowledge makes its way all the way through to the linker these days).
Historically, modules were always just "smash these two files together in this order" and it was only later that languages started asking hard questions such as "which symbols from this other file are available".
The modern concept of modules as a fixed "shape" of self-describing public symbols is as much a later abstraction out of Object-Oriented Programming as anything else. Our idea of what a module should be in 2020 is much more informed by "recent" notions such as Component systems like COM and languages such as Java and C#.
And for tools that want end to end esm, like vite. Won't it cause degradation in behaviour?
Correct. I wouldn't follow the article's advice wrapping ESM with CJS as it leads to larger bundles.
> Won't it cause degradation in behaviour?
For web apps - certainly. The forced sync behavior of CJS wouldn't be very noticeable for local applications however.
Given the widespread popularity of CJS they should have prioritized CJS backward compatibility over things like better support for circular dependencies and mutable exports. Many languages are doing quite fine without them.
Then everything just works. Top level await and all.
Everything used to work until Node 12.15, then 12.16 started throwing when trying to require type:module packages with `require` (which is what `esm` does)
> You can’t require() ESM scripts
That’s one way to put it! Another would be to say that you don’t have to. require() came about because we didn’t have the import keyword.
You were not using native ESM in Node years ago. Node 14 shipped with unflagged ESM support just this year. (If I had to take a guess, you were probably transpiling "import" statements into "require" statements and/or bundles, which honestly works better than native ESM, because there's no interop issue.)
Almost all of the most depended on NPM packages are CJS, and provide no ESM exports. https://www.npmjs.com/browse/depended
lodash, react, chalk, commander, moment, express, axios, and vue are all CJS.
Or if you prefer this Hessian article:
https://www.faz.net/podcasts/wie-erklaere-ich-s-meinem-kind/...
(Central Europe)
One of my earliest memories is of sitting on the open tailgate of my mom's truck, carefully holding a slice of pizza so as not to disturb the black-and-yellow mud dauber who had landed on it to nibble daintily at the edge of a bit of pepperoni. In the general run of human theory, I gather, such a moment should require all sorts of histrionics, but neither she nor I saw any need for them, and no one else was around to disturb her appetite or my fascination. On reflection, I suppose that moment must have done something to set a tone for my life in terms of my relationship with aculeates.
And it's not as if they eat enough for a human appetite to notice, anyway. They're hell on caterpillars - I've seen a few P. metricus foragers entirely depopulate a thriving fall webworm nest in the course of a couple of days - but, having never developed a taste of my own for such things, that really doesn't bother me. And I do very much enjoy watching them hunt!
Squirrels, on the other hand, can all go straight to hell.