Goodbye, Node.js Buffer
sindresorhus.com
sindresorhus.com
For a status update, for the last year or two the main blocker has been a conflict between a desire to have streaming support and a desire to keep the API small and simple. That's now resolved [1] by dropping streaming support, assuming I can demonstrate a reasonably efficient streaming implementation on top of the one-shot implementation, which won't be hard unless "reasonably efficient" means "with zero copies", in which case we'll need to keep arguing about it.
I've also been working on documenting [2] the differences between various base64 implementations in other languages and in JS libraries to ensure we have a decent picture of the landscape when designing this.
With luck, I hope to advance the proposal to stage 3 ("ready for implementations") within the next two meetings of TC39 - so either next month or January. Realistically it will probably take a little longer than that, and of course implementations take a while. But it's moving along.
[1] https://github.com/tc39/proposal-arraybuffer-base64/issues/1...
[2] https://gist.github.com/bakkot/16cae276209da91b652c2cb3f612a...
I had to convert a ReadableSteam to base64 recently, and I was shocked at how much boilerplate it required in 2023:
if (isReadableStream<Uint8Array>(value)) {
const chunks = [];
for await (const chunk of value) {
chunks.push(chunk);
}
if (chunks[0].byteLength) {
const length = chunks.reduce(
(agg, next) => agg + next.length, 0
);
// Make the same species TypedArray that ReadableStream gave us.
value = new (chunks[0].constructor as Uint8ArrayConstructor)(length);
for (let i = 0, offset = 0; i < chunks.length; i++) {
value.set(chunks[i], offset);
offset += chunks[i].length;
}
} else {
throw new Error(`Unrecognized readable stream type: ${ chunks[0].constructor.name }`);
}
}
return `data:application/octet-stream;base64,${
btoa(
(value as Uint8Array).reduce(
(result, charCode) => result + String.fromCharCode(charCode),
''
)
)
}`;
I had to learn/implement a lot of little details to simply change the encoding from binary to its most common text representation. That's a day I would have rather spent doing product work. return `data:application/octet-stream;base64,${(value as Uint8Array).toBase64()}`;
but you'll still have to do the work of reading the stream to a buffer yourself.There is a _very_ early stage (as in, it's literally just an idea one person had, which may never happen) proposal [1] to do zero-copy ArrayBuffer concatenation, which would further simplify this - once you'd collected the chunks you could `value = new Uint8Array(ArrayBuffer.of(chunks.map(chunk => chunk.buffer))` instead of manual concatenation.
Finally, there's the Array.fromAsync proposal [2] and/or async iterator helpers proposal [3] (which I am also working on), which would make it easier to collect the chunks. Putting these together, you'd get something like
if (isReadableStream<Uint8Array>(value)) {
const chunks = await Array.fromAsync(value);
if (chunks[0].byteLength) {
value = new Uint8Array(ArrayBuffer.of(...chunks.map(chunk => chunk.buffer)));
} else {
throw new Error(`Unrecognized readable stream type: ${ chunks[0].constructor.name }`);
}
}
return `data:application/octet-stream;base64,${(value as Uint8Array).toBase64()}`;
[1] https://github.com/jasnell/proposal-zero-copy-arraybuffer-li...Why does de-chunking a byte array need to be complicated:
new Uint8Array(ArrayBuffer.of(...chunks.map(chunk => chunk.buffer)))
esp when chunking is specified by the platform in ReadableStream?-----
You have made me realize I don't even know what the right venue is to vote on stuff. How should I signal to TC39 that e.g. Array.fromAsync is a good idea?
The thing that I'd actually want for your case is either a TransformStream for byte stream <-> base64 stream (which I expect will come eventually, once the simple case gets done; it's also easy in userland [1]), or something which would let you read the entire stream into a single Uint8Array or ArrayBuffer, which is a long-standing suggestion [2].
---
> Why does de-chunking a byte array need to be complicated
Keep in mind the concat proposal is _very_ early. If you think it would be useful to be able to concat Uint8Arrays and have that implicitly concatenate the underlying buffers, [3] is the place to open an issue.
---
> You have made me realize I don't even know what the right venue is to vote on stuff. How should I signal to TC39 that e.g. Array.fromAsync is a good idea?
Unfortunately, it's different places for different things. Streams are not TC39 at all; the right place for suggestions there is in the WHATWG streams repo [4]. Usually there's already an existing issue and you can add your use case as a comment in the relevant issue. TC39 proposals all have their own Github repositories, and you can open a new issue with your use case.
Concrete use cases are much more helpful than just "this is a good idea". Though `fromAsync` in particular everyone agrees is good, and it mostly just needs implementations, which are ongoing; see e.g. [5]. If you _really_ want to advance a stage 3 proposal, you can contribute a PR to Chrome or Firefox with an implementation - but for nontrivial proposals that's usually hard. For TC39 in particular, use cases are only really valuable pre-stage-3 proposals.
[1] https://github.com/lucacasonato/base64_streams/blob/7c4ed815...
[2] https://github.com/whatwg/streams/issues/1019
[3] https://github.com/jasnell/proposal-zero-copy-arraybuffer-li...
If it's faster than node and the API to do what I need is better (Bun.FileSystemRouter comes to mind), I don't see why not to switch over.
Heck, I would accept even less retro-compatibility to get rid of transpiling, that completely oscene political mess they did with modules and treating TS as a 2nd class citizen.
The amount of time I'm still hit by ESM vs CJS is insane. Truly a Python 2vs3 moment for the node.js community.
The wiki article for it in the context of software (https://en.wikipedia.org/wiki/Proprietary_software) seems to pretty much stand it in opposition to open licenses, but I also have a sort of gut feeling that it could also refer to interfaces in OSS that are not designed to re-implementable. Not sure about that though.
XBL, a precursor to Web Components in use in Firefox for 15+ years before any ordinary webdev was talking about shadow DOM, was open source in every sense, for example. It was also proprietary.
See also: "Problems with XUL".
> XUL is a proprietary technology developed by Mozilla and only used by Mozilla.
<https://mozilla.github.io/firefox-browser-architecture/text/...>
The documentation of it, and the jump in complexity when it came to XPCOM (which were written in C++), were the reason (in my opinion) why platforms like Electron got popular instead of XUL.
In the context of not software licensing, I would define it roughly like you said. Interfaces that are designed in a way that is not compatible with existing software or easily implemented by others.
An example would be QT, the GUI framework for C++. It has it's own implementation of stuff that already exists in the standard library, like the string container for example (std::string). You can't use standard C++ types with QT and you can't use QT types with non-QT C++ libraries or types.
I am not a C++ or QT expert so there may be some level of compatibility between the types when using generics or in general, but from my very tiny use of QT it seemed like you had to use their types.
The problem, with Bun but really with the ecosystem at large, is that shipping stuff is (still) the de facto way that standards kind of congeal into something resembling actual standards.
it's always vendor lock-in. can never trust companies, mate!
Was very confused on what to use!
- File: Like a Blob, but with additional file-specific properties (e.g., filename).
- ArrayBuffer: Fixed-length raw binary data in JavaScript, not directly accessible.
- Uint8Array: Interface for reading/writing binary data in ArrayBuffer, showing them as 8-bit unsigned integers.
- Buffer: Readable/writable raw binary data container in Node.js (subclass of Uint8Array)
Buffer is a subclass of Uint8Array.
High level languages will still want to do low level things
You can see it in action on this WebGL fluid simulator[0] by PavelDoGreat.
So a list of floats? No! Let's not use floats everywhere just because JavaScript does.
This is what I feel the node people get right over the ES people, the ES ideology is so abstracted and pushes everything out into small utility classes so you have to create a handful of different objects with odd combinations of methods to get one useful conversion done.
The ES people also seem to have no love of the CLI or for type types of debugging and testing done there. As a result, I almost always choose the node created abstractions over the ES specified ones.
Does JavaScript's security model let you effectively sandbox scripts running in the same context from each other? If not, then why does this matter?
Not a security expert etc.
If the buffer doesn't have any cross-contamination with global state, there's no way one user could access another's data (because it's behind an object reference that never comes into the scope of the request logic for the other user). But if it did, and a malicious user found some other kind of vulnerability, they could potentially access data across-scopes
There's something other at play:
> buffers expose private information through global variables, a potential security risk.
This links to following piece of code [1]:
> // Somewhere in your code
> const privateBuf = Buffer.from(privateKey, 'hex');
> // Rogue package can access
> Buffer.from('1').buffer
I've just run it in node, and my god am I shocked!
> const privateBuf = Buffer.from('DEADBEEF', 'hex');
> Buffer.from('1').buffer
> ArrayBuffer { [Uint8Contents]: <2f 00 00 00 00 00 00 00 de ad be ef 00 > ....
[1] https://github.com/nodejs/node/issues/41588#issuecomment-101...
See the whole section on converting arbitrary binary data and the complex ways to do it.
I also provide a package to make the transition easier: https://github.com/sindresorhus/uint8array-extras (Feel free to copy-paste the code if you don't want another dependency)
It's insane to me that something as simple as concatenating an array needs a library, but as I've shown upthread, Uint8Arrays are way too complicated to work with.
https://github.com/sindresorhus/uint8array-extras/blob/cbf24...
[0] https://developer.mozilla.org/en-US/docs/Web/JavaScript/Refe...
Feel free to copy-paste the function to your own code base if you don't want the dependency:
``` const objectToString = Object.prototype.toString;
export function isUint8Array(value) { return value && objectToString.call(value) === '[object Uint8Array]'; } ```
Is that just another way of saying `Uint8Array.prototype.isPrototypeOf` and `instanceof Uint8Array` are not available in all JS environments?
I guess what I'm asking is the definition of a "Javascript Realm" in case I'm thinking it's something different.
Examples of this are frames in the browser and the `vm` module in Node.js.
This isn't meant as a personal attack on anyone, but we really need to frown upon needless dependencies, especially given the growing number of malicious NPM packages.
I suspect your philosophies are irreconcilable.
That sounds like a bug in those implementations.
I'm investigating https://arrow.apache.org/docs/js/ https://github.com/vega/falcon https://github.com/pola-rs/nodejs-polars https://github.com/uwdata/arquero
As for the fastest way to serialize data to Pandas data to the browser, you should use Parquet; it's the fastest to write on the Python side and read on the JS side, while also being compressed. See https://github.com/kylebarron/parquet-wasm (full disclosure, I wrote this)
I wonder what unmaintaned libraries will have to be dropped this time as the ecosystem grinds on and causes them to break?
(How did we get to a place where everything is this bad?)
Not vetting the bedrock you chose to build upon; deciding to rely on the unreliable.
("Everything" is an overstatement.)
You call a function that returns a buffer. The interfaces changes and you now get a UInt8Array.
You will never know.
It's one of those names where I tend to raise an eyebrow when I hear it. This looks like something similar to the module/require changes, where the overall change is good, but a slower transition may help people more.
I do have to wonder how many of those projects are essentially a wrapper around some native functionality just with their preferred API. Regardless, that level of scale is impressive.
Thanks for the heads up though, I’ll have approach with caution.
Like, I would think that replaces a lot of Buffer usage....but the articles mentions it exactly zero times.
This ideological pedantic purism which eliminates practical use cases is in my book just that - impractical pedantic purism. It’s easy to advocate for when your job is shaping landscapes without having to walk these.
And yes, I’m writing this still having flashbacks from ESM-only “transition”, which was more like throwing everyone into freezing waters with the promise it will warm up eventually.
No, it also (from the article):
> introduces numerous methods that are not available in other JavaScript environments. Consequently, code leveraging Buffer-specific methods needs polyfilling, preventing many valuable packages from being browser-compatible.
I think you are missing the point. Mutable and copied slices both have their use, and maybe a method that sometimes copies and sometimes shares does too. The problem with Buffer is that it overrides a method with well-defined semantics, then violates those. Buffer is a subclass of UInt8Array, but not a subtype nor Liskov-substitutable.
In practical terms, when you get a UInt8Array and even check that it in fact is a UInt8Array, you no longer know that .slice() does even though UInt8Array's documentation explains how every UInt8Array behaves, because that documentation is now wrong. If the UInt8Array you get is a Buffer, then it is both a buffer and a UInt8Array but its .slice() behaves differently than expected.
It looks like I could accomplish the same thing using a DataView of an ArrayBuffer, but I don't see enough of a benefit to justify converting everything to this approach.
I can see how the Buffer behavior here would be preferable in cases where performance is important or memory constrained environments.
But this is an implementation detail, not specified behavior. Changing method behavior in subclasses is a key aspect of inheritance.
Changing method behavior in subclasses is part of inheritance, but it shouldn't confuse or mislead. In the case of Buffer and Uint8Array, the altered `.slice()` functionality isn't a mere implementation detail; it's a significant deviation. This inconsistency can lead to unexpected bugs, especially for those who assume similar behavior based on the inheritance hierarchy. It's crucial for reliability that such fundamental behaviors remain predictable across subclasses.
[0] https://tc39.es/ecma262/multipage/indexed-collections.html#s...
Correct spec: https://tc39.es/ecma262/multipage/indexed-collections.html#s...
Steps 14.g.i to 14.g.ix detail the transfer of data from the original TypedArray (O) to the new TypedArray (A). It involves reading values from the original and writing them to the new array's buffer, effectively duplicating the data segment. The process ensures both arrays are distinct with separate memory spaces.