Async Iterators: These Promises Are Killing My Performance
medium.com
medium.com
> The problem is that Promises have to be resolved in the next turn of the event loop. This is part of the spec.
That statement is false unless I'm missing something major here. Promises aren't resolved in the next turn of the event loop. They should be resolved as part of the microtask queue which can happen at the end of the tick (recursively).
Example (tested on node 8):
Promise.resolve()
.then(() => console.log('1'))
.then(() => console.log('2'))
.then(() => console.log('still the same tick!'));
setTimeout(() => console.log('timeout happening in the next tick'), 0);
Result: 1
2
still the same tick!
timeout happening in the next tick
The same will happen with async functions: (async () => {
console.log(await Promise.resolve('1'));
console.log(await Promise.resolve('2'));
console.log(await Promise.resolve('still the same tick!'));
})();
setTimeout(() => console.log('timeout happening in the next tick'), 0);This doesn't seem to follow to me. Synchronous reading means you have to read a given chunk in before you can process it, which means that if you're not reading out of cache that you're going to stall on the spinning disk, but synchronous code doesn't require you to read the whole file at once. And you can easily write something that will process it into nice lines for you. The "standard library" may not do that because there isn't much standard library for Javascript, but it's not that hard.
The proof is, that's what the Python snippet is doing. It's built into Python in this case, and I would expect it does some magic internal stuff to accelerate this common case (in particular doing some scanning at the C level to see how long of a Python string to allocate once, instead of using Python-level string manipulation), but the API ought to be one you can do in Node, even synchronously, with generators.
I don't know much about the author, Dan, but this should be able to be written to be a LOT faster. Perhaps not as fast as the synchronous example (reading something into memory then processing it is always going to be very fast if you can do it) but faster than the code example they give.
(I assume there are 100s of npm packages that do this already)
Even if you have a synchronous file read function that's doing partial file reads, you can not use it inside a synchronous loop to process the returned data because it will hold up the entire Node.js event loop until the whole file has been processed.
You can use the synchronous file read function in a way that doesn't hold up the main event loop but to do that you re-introduce all the complexity the author was trying to avoid.
The ideal would be just as performant as the sync Python code, but effectively async. And his observation that async iterators by design can't do it without severe overhead for such tight inner loops is very good.
It's implicit that stalling the node.js event loop is not allowed, so a synchronous "line" generator or iterator would block the event loop when it has exhausted the data read so far. The node.js event loop cannnot suspend the current task and process other tasks while waiting for more data to be read.
Edit: The point of the author's post is that async iterators generate a Promise (and thus a yield to the event loop) for every element, even when the element is already available synchronously. Even if the generator has read three lines in the latest call to read(), it must still yield three promises of one line each (each involving one yield back to the event loop) before all three lines are processed.
Edit 2: That said, it's not hard to fix this, and the author actually comes close to the solution but doesn't realize it. They say the result of the async iterator in the spec should contain an array of values, but it's just as easy to write the "line" iterator to yield an array of lines as its values without changing the spec.
What he was proposing was a solution that looked like you were only returning one value per iteration (so none of the code doing the processing had to be aware there are arrays of values) but underneath the covers was actually being processed in batches on each run through the event loop.
It's actually a very elegant solution. The only code that would have to be aware that it's actually dealing with batches of values is the code producing those batches right at the top of the stack of calls.
what's the point using nodejs if you can't do a simple task such as reading a file? Use something else to begin with. NodeJS wasn't designed to be "stateless", no more than Javascript was. If you want "stateless" use clojure or Haskell.
var fs = require(‘fs’),
csv = require(‘csv-parser’);
var start = new Date().getTime();
fs.createReadStream(‘./Batting.csv’)
.pipe(csv())
.on(‘data’, function() {})
.on(‘end’, function() {
console.log(new Date().getTime() — start);
});
This takes about 400 ms on my machine to parse a 6.1 MB CSV (15.25 MB/s). I didn’t do anything useful with the data, but I did verify it exercises the parsing function.Clearly all the work/time is spent in the actual parsing and handling corner-cases in CSV trying to be correct over high performance/throughput.
This is your problem, TypeScript's Promises polyfill is really really slow. Switch TypeScript to emit ESNext and run it on Node 8 and you'll see a much much different situation.
[1] This applies to sockets and pipes (incl. stdin), not regular file I/O which usually ends up being off-loaded to a thread/process pool (if needed)
https://www.tedunangst.com/flak/post/accidentally-nonblockin...
There's one more thing you should be aware of here that can catch a lot of developers out. On most POSIX systems, file reads from disk do not return EAGAIN but will in fact block if the disk is too slow returning the data back. EAGAIN is only meaningful on things like pipes, sockets, etc. The only way I'm aware of to do disk access truely asynchronously is using the POSIX AIO interface.
This seems imprecise? There are multiple queues and the details seem to be implementation-specific [1]. The Promises/A+ specification doesn't seem to specify a particular implementation [2].
The reason this might matter is that if you use promises to do something CPU-intensive without doing I/O, it might be possible to starve the event queue, similar to what happens if you're just writing a loop.
On the other hand, the example in the article is waiting on reads, so they'll be delivered through the event queue. Can reading larger chunks at a time and buffering be used to speed that up?
[1] https://github.com/nodejs/node/issues/2736#issuecomment-1386...
[2] "This can be implemented with either a 'macro-task' mechanism such as setTimeout or setImmediate, or with a 'micro-task' mechanism such as MutationObserver or process.nextTick. Since the promise implementation is considered platform code, it may itself contain a task-scheduling queue or 'trampoline' in which the handlers are called."
console.log("First");
promise.then(_ => console.log("Third"));
console.log("Second");
This means that even if `promise` is fulfilled already the callback must be delayed at least until the next turn of the event loop.
This is fine, unless you're:
1. Modifying the DOM and hoping to avoid jankiness 2. Trying to work with popup windows, in a browser which deprioritizes setTimeout in all but the focused window.
I ended up implementing https://github.com/krakenjs/zalgo-promise to get around this. It intentionally "releases Zalgo" by allowing promises to be resolved/then'd synchronously, but my belief is this doesn't have to cause bugs if it's used consistently.
Invoking a Promise's then() callback synchronously is a violation of the spec. IIRC there are some tests to prevent that in the test suite.
[1]: https://github.com/petkaantonov/bluebird/blob/2c9f7a44/src/s...
Anyway it sounds like we should try to fix the spec.
From all I know, that statement is false. Promises aren't resolved in the next tick. They should be resolved as part of the microtask queue which can happen at the end of the tick (recursively).
Promise.resolve()
.then(() => console.log('1'))
.then(() => console.log('2'))
.then(() => console.log('still the same tick!'));
setTimeout(() => console.log('timeout happening in the next tick'), 0);
Result: 1
2
still the same tick!
timeout happening in the next tickThere's no reason the source of the promise can't implement a cancel() method, though.
I really want to compute things if and only if there are listeners.
The number of listeners can go up and down during the computation. When that number reaches zero, the computation needs to be stopped.
When the number becomes nonzero again, the computation needs to be restarted.
Those are the requirements, basically.
If you use perl regularly, then the $_ syntax stuff is super fast to type and do all kinds of cool stuff, but I find as a irregular scripter that I forget after 6 months or so. Python's monotonous syntax means stuff is generally still understandable even by non programmers.
[1]: https://nodejs.org/api/readline.html#readline_example_read_f...