fs.promises.readFile is 40% slower than fs.readFile
github.com
github.com
,,But no one said fs.promises.readFile should perform similarly to fs.readFile ''
I'm happy that it was downvoted at least, but there should be a better way to set expectations for excellence in development.
I think as a general rule it is good to assume the best, even though I myself might not always be able to do so because of various limitations of understanding.
I think his point is valid even if slightly tongue in cheek. Bugs are defined by the contract / specification not wishful thinking.
Do you categorize thinking that JSON parsing takes at most O(n log n) time as wishful thinking as well?
It's humanly impossible to code for contract, that's why generally it's important to strive for no surprises in an API. This is the only way to build more and more complex software over time.
If this is how you do your work, you are asking for a computer to replace you! We should aspire to be better than this. It’s crazy to me that people take massive salaries and expect other people to spoon feed their tasks to them.
I remember being one of two software engineers working on a project in a small company. We were cooperating with two bigger companies delivering other parts of the system.
I kept asking myself: "These guys have hundreds of well paid engineers working on this, and they never deliver! Are they sabotaging the project?".
In the end we delivered the finished prototype and the customer ran a 48h acceptance test, their parts fizzled out almost immediately, while ours just kept running.
It took me a few years and joining a bigger company to notice that thinking about what you're doing is not expected in bigger corporations, where everything is bureaucratized. Also if the company is big enough it can shift the deadlines set by its customers. It drove me near madness whenever a colleague said something "wasn't their job". I never heard such an attitude before, and these people were paid twice as much as in my previous company.
I still, to this day, hope that a lot of these people get replaced by computers. Defective programs can at least be improved.
“Not my job” people tend to be the most frustrating folks to work with. While I understand wanting to mainly do the work you were hired for, there are always going to be times where you have to roll up your sleeves and do something different.
I wish I was the kind of person who would thrive in that environment because it sounds nice, but I’m not wired that way.
"When I pass X in, I'm not getting anything back - no error, I can't see the logs. Can you assist?"
"Not my job - you figure it out. We can't be expected to know how everything works."
"Yes but... this is literally something your team wrote in the last 3 weeks to support the larger project we're all on, and I was told to make my code work with yours. Have you ever tested taking input? If so, what am I supposed to pass in to your system?"
"Not my job - we're not here to babysit you - go read the code."
Later... someone else...
"Why are you spending all the time trying to trace through this code? You just have to ask. You have to learn to talk to people."
I sent over screenshots of the previous 'conversation' and got nothing back.
Thankfully this doesn't happen very often, but I'd realized that this org had segmented some teams so poorly, and put such deadlines on, that the culture of half the teams was "do minimal work and throw it over to someone else". It was more confusing because I interacted between a few internal teams, and some teams literally did not believe other people behaved this way. Even with written evidence, the problem was always assumed to be "oh, you're just not asking the correct way".
I’m not sure where in my comment you got the idea that I was advocating for engineers to have their tasks spoon fed to them by other people. My comment is simply about reading comprehension and not arguing against strawmen.
Don't take this for defending him, but it's in the realm of possibilities ¯\_(ツ)_/¯
The programmers didn't bother to free() any of their memory - instead they just calculated the maximum amount of RAM their program could consume while the rocket was in the air, doubled it, then installed that much ram in the rockets.
The author owes me a new keyboard.
Am I the only one who is terrified by reading an "amusing story" about "on-board software for a missile"? The amusement sort of requires us to completely ignore the individuals being blown to pieces.
Presumably, the RAM was a rounding error in the total budget.
And this move eliminated an entire class of bugs (use after free).
edit: And, you don't have to worry about memory fragmentation, or how long free() might take in the worst case.
I think a better approach (and one which NASA uses from what I've read) is to only use static memory.
In 1996, Ariane 4's successor, Ariane 5, took its maiden voyage to carry the Cluster constellation into orbit, when this happened https://www.youtube.com/watch?v=gp_D8r-2hwk&t=54s
Ariane 5 used a different launch profile. The impossible was now possible and the 16 bit value overflowed. The inertial reference system failed, along with its backup (which was running the same code), causing the rocket to receive wrong data.
At the same time, having added big uppercase warnings in the Ariane 4 source code and architecture docs, might have prevented this.
So, important, in safety critical code, to assume that others might blindly copy paste in bad ways -- and then try to warn them, although you might not be there any longer
The new one did not have these interlocks, instead doing everything in software. The result of this was a few people being fatally irradiated when a certain race condition was triggered.
It's not going to malloc randomly...
In fact, it's common in real time applications to pre-allocate everything you need at the start of the program or use static memory, and never allocate afterwards.
Then again, 206 downvoters didn't seem to think so (unless their reason for downvoting was objection to such a low form of wit).
Now when I think about it
> such a low form of wit
I think it's fun :-)
const readFilePromise = (...args) => new Promise((resolve, reject) => readFile(...args, (err, result) => err ? reject(err) : resolve(result)))
How can the native implementation be much more complex than that, beyond maybe some argument assertions? That will take almost the exact same amount of time as readFile, the additional time cost is a small constant value.I’m guessing as disk speeds have increased the relative overhead of dispatching the task to a thread pool has also increased.
At least, that’s how I think it works based on some half-remembered blog post.
Edit: yep - http://docs.libuv.org/en/v1.x/design.html#file-i-o
You’d need to benchmark another non-io bound function to find how much overhead comes from promises themselves.
"My guess is that the promise version is idling when IO capacity is available due to scheduling? Is it implemented in an event driven way internally, and if it is (for example using epoll) is the epoll event able to interrupt and cause the promise to continue to be evaluated or does it have to wait for the vm to get around to scheduling it again?"
It didn't occur to me that the async promise is rescheduled. That right there is a concern.
The OP also points out buffer concatenation, which is an odd choice. Reading the file into memory in one go should start with an stat and then a single buffer alloc. But that is probably naive.
Looking at you, svn team.
I think this is where fs.promises.readFile reads the actual file https://github.com/nodejs/node/blob/master/lib/internal/fs/p..., with many layers of indirection before that
fs.readFile is this one function: https://github.com/nodejs/node/blob/master/lib/fs.js#L327-L3...
The 'one function' you linked to just passes a ReadFileContext into binding.open[1]. The ReadFileContext appears to have its own read that uses its own allocated Buffers[2] under some circumstances, but under others they are mysteriously already there[3]. Then binding.open refers to an internalBinding of fs[4], which I'm unsure of, is that js or C++ code? This[5] suggests C++ is involved, but there is also an open in lib/fs.js[6] with that signature, though it also just calls binding.open at the end with the same signature, so I think it's off into C++ by then. At that point, I lose the trail. Maybe C++ is even doing the allocating.
This is not easy code to grok in its full extent.
1 - https://github.com/nodejs/node/blob/master/lib/fs.js#L356
2 - https://github.com/nodejs/node/blob/master/lib/internal/fs/r...
3 - https://github.com/nodejs/node/blob/master/lib/internal/fs/r...
4 - https://github.com/nodejs/node/blob/master/lib/fs.js#L74
5 - https://github.com/nodejs/node/blob/3b2863db12818778671713b5...
6 - https://github.com/nodejs/node/blob/master/lib/fs.js#L472
edit: I was curious so I kept reading, It seems like open is probably C++ here[7], which makes an AsyncCall or SyncCall to "open", which are just wrappers around syscalls[8]. Still not sure about other places the buffer comes from, though.
7 - https://github.com/nodejs/node/blob/master/src/node_file.cc#...
8 - https://github.com/nodejs/node/blob/606df7c4e79324b9725bfcfe...
fs.promises.readFile = util.promisify(fs.readFile);- The readFileSync in this benchmark has UTF-8 decoding (produces a string instead of a Buffer). That’s why it’s slower; not sure why it’s included.
- The first guess/suggestion,
> I suspect the cause is right here: [https://github.com/nodejs/node/blob/3b2863db12818778671713b5...]
> Instead of creating a new Buffer for each chunk, it could allocate a single Buffer and write to that buffer. I don't think Buffer.concat or temporary arrays are necessary.
is already what the code does. Multiple chunks and Buffer.concat are only involved if more data gets read than stat initially reports.
Do you have benchmark results to show this?
const readFileP = (...a) => new Promise((res,rej)=>readFile(...a, (e,r)=>e?rej(e):res(r)));
This is a one-line implementation and is probably way faster. I would totally just use this one-liner anyway.Results:
fs.readFileSync x 696 ops/sec ±1.96% (76 runs sampled)
fs.readFile x 2,045 ops/sec ±6.43% (72 runs sampled)
fs.promises.readFile x 435 ops/sec ±1.76% (78 runs sampled)
util.promisify(fs.readFile) x 2,053 ops/sec ±6.58% (67 runs sampled)
readFileP x 2,182 ops/sec ±4.00% (67 runs sampled)
fs.promises.readFile CACHED readFile x 438 ops/sec ±1.56% (79 runs sampled)
Fastest is readFileP,util.promisify(fs.readFile),fs.readFile
Slowest is fs.promises.readFile,fs.promises.readFile CACHED readFile
index.js: const Benchmark=require("benchmark"),fs=require("fs"),path=require("path"),chalk=require("chalk"),util=require("util"),promisifed=util.promisify(fs.readFile),bigFilepath=path.resolve(__dirname,"./big.file"),suite=new Benchmark.Suite("fs"),readFileP=(...e)=>new Promise((i,r)=>fs.readFile(...e,(e,l)=>{e?r(e):i(l)})),readFileC=fs.promises.readFile;suite.add("fs.readFileSync",e=>{fs.readFileSync(bigFilepath,"utf8"),e.resolve()},{defer:!0}).add("fs.readFile",e=>{fs.readFile(bigFilepath,(i,r)=>{e.resolve()})},{defer:!0}).add("fs.promises.readFile",e=>{fs.promises.readFile(bigFilepath).then(()=>{e.resolve()})},{defer:!0}).add("util.promisify(fs.readFile)",e=>{promisifed(bigFilepath).then(()=>{e.resolve()})},{defer:!0}).add("readFileP",e=>{readFileP(bigFilepath).then(()=>{e.resolve()})},{defer:!0}).add("fs.promises.readFile CACHED readFile",e=>{readFileC(bigFilepath).then(()=>{e.resolve()})},{defer:!0}).on("cycle",function(e){console.log(String(e.target))}).on("complete",function(){console.log("Fastest is "+chalk.green(this.filter("fastest").map("name"))),console.log("Slowest is "+chalk.red(this.filter("slowest").map("name")))}).run({defer:!0});
This little bit is my "implementation": readFileP=(...e)=>new Promise((i,r)=>fs.readFile(...e,(e,l)=>{e?r(e):i(l)}))
Run: npm install --save benchmark chalk && node ./index.js> readFileP x 2,182 ops/sec ±4.00% (67 runs sampled)
This means readFileP took between 2,094 and 2,270 ops/second.
> fs.readFile x 2,045 ops/sec ±6.43%
This means fs.readFile took between 1,913 and 2,176 ops/second.
So its possible it did run faster, but it's also possible it ran slower. You'd probably need to run it a lot more times to have an answer.
I do think the data here is enough to know that a promise wrapper here is insignificant in the grand scheme of things. This fs.promises library must do way more under the hood.