Fastcat – A Faster `cat` Implementation Using Splice
matthias-endler.de
matthias-endler.de
I know this is referring mostly to the `cat` portion and not the `splice` portion of the article, but I'll throw in a quick shoutout to `splice` for giving me one of the single biggest build performance wins in my time at Zynga (and possibly across most teams at the company at the time).
We had a ruby script which ran the majority of the build, and as the game grew we found that by far the slowest part was a loop which MD5 hashed each individual asset and used that as its filename on our CDN for per-asset-versioning.
At its worst it was taking nearly an hour and a half; the code was basically as inefficient as you could make it - multiple shell calls for each file rather than any sort of inlining of the hashing process.
I wrote a basic C program using splice and an MD5 library which took the whole process to under 10s. A bit overkill, perhaps, but the naive speedup I tried first still took over 1-2 minutes, and I figured 99.99% was worth the extra few hours to put it together knowing how many builds we ran each day.
Definitely gave me a healthy appreciation for the cost of transferring to user space that has stuck with me.
"Useless Use of Cat Award" [0] is the canonical text for avoiding unnecessary use of cat, for those who haven't come across it yet.
[0] http://porkmail.org/era/unix/award.html (2000)
cat file | this | that | other | out
vs
this < file | that | other | out
In the cat example, it's easy to change the head of the pipeline, by adding things before "this" or deleting "this", which is less so in the non-cat example. (The use case I have in mind is experimental commands that take probably <10s to complete, where editing time is a significant fraction of the time you spend.) < file this | that | other | outMaybe there are cases where a long string of awk|something|other|sort|uniq is not the problem, but forking an extra process for cat is.
And maybe there's a mismatch between pipes, files and mmap today. Splice seems like a reasonable fix (if we splice all the things, awk, grep etc).
Finally, I just think:
cat input.txt \
| filter1 args \
| filter2 args \
| reduction \
| ouput-formater
Reads better than having to tack on an <input.txt at the end, or special-case the first filter to be (... And also open a file). < input.txt filter1 args # ... rest of pipeline
(But I agree that the cat version is more readable.)I'd say that is somewhat of a harsh premise, especially since the in-place editing of files available e.g. in many GNU tools (awk, sed, sort) is really useful and based on exactly that possibility.
I do agree that cat often makes pipes easier to read, though. And yes, obsessing over that one additional process seems to be somewhat silly. Unless, of course, it introduces a real bottleneck and the whole thing is time sensitive.
With `cat`, a new process must be created. But with shell redirection, no new process is necessarily created, so that is going to be faster.
But aside from that it works just as the Rust version.
#define _GNU_SOURCE
#include <fcntl.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <unistd.h>
#define BUF_SIZE 16384
static void unset_flag(int fd, int flag) {
int flags = fcntl(fd, F_GETFL, 0);
flags &= ~flags;
fcntl(fd, F_SETFL, flags);
}
int main(int argc, char** argv) {
int pipefd[2];
pipe(pipefd);
unset_flag(STDOUT_FILENO, O_APPEND);
for (int i = 1; i < argc; ++i) {
int fd = strcmp(argv[i], "-") ? open(argv[i], O_RDONLY) : STDIN_FILENO;
if (fd < 0) {
fprintf(stderr, "%s: No such file or directory\n", argv[i]);
exit(1);
}
while (splice(fd, NULL, pipefd[1], NULL, BUF_SIZE, 0))
splice(pipefd[0], NULL, STDOUT_FILENO, NULL, BUF_SIZE, 0);
close(fd);
}
return 0;
}
WTFPL if anyone cares.With 32kiB buffers I get double the throughput than with 16k, the peak appears to be at 64k, after that it levels off.
https://rubygems.org/gems/io_splice/versions/4.4.0
http://www.bigfastblog.com/zero-copy-transfer-data-faster-in...
EDIT: source to current stable coreutils’ cat http://git.savannah.gnu.org/gitweb/?p=coreutils.git;a=blob_p...
Anybody knows if the Windows TransmitFile API can also be used to make file-to-file copies?
Interesting, but less than earthshaking.