Author here. I wasn't suggesting it would have any perf impact. Just that it was an interesting change set.
> profile of a sample word count program I was writing, which showed the program was spending way too much time in the syscall module. That in this context can only mean one thing: way too many read syscalls were getting called.
I find it hard to believe that the profile would look any different with 1 vs 2 syscalls per 2GB chunk. The syscall overhead is going to be insignificant compared to actually copying the data. The program is going to be spending a lot of time doing syscalls no matter how many there are, because the individual syscalls will just start to take more time as you increase the size of the buffer.
Edit:
Compare:
strace -e read -T perl -MFcntl -e 'sysopen FD, "foo", Fcntl::O_RDONLY; while (sysread FD, $buf, 1*1024*1024*1024) {} '
read(3, ..., 1073741824) = 1073741824 <0.455901>
read(3, ..., 1073741824) = 1073741824 <0.219711>
read(3, ..., 1073741824) = 1073741824 <0.213923>
read(3, ..., 1073741824) = 1073741824 <0.211783>
Vs. strace -e read -T perl -MFcntl -e 'sysopen FD, "foo", Fcntl::O_RDONLY; while (sysread FD, $buf, 2*1024*1024*1024) {} '
read(3, ..., 2147483648) = 2147479552 <0.921789>
read(3, ..., 2147483648) = 2147479552 <0.487007>
read(3, ..., 2147483648) = 8192 <0.000031> dd if=/dev/zero of=foo bs=1024 count=$((1024*1024*4))
You probably don't have a file of that name in the current working directory?