(I had actually hoped Futhark would be slower sequentially, just so this wouldn't be the focus of the discussion!)
(I had actually hoped Futhark would be slower sequentially, just so this wouldn't be the focus of the discussion!)
I used the "reference" source code linked from the original Haskell post, a BSD version hosted by Apple: https://opensource.apple.com/source/text_cmds/text_cmds-68/w...
It uses raw read() from a file descriptor and works with pipes as well. I think the only special handling for stdin vs. an actual file it has is calling fstat() if only the number of characters is requested, which shouldn't apply here.
So yes, this version does need to do more complicated I/O than a simple mmap(). And (broken record, but I'll stop after this) it's 2x as fast as my system's GNU wc (when compiled with -O3 vs. however the system wc was compiled).
> I had actually hoped Futhark would be slower sequentially
It might still turn out to be, if you see if you can get a faster C version of wc.
You definitely can, at least if you allow manually vectorized code.
On my system, with a 1.661GB file (256 times big.txt from the original Haskell post) GNU wc takes about 6.5s (real time), a stripped down version of Apple's implementation about 4.1s, and a single-threaded vectorized wc (written in C) only 0.27s. (These times are of course only with a hot cache. For reference, catting the same file to /dev/null takes about 0.18s.)
edit: corrected the time for the BSD-derived implementation
You're going to be intially page faulting every 4096 bytes if you mmap the file. The fact that you're accessing the mapped range sequentially in this case may help, I guess.