FWIW, when this was going around for the first time I took this Darwin version of wc and experimented with setting domulti to const 0, statically removing all paths where it might do wide character stuff. I didn't measure any performance difference to just running it unmodified.
if (iswspace(wch))
by if (wch == L' ' || wch == L'\n' || wch == L'\t' || wch == L'\v' || wch == L'\f')
And I get a ~1.7x speedup: $ time ./wc ../wiki-large.txt
854100 17794000 105322200 ../wiki-large.txt
./wc ../wiki-large.txt 0.47s user 0.02s system 99% cpu 0.490 total
time ./wc2 ../wiki-large.txt
854100 17794000 105322200 ../wiki-large.txt
./wc2 ../wiki-large.txt 0.28s user 0.01s system 99% cpu 0.293 total
Remove unnecessary branching introduced my multi-character handling [1]. This actually resembles the Go code pretty closely. We get a speedup of 1.8x.: $ time ./wc3 ../wiki-large.txt
854100 17794000 105322200 ../wiki-large.txt
./wc3 ../wiki-large.txt 0.25s user 0.01s system 99% cpu 0.267 total
If we take the second table from the article and divide the C result (5.56) by 1.8, the C performance would be ~3.09, which is faster than the Go version (3.72).Edit: for comparison, the Go version from the article:
$ time ./wcgo ../wiki-large.txt
854100 17794000 105322200 ../wiki-large.txt
./wcgo ../wiki-large.txt 0.32s user 0.02s system 100% cpu 0.333 total
So, when removing the multi-byte character white space handling, the C version is indeed faster than the (non-parallelized Go version).[1] https://gist.github.com/danieldk/f8cdaed4ba255fb2954ded50dd2...
if (iswspace(wch))
to something like if (domulti && iswspace(wch))
...
else if (!domulti && isspace(wch))
...
got something like a 10% speedup on my machine. And replacing isspace with an explicit condition like yours is much faster still. I checked, isspace is macro-expanded to a table lookup and a mask, but apparently that's still slower than your explicit check. I'm a bit surprised by this but won't investigate further at the moment.I am sorry for the unclear comments. I'll stop commenting on a phone ;).
Indeed, the code uses iswspace to test all characters, wide or normal. Strange design choice.
I agree, it's really strange. This seems to be inherited by the FreeBSD version, which still does that as well:
https://github.com/freebsd/freebsd/blob/8f9d69492c3da3a8c1ea...
It has the worst of both worlds: it incorrectly counts the number of words when there is non-ASCII whitespace (since mbrtowc is not used), but it pays the penalty of using iswspace. It's also not in correspondence with POSIX, which states:
The wc utility shall consider a word to be a non-zero-length string of characters delimited by white space.
[...]
C_CTYPE
Determine the locale for the interpretation of sequences of bytes of text data as characters (for example, single-byte as opposed to multi-byte characters in arguments and input files) and which characters are defined as white space characters.