It's about
iswspace as I mentioned in the parent comment. Replace the line
if (iswspace(wch))
by
if (wch == L' ' || wch == L'\n' || wch == L'\t' || wch == L'\v' || wch == L'\f')
And I get a ~1.7x speedup:
$ time ./wc ../wiki-large.txt
854100 17794000 105322200 ../wiki-large.txt
./wc ../wiki-large.txt 0.47s user 0.02s system 99% cpu 0.490 total
time ./wc2 ../wiki-large.txt
854100 17794000 105322200 ../wiki-large.txt
./wc2 ../wiki-large.txt 0.28s user 0.01s system 99% cpu 0.293 total
Remove unnecessary branching introduced my multi-character handling [1]. This actually resembles the Go code pretty closely. We get a speedup of 1.8x.:
$ time ./wc3 ../wiki-large.txt
854100 17794000 105322200 ../wiki-large.txt
./wc3 ../wiki-large.txt 0.25s user 0.01s system 99% cpu 0.267 total
If we take the second table from the article and divide the C result (5.56) by 1.8, the C performance would be ~3.09, which is faster than the Go version (3.72).
Edit: for comparison, the Go version from the article:
$ time ./wcgo ../wiki-large.txt
854100 17794000 105322200 ../wiki-large.txt
./wcgo ../wiki-large.txt 0.32s user 0.02s system 100% cpu 0.333 total
So, when removing the multi-byte character white space handling, the C version is indeed faster than the (non-parallelized Go version).
[1] https://gist.github.com/danieldk/f8cdaed4ba255fb2954ded50dd2...