time LANG=C grep asdf < les_misérables.txt > /dev/null
versus time LANG=en_US.UTF-8 grep asdf < les_misérables.txt > /dev/null time LANG=C grep asdf < les_misérables.txt > /dev/null
versus time LANG=en_US.UTF-8 grep asdf < les_misérables.txt > /dev/nullMy guess is that p9idf has it right, and grep is just converting everything to wchar_t first, rather than trying to do any sort of clever searching directly on the UTF-8 byte stream.
Or if your input data is sufficiently ASCII-ish, and so is your search pattern, then why not just force the process locale to C and avoid the whole mess to begin with.
I'm suddenly left wondering how the "." regex syntax functions in the face of surrogates when handling UTF-8.
However, if you find four bytes 'asdf' in the input, you still have to check whether a combining mark follows the 'f'. For this example, that is simple, but I guess things get hairy for many regexes found in real life, such as ones containing even a single period.
The UTF standard used to allow non-conformant representations of ASCII characters. As http://www.schneier.com/crypto-gram-0008.html notes, this lead to security problems. Now the standard says that you can't allow non-conformant representations of ASCII characters. And if you look at http://www.unicode.org/versions/Unicode6.0.0/ch03.pdf and scroll to page 94 you'll find that you can't be said to be conformant unless you explicitly reject non-conforming input.
Therefore UTF-8 decoders cannot be considered conformant unless they actually examine each and every byte to verify that there is nothing dodgy.
I no longer notice this behavior:
dfc@motherjones:~$ grep --version
grep (GNU grep) 2.9
dfc@motherjones:~$ time LANG=C grep asdf < pg135.txt > /dev/null
real 0m0.017s
user 0m0.008s
sys 0m0.004s
dfc@motherjones:~$ time LANG=UTF8 grep asdf < pg135.txt > /dev/null
real 0m0.017s
user 0m0.012s
sys 0m0.004s
dfc@motherjones:~$ time LANG=en_us.UTF8 grep asdf < pg135.txt > /dev/null
real 0m0.012s
user 0m0.004s
sys 0m0.004s
There is not a lot of info about this in debian bug 604408http://bugs.debian.org/cgi-bin/bugreport.cgi?bug=604408
If memory serves me correctly this upstream fixed this sometime after 2.7.1 or 2.7.3
Funner fact: GNU grep used to be slow with UTF.
; grep --version
GNU grep 2.5.3
; time LANG=C grep asdf < lesms10.txt > /dev/null
real 0m0.025s
user 0m0.011s
sys 0m0.014s
; time /usr/local/plan9/bin/grep asdf < lesms10.txt > /dev/null
real 0m0.082s
user 0m0.043s
sys 0m0.013s
; time LANG=en_US.UTF-8 grep adsf < lesms10.txt > /dev/null
real 0m1.209s
user 0m0.818s
sys 0m0.018s
Those are the only two grep implementations I have handy. GNU grep 2.6.3 takes the same amount of time searching for 'asdf' in both locales, but searching for '.' is still slow. Thanks for pointing that out.