Make grep 50x faster
blog.x-way.org
blog.x-way.org
His estimate accounting for that was 7x, but this is clearly not a benchmark that was carefully thought through.
Actually I would recommend people to give it a try to those alternatives, I haven't had to look back to grep again since I am using ack-grep (and now ag)
[2] http://geoff.greer.fm/2011/12/27/the-silver-searcher-better-...
stuff$ du -sh big.log
2.8G big.log
stuff$ time grep -i e big.log > /dev/null
real 0m30.228s
user 0m12.213s
sys 0m3.228s
stuff$ time LANG=C grep -i e big.log > /dev/null
real 0m30.130s
user 0m12.105s
sys 0m3.308s
What is LANG=C supposed to do?I'm still surprised that TFA can claim such a speedup, I would have thought IO speed was the bottleneck when you grep through a big amount of data.
As an other poster mentioned I wonder if the speedup is not mainly disk caching in RAM during the 2nd run.
The built-in BSD grep (2.5.1-FreeBSD) also runs in 30% of the time GNU grep does.
$ time grep -i blah stuff
bLAH
blAH
real 0m7.227s
user 0m7.194s
sys 0m0.011s
$ time LANG=C grep -i blah stuff
bLAH
blAH
real 0m0.486s
user 0m0.467s
sys 0m0.016s
So a pretty big difference. My default lang is en_US.utf8.[0] the locale should have an impact on what character ranges match. See http://stackoverflow.com/questions/6799872/how-to-make-grep-... for an example
Another note is that this is triggering fgrep which is already fast due to it's fixed length expression (i.e. no recursion is involved)
It changes the charset to do not use utf-8.
The improvement mentioned here also has to do with the Boyer-Moore algorithm. When switching the locale from LANG=whatever to LANG=C, we're reducing the size of the lookup table to a fraction of what it previously was. In this case, the fraction is 1/50th, but, as the author said, this will vary between patterns and platforms.
[1] http://lists.freebsd.org/pipermail/freebsd-current/2010-Augu...
It shouldn't that simple – it'd also need to confirm that the pattern wouldn't match any combining characters or normalization would still be necessary.
http://rg03.wordpress.com/2009/09/09/gnu-grep-is-slow-on-utf...
"Update on 2010/10/28: GNU grep is no longer slow on UTF-8. The problem was fixed with the release of GNU grep 2.7. The rest of the article can now be considered obsolete."
That being said there was a bug with grep and UTF a little while back. Debian lists the bug as present in 2.6 and fixed in 2.8:
"grep ." pathologically slow in UTF-8 locales -- http://bugs.debian.org/cgi-bin/bugreport.cgi?bug=604408