GNU grep is 10x faster than Mac grep
jlebar.com
jlebar.com
It's grep, just better. It highlights the selected text, it shows which files, and in what line the text was found (and uses vivid colors so you can distinguish them easily), ignores .git and .hg directories (among others, that shouldn't be searched) by default, you can tell it to search, for example for only `--cpp` or `--objc` or `--ruby` or `--text` files (with a flag, not a filename pattern), and many many other neat features that I'm sure grep has, but you have to remember and memorize them. ack has sensible defaults.
Why ack? http://betterthangrep.com/why-ack/
manpage: http://betterthangrep.com/documentation/
Oh, and ack is written in perl and doesn't require admin privileges to install.
gfind . -type f -exec grep -i mbr {} \; >| /dev/null
1.10s user 0.81s system 90% cpu 2.113 total
gfind . -type f -exec ack -i mbr {} \; >| /dev/null
24.34s user 4.17s system 96% cpu 29.678 total
(Yes, I know about the flag to search recursively. This is the most fair comparison.)I spared no effort in optimizing. Pthreads, mmap(), boyer-moore-horspool strstr, it's all there. Searching my ~/code (5.2GB of stuff), I get this:
ag blahblahblah 1.93s user 3.54s system 313% cpu 1.749 total
ack blahblahblah 9.75s user 2.79s system 98% cpu 12.690 total
Both programs ignore a lot of extraneous files by default (hidden files, binary files, stuff in .gitignore, etc). The real amount of data searched is closer to 500MB.So re writing this in C will fundamentally mean endlessly growing a language which will look similar to the Perl implementation. Or a Perl DSL.
Not that its a bad thing, I find it interesting though. I would say you better start with a specification.
"Perl Compatible" isn't really Perl compatible, see http://en.wikipedia.org/wiki/PCRE for details.
EDIT: The last precise build works just fine, though.
https://github.com/ggreer/the_silver_searcher/wiki/Windows
The author forewarns that "[i]t's complicated".
ack --ruby --js foo_bar
will search only ruby and javascript files, which means .rb+.erb+.rhtml+.js+...Also exclusion with --no-* is very useful (especially --no-sql).
This is markedly different from 'simply' ignoring irrelevant files, besides the fact that it does not need a 'project' to work (ack --ruby foo_func $(bundle show bar_gem)).
The better part being it is extendable so that I can create --stylesheets covering css+sass+scss+less, or add say .builder to --ruby.
(BTW, love the name/command)
In fact looks like GNU grep has --mmap switch and it's a little bit faster in the simple case than default on my Ubuntu system. But -i makes mmap slower. Maybe GNU grep just avoids mmap because of error handling (you get a segfault/bus error instead of an io error return when things go wrong).
ack -i mbr > /dev/null
I think that starts up perl once, not once per file. If so timing should be much better.ack searches recursively by default; I don't think it can search non-recursively (why would you want to? That is what grep is for)
Also: try comparing grep and ack in a directory tree that has 'garbage' such as .svn or .git directories or .o files.
For every single file found you start ack again. You compare startup times here. Ack is so slow here because it's a perl script. For every single file you start the perl interpreter, and the perl interpreter compiles and interpretes ack every time.
First, without knowing the makeup of files he has, you can't tell how much a corner case this is. It could be 100K small files or 10 large ones. Few care about runtimes for small files, but many care about runtimes for large ones.
Also, and probably more importantly, you'd use ack differently in a recursive-find situation. You just "ack" from the top of the tree. The perl interpreter starts only once.
I don't think this is a useful benchmark for typical uses of ack.
http://travisjeffery.com/b/2012/02/search-a-git-repo-like-a-...
function g! { grep -nr --exclude=.git --exclude=.hg --exclude=.svn --include="*.$1" "$2" ${3-"."} --colour; }
And then: g! py "some python code"
https://coderwall.com/p/uhzc0aTo be fair, neither does GNU Grep - just do `make' (without `make install') and you're good to go.
grep --color
> it shows which files, and in what line the text was found (and uses vivid colors so you can distinguish them easily), grep -rn --color pattern ./files/
files/foo.sh:123: echo "Look at the floral pattern on this dress!"
> ignores .git and .hg directories (among others, that shouldn't be searched) by default, git --exclude=.git --exclude=.hg --exclude=.svn
> you can tell it to search, for example for only `--cpp` or `--objc` or `--ruby` or `--text` files (with a flag, not a filename pattern),You would use `find` in conjunction with `grep`. "Art of Unix Programming", modularity, and all that jazz. Presumably you would just modify your own grep alias or define a function to avoid retyping. The end result pretty much looks like my grep alias:
alias grep='grep -Ein --color --exclude=.git --exclude=.hg --exclude=.svn'
I still fail to see a reason to use ack, especially when I can assume grep is always available for portability.That's why.
% /usr/local/bin/grep --version
/usr/local/bin/grep (GNU grep) 2.14
<snip>
% time find . -type f | xargs /usr/local/bin/grep 83ba
find . -type f 0.01s user 0.06s system 8% cpu 0.870 total
xargs /usr/local/bin/grep 83ba 0.66s user 0.31s system 95% cpu 1.017 total
% /usr/bin/grep --version
grep (BSD grep) 2.5.1-FreeBSD
% time find . -type f | xargs /usr/bin/grep 83ba
find . -type f 0.01s user 0.06s system 0% cpu 28.434 total
xargs /usr/bin/grep 83ba 31.65s user 0.40s system 99% cpu 32.113 totalIncidentally, on OS X, you can commonly get another order of magnitude improvement over even GNU grep with Spotlight's index: use xargs to grep only through files that pass a looser mdfind "pre-screen".
If the author were to set LANG to c. He would find that BSD grep suddenly speeds up tremendously.
brew install https://raw.github.com/Homebrew/homebrew-dupes/master/grep.rb brew tap homebrew/dupes
brew install grepThe two I end up using semi-frequently are gcc and apple-gcc, for those projects that Clang just won't compile.
One common case: I have a Perl script processing a giant file, but it only processes certain lines that match a test. You can move that test to grep, to remove nonmatching lines before Perl even hits them, which will typically be much faster than making Perl loop through them.
Say your script.pl is doing something like:
next unless /relevant/;
You can replace that with: grep "relevant" filename | perl ./script.pl :vimgrep
is slower because it loads each file in memory with all the filetype-specific stuff going on each time before the actual searching.Seriously though, it's really amazing what performance they squeezed of that tool. Always amazing to grep through gigabytes of files in a few seconds.
Someone commented on the article that this might be caused by missing off the -F flag; I tried this, and -F makes both versions slightly faster again.
brew install ack