The Silver Searcher: An attempt to make something better than ack
github.com
github.com
[@] Technically git grep has five modes of operation:
1. Search the contents of the tracked files as they currently are on disk. This is the default.
2. With --cached, search the contents of the tracked files as they are in the index (i.e. ignore any un-added changes).
3. With --no-index, search all files recursively from the current directory down. This allows you to use "git grep" as a "grep -R" replacement even when your CWD is not inside a repo.
4. With --untracked, search all files recursively from the current directory down in addition to files in the index. (The difference between this and --no-index when used inside a repo is that --untracked honors the .gitignore mechanism by default, i.e., --untracked is a synonym for --no-index --exclude-standard when inside a repo.)
5. With a tree'ish (commit, tag, branch name, etc), search all files in the tree.
It also only works on Git repos; having to think before choosing between "git grep" and "ack" just adds unnecessary mental context switching.
Personally, I prefer ack's output, which puts the file name on a separate line, and includes line numbers by default. Together with -C you get a much more readable output.
Mostly I end up just using Sublime Text's built-in file search, which is like ack, but has a generous -C setting enabled by default, and supports replacing.
That said, the readme there lists five reasons why Silver Searcher is better than ack. Two of them are nonsense (who cares what language it's written in, or same-order-of-magnitude differences in how big the executables are?), two sound like they would be pretty trivial changes to ack, and the last is a significant speed improvement. But then you read further down the readme, and it says the current development state is somewhere between "Runs" and "Behaves correctly". Isn't it kind of premature to be bragging about how fast you are before your code actually behaves correctly?
Also, it makes me wonder how much of the speed increase is based on the easy changes filtering out more files...
(who cares what language it's written in, or same-order-of-magnitude differences in how big the executables are?)
That last reason there is about the length of the filename, not about the size of the file. I don't think anyone's supposed to really care, no.It's a little inside joke, because one of the points in favor of "ack", offered jokingly years ago by its developer, is that "ack" is 25% fewer letters to type compared to "grep".
(This feature is still there as point #10 in http://betterthangrep.com/why-ack/, now not offered 100% jokingly.)
Defaults matter. ack is all about having sensible defaults for your most common uses.
- Literal matches use Boyer-Moore-Horspool strstr.[1]
- Files are mmap()ed instead of read into a buffer.
- If you're building with PCRE 8.21 or greater, regex searches use the JIT compiler.[2] Also I call pcre_study() before executing the regex on a jillion files.
- Ag reads your .gitignore and .hgignore files to ignore code you don't care about.
- Instead of calling fnmatch() on every pattern in your ignore files, non-regex patterns are loaded into an array and binary searched.
I wrote a couple of blog posts about profiling The Silver Searcher and improving performance. http://geoff.greer.fm/2012/01/23/making-programs-faster-prof... is the most informative one, IMO.
1. http://en.wikipedia.org/wiki/Boyer%E2%80%93Moore%E2%80%93Hor...
Also, obligatory link whenever Boyer-Moore is mentioned: http://ridiculousfish.com/blog/posts/old-age-and-treachery.h...
http://www.pixelbeat.org/scripts/findrepo
A quick test on a moderately big repo:
$ time findrepo test '*' | wc -l
158819
real 0m0.532s
$ time ack -a test | wc -l
76526
real 0m8.762sMy concern is that someone who's new to all of this isn't going to understand what to do if I just link to http://www.pixelbeat.org/scripts/findrepo
If your find command looks like
find . -name '*.pl' -o -name '*.pm' | xargs grep foo
and it takes 1 second to finish, and your ack command is ack foo --perl
and it takes 1.5 seconds to finish, you can say the grep is faster.But I just timed the time it takes to type those, and they took me 9.2 vs. 2.3 seconds.
So which is faster: 9.2+1.0 or 2.3+1.5?