Ack is a grep-like tool, optimised for programmers
betterthangrep.com
betterthangrep.com
Honestly though, I don't think ag is the right tool for that job. For a single huge file, grep is going to be the same speed. Possibly faster, since grep's strstr() has been optimized for longer than I've been alive.
DNA files don't change very often, which makes building an index worthwhile. Apparently, sequencing isn't perfect and neither are cells, so you'd want fuzzy matching. But repeats in DNA are also common, so that means fuzzy regex matching. There is already a fuzzy regex library[1], but I have no idea how fast it is. If the application requires performance above everything, an n-gram index sounds like the right tool for the job.
After writing the paragraph above, I searched for "DNA n-gram search." The original n-gram paper from 2006 used DNA sequences in their test corpus.[2] I don't know much about DNA or the applications built around it, so I'm glad I managed to recommend a tool that was designed for the job.
1. https://github.com/laurikari/tre/ (used by agrep)
2. Fast nGram-Based String Search Over Data Encoded Using Algebraic Signatures http://cedric.cnam.fr/~rigaux/papers/LMRS07.pdf
Silently skipping parts of it seems like the worst thing to do.
https://github.com/ggreer/the_silver_searcher/blob/master/sr...
I built ag for myself; both as a tool and to improve my skills profiling, benchmarking, and optimizing. Had I known how popular it would become, I would have definitely held myself to a higher standard, or any standard. Most importantly, I'd have written tests. These days, I'm busy with a startup so progress on those fronts has been slow.
So.. thanks :)
One thing though, it skips certain source files seemingly arbitrarily without the -t param and I haven't figured out why... Doesn't seem related to any .gitignore entries that I have been able to identify.
ack turns out to be much faster than grep on these large files, FWIW.
Thanks for making this superb tool :)
[0] https://github.com/ggreer/the_silver_searcher/issues/367 for example
ag is less picky, and the increased speed is a nice bonus.
ack --java "foo"
while with ag you write: ag -G"\.java$" "foo"
But yes, ack and ag feel pretty identical except for the speed. Most of the time the speed improvement is irrelevant to me, except sometimes now I'll use ag in my home folder, and it's still fairly snappy.(I've never needed to use the rackup "rack" command directly, fortunately, if you do you ought to use a different alias)
Or escape alias \rack
(Presumably because in zsh, `which which` says it's a shell built-in, whereas in bash it finds `/usr/bin/which`, so bash doesn't seem to be caring about your aliases.)
That's not a knock on ag at all, and if ag fits your needs, then by all means use it.
It's hard to believe this would give a significant performance boost. Is there evidence of this?
Also, mmap() has the disadvantage that it can segfault your process if something else makes the underlying file smaller. In fact, there have been kernel bugs related to separate processes mmapping and truncating the same file.[1] I mostly use mmap() because my primary computer is a mac.
A side note: Parts of OS X's kernel seem... not very optimized to say the least. See the bar graph at http://geoff.greer.fm/2012/09/07/the-silver-searcher-adding-... for an example.
http://lists.freebsd.org/pipermail/freebsd-current/2010-Augu...
He mentions: "So even nowadays, using --mmap can be worth a >20% speedup."
- replicate the experiment, confirm --mmap shaves off a non-negligible amount of time. It could be that his computer happened to be running something in the background that was using his harddrive, for example, which would skew the results.
- look at the code, figure out the exact difference between what --mmap is doing and what it does by default. Confirm that the problem isn't in grep itself (it's probably not, but it's important to check).
- dig into the kernel source to figure out the difference under the hood and why it might be faster.
edit: I've been reading your posts for a while and I like them, but I keep wondering, why do you have sillysaurus1-2-3?
Regarding my ancestry, I'm sillysaurus3 because I've (rightfully) been in trouble twice with the mods for getting too personal on HN. I apologized and changed my behavior accordingly, and additionally created a new account both times to serve as a constant reminder to be objective and emotionless. There's rarely a reason to argue with a person rather than with an idea. Debating ideas, not people, has a bunch of nice benefits: it's easier to learn from your mistakes, it makes for better reading, etc. It's pretty important, because forgetting that principle leads to exchanges like https://news.ycombinator.com/item?id=7700145
Another nice benefit of creating a new account is that you lose your downvoting privilege for a time, which made me more thoughtful about whether a downvote is actually justified.
...
I just skimmed the bsd mailing list email on why grep is fast which was linked up-thread, and it seems that's somewhat the case. It sounds like since they are doing advanced search techniques on what matches or can match, they use mmap to avoid requiring the kernel copy every byte into memory, when they know they only need to look at specific ranges of bytes in some instances. At least that was the case at some point in the past.
Finally, when I was last the maintainer of GNU grep (15+ years ago...), GNU grep also tried very hard to set things up so that the _kernel_ could ALSO avoid handling every byte of the input, by using mmap() instead of read() for file input. At the time, using read() caused most Unix versions to do extra copying.
P.S. Nice attitude, it earned an upvote from me. Which is probably one reason why your third account has more karma than my first.
So the assumption is that those pages don't even ever get swapped in, but I think that'd only be the case when the pattern size is at least as large as the page size (usually 4KB!), which is not the case in the example in the mailing list. So the mystery continues!
read(fd, buf, 100 megs)
can, in the kernel, do something like: read(fd, buf[0:first-page-boundary])
remap(fd, buf[first-page-boundary:last-page-boundary])
read(fd, buf[last-page-boundary:end])
There you go, zero copy reads. Or at least, minimal copy reads -- at most you will get 2 pages worth of copying.Deleted comment
EDIT: The post was http://geoff.greer.fm/2012/08/25/the-silver-searcher-benchma...
Why is that so hard to believe? It's a standard optimization—the kernel can almost certainly coordinate reading better than your userspace C can.
I don't think that anybody is claiming that mmapping actually changes the algorithmic complexity of the actual search operations.
[1]: http://unix.stackexchange.com/questions/108471/no-output-usi...
function ffjar() {
jars=(./**/*.jar)
print "Searching ${#jars[*]} jars for '${*}'..."
parallel --no-notice --tag unzip -l ::: ${jars} | ag ${*} | awk '{print $1, ":", $5}'
}
Because it uses parallel it spreads the workload across CPUs. I use this frequently when I have to update/rewrite/create build scripts, and I know a class exists but not which jar file it lives in. find . -name "*.jar" | xargs -tn1 unzip -l | grep SomeClassI've tested grep against ack and ag for large text files and grep won handily, especially the latest version of grep.
Also note you can use GNU Parallel to run multiple greps
Exactly. There's no reason that you can't have grep AND ack AND ag in your toolbox to choose from.
alias ack="grep --include=*.c"
Anyway, neat tool. Will check it out soon.
Andy Lester, the primary author of ack, is one of the nicest guys I know of. You wouldn't see any trash-talking on that site. He even changed the name of the site from "better than grep" to "beyond grep" [1].
In fact, he gives props to similar tools like ag and others [2].
[1] https://news.ycombinator.com/item?id=5578304 [2] http://beyondgrep.com/more-tools/
"ack versions 2.00 to 2.11_02 are susceptible to a code execution exploit. Please upgrade to 2.12 or higher ASAP."
From what I can observe, I am generally faster than my co-workers. But it could be possible that with a great IDE, I could be faster yet. I don't feel any tug to leave, but could just mean I'm ignorant of that truly better way.
Nested Search:
ag functionName | ag moreSpecificContextLikeArgs
Find variable changed yesterday: git log -p --since yesterday | ag varName
Find controllers changed yesterday: git log --oneline --showfile | ag controllers
What files did I work on last week: git log --name-only --oneline --author me --since 1.weeks
How many JS commits did I do last month? git log --since 1.months --author me --name-only | ag -i '\.js$' | wc -l
How many JS commits did I do on each file last month? git log --since 1.months --author me --name-only | ag -i '\.js$' | awk '{arr[$1]++} END {for(i in arr) print arr[i]," - ",i}' | sort -r -n
Change a "classname" from MyClass to BetterName: ag MyClass # verify it only finds what you think it will
ag MyClass | awk -F':' '{print $1}' | sort | uniq | while read line
do
sed -i' ' 's/MyClass/BetterName/g' $line
done ack -C5 firstThing | less
/secondThing
One bonus ack trick that I like also like is bulk loading into Emacs for further manipulation (e.g., multi-occur and then occur-edit-mode): emacsclient -n `ack -l functionName`However, when you just want to find stuff fast, it's annoying to have to deal with Perl/CPAN or RVM/Rubygems, especially when the dependencies are not installed on your server/workstation.
That's why I've switched to silver searcher (ag) [2], as it can be installed with any OS package manager (brew, apt, yum).
[1] http://rak.rubyforge.org [2] https://github.com/ggreer/the_silver_searcher
The benefit of grep's ubiquity outweighs any small advantage ack has in usability.
I suggest that you need not limit yourself to only one tool for your code searching. Toolboxes FTW.
I have a simple wrapper over egrep (see https://github.com/sitaramc/ew ) that adds those little extras (ignoring binary files, ignoring VCS directories...).
I'm sure it's improved since the days I tried it, but I tend to be permanently prejudiced against tools where the author can't/won't document the file selection logic and says "there's really no English that explains how it works" when someone asks.
The manual explains: "This is done with command line options that are best put into an .ackrc file - then you do not have to define your types over and over again." Then comprehensively describes options for both command line and .ackrc.
[0] http://beyondgrep.com/documentation/ack-2.12-man.html#faq
[1] http://beyondgrep.com/documentation/ack-2.12-man.html#defini...
(since it's in a sorry state I won't post it here, and it will attract the rage of people for not being compatible with grep/in python/no docs/etc)
"The fuzzy parser supports C, but is flexible enough to be useful for C++ and Java, and for use as a generalized 'grep database' (use it to browse large text documents!)"[0]
Exuberant Ctags supports 41 languages[1], incl. javascript, Tcl, Ruby, TeX, awk, etc.I was wondering if there are obvious ack killer features or areas where it's remarkably superior.
[1] Things might have changed since the last time I personally tried this, at the time grep was significantly faster, especially for fixed string searches -- but then again, I never tried to coerce up a command line that gave the same kind of output that ack/ag does (which could probably be hammerd out with help of awk). So don't take my comment to suggest that these tools aren't valuable, just maybe not for the reason some people (notably not the authors of said tools) claim.
Not just that but an extensible set of file type filters that are simple to invoke is what I had in mind. E.g., the tool would let you perform searches like
find++ --Python projects/archive/200?
or find++ --video trailer
where in the latter case the hypothetical find++ would refer to my config to get a list of video file extensions and then print a list of all files in the current directory and its subdirectories with the word "trailer" in their name. For better effect it would ship with useful filters like "--video" by default.But first getting all files via find, then testing with file, and finally matching against mime-type doesn't sound like something that's going to be as fast as possible...
I tried to see if maybe gvfs (gio - gnome io) could help, but couldn't really find anything directly applicable (although there is a set of gvfs command line tools, like gvfs-ls, gvfs-info, gvfs-mime).
That's one of the big features of ack that the find/grep combo can't replicate is checking the shebang of the file to detect type. In ack's case, Perl and shell programs are detected both by extension:
--type-add=perl:ext:pl,pm,pod,t,psgi
--type-add=shell:ext:sh,bash,csh,tcsh,ksh,zsh,fish
And by checking the shebang: --type-add=perl:firstlinematch:/^#!.*\bperl/
--type-add=shell:firstlinematch:/^#!.*\b(?:ba|t?c|k|z|fi)?sh\b/
Run `ack --dump` to see a list of all the definitions.Thanks for suggesting gvfs. I'll investigate it and similar databases from other packages (I know at least KDE has its own).
I'll send Andy a PR about putting the Github link on the site somewhere.
ack -f --css | parallel grep search_termDidn't read further.