Use ack instead of grep to parse text files
stevengharms.com
stevengharms.com
Things that do sell ack, for me:
ack css_class --sass # search .sass and .scss
ack some_method --no-flash # ignore .as and .mxml
# ignore compiled css in every Rails project on
# my system (as long as I `ack` from the root)
--ignore-dir=public/stylesheets/compiled
And the fact that it prints out like this: path/to/file.ext
123: some text matching
234: more text matching
path/to/other/file.ext
480: a match
instead of like this (with `-n`): path/to/file.ext:123: a match
path/to/other/file.ext:567: another match
path/to/that/file/you/didnt/know/you_had.ext:32: yet another match
makes it massively more useful for human-viewing of the results than the normal behavior of grep. And it reverts to grep-like output when you pipe it into something, so you can go from exploration to composition with no effort.For the curious: It's "ack-grep" in Ubuntu's package manager (and presumably Debian, though I can't say for sure); I stick it on every machine/server I set up just to have it handy. Queries to the effect of ack-grep --python ClassName yield fast, readable, extremely useful output, as you mention. That's why I use it in addition to grep.
ack '(?=silver).*needle' haystack
is really not the same as: grep needle haystack|grep silver
Because with the double grep method, a match will be made whether "silver" is before or after "needle"; while with the single ack command shown above, a match will only be made when "silver" comes before "needle".Also, I'm not sure why the author uses
ack '(?=silver).*needle' haystack
instead of simply: ack 'silver.*needle' haystack grep 'silver.*needle' haystackNow, there's stuff that never sticks in my brain (tests in shell, sigh). But generally there's less syntax and therefore less to remember in a chain of greps. Composition of simple piece is easier to understand than one equivalent and therefore more complex piece. Heck, the power of the shell is predicated on this idea.
Perhaps the best part about ack is that it's simple to restrict your search of files to a given pattern with a command line flag rather than using shell globbing. You could wrap invocations of grep with a shell function or another script, but that's still not great.
But that too seems like a demonstration of something. The more "simple" methods with obscure names that populate the Unix toolbox, the more confusing it gets. I've gone from find to locate recently, for example, but their functionalities kind of overlap and so when I do find, I'm rusty with it.
You mean the engine that lets you write pathological regular expressions[1] and accidentally ReDoS[2] yourself? To be fair, it's fine if you understand how the engine works well enough to avoid these cases. But how many people can actually say this?
Why not simply $ grep "silver needle" haystack?
When I use double grep in real life, I often tend to do so on a relatively large haystack, where I don't necessarily know what the second search term will be. In that situation, I'll usually do the first grep, look through its output, and add on the second grep once I see something in the first grep's output that I want to narrow the results down to.
Of course, instead of adding on a second grep, I could modify the original regex (and sometimes I do); but if the original regex is complicated, then modifying it is error prone. And, anyway, using a shell abbreviation, it's very easy to type " G " and have that expand to " | grep " to simply add on another grep, without touching the first regex.
A second, quite common use case for a double grep is when I want the second search term to match whether it's before or after the first term. There's probably some convoluted way to get the same effect using a single regex, but it probably won't be nearly as easy or intuitive as a double grep.
cp /usr/bin/grep ack
Find the needles ./ack needle haystack
Find the silver needles ./ack silver.*needle
Find all needles except lead ones ./ack '[^^][^e.][^a.][^d.] needle' haystack
That last one could be tricky if there's other types of needles with names like "ead needle" or "mead needle". But using the haystack he gives us BRE can do the job, easily.Perl regex may be easy to use but they are inferior from a performance perspective. As someone else said, they're slower than BRE or ERE. Moreover, even if speed is not an issue, you pay a price in the amount of memory you will need compared with line-based utilities like, e.g., sed and awk.
Find the needles
sed '/needle/!d;/needle/q' haystack
Find the silver needles sed '/silver needle/!d;/silver needle/q' haystack
Find all needles except lead ones sed '/lead needle/d;/needle/!d/needle/q' haystack
My preference is to use (f)lex if I want a fast "parser" (scanner). Its regex is more than adequate.- grep uses by default the same regular expressions as sed, which is another frequently used tool.
- grep also supports perl regular expressions.
- grep is available on every linux/bsd/*nix system out there, so it just works and make your scripts work.
- We use grep to search through gigabyte sized files (ie logs). You didn't show us how well ack performs there.
moses@deunan:~$ </etc/mime.types grep application |grep x-ruby
application/x-ruby rb
moses@deunan:~$ </etc/mime.types ack application x-ruby
application: No such file or directory
x-ruby: No such file or directory
moses@deunan:~$ ack -h
ack v1.39 Copyright 1993,94 Ogasawara Hiroyuki (COR.)
usage: ack [-{e|s|j|c[c]}] [-{a|A|o<file>}] [-zCntud] [-{E|S}] [<file>..] ack -C5 'scope(?!.*lambda)' app/models
Better than: grep -C5 scope app/models | grep -v lambda
? $ grep needle haystack|grep silver
This sucks.
It would be nice if the argument against this had some sort of substance.In most cases grep is going to be faster that ack. If you are searching large files this can make quite a difference.
Also, `ack` is not installed by default, which is reason enough to not get too used to it. Some people will say "optimize for being on your own machine, since you are 99% of the time", but I'm not. Installing additional utilities on multiple production servers is annoying enough, and can actually become problematic in a PCI-compliant environment as mine is. I'm also frequently helping out other members of my team, and having a magic one-liner that often results in "-bash: ack: command not found" is not terribly useful to me. YMMV.
When I need to do these sort of tasks, I do them in a scripting language with some combination of split() and regex instead of using command line tools. But, I'm just doing that because it's what I know.
Would I end up saving a significant amount of time if I learned to use grep instead?
Shell tools in general are relatively simple or at least specialized, and they're built to be composed in novel/useful ways. The interface between all of these is text, aka data, aka what is arguably the simplest interface.
Well, let's apply that same criteria to searching for two terms:
> grep needle haystack|grep silver
> ack '(?=silver).*needle' haystack
Look at that one character shorter than ack and just 10 times easier.
grep wins, by a knockout.
Saying "You should stop using them. Now." is ridiculous. Why do I need to stop anything if it does what I need?
dfc@ronin:~$ apt-file search bin/ack
ack: /usr/bin/ack
ack-grep: /usr/bin/ack-grep
...
dfc@ronin:~$ apt-cache search --names-only ^ack
ack - Kanji code converter
ack-grep - grep-like program specifically for large source trees
dfc@ronin:~$[shameless plug] Try also the pure-Python based alternative to ack called "pss" (pip/easy_install or https://bitbucket.org/eliben/pss).