Feature comparison of ack, ag, Git-grep, grep and ripgrep
beyondgrep.com
beyondgrep.com
Ack is also nice, I've used that quite a bit too. It has the advantage of being in Perl, so if you're on a "secure" computer (no compiler), you can still use a fast + featureful search tool.
It's a collection of improved shell tools, organized by the tool they supplement.
As with this feature comparison chart, patches and suggestions are welcome: https://github.com/petdance/altbox
In particular "Show proximity of matches to other matches" would be a huge boon to replace `grep -C 5 foo | grep bar`.
So it's no comparison for me.
So far, I have been unsuccessful in finding a grep replacement that can read patterns from a file, and which also uses a DFA engine. Does one exist? From the table, it looks like ripgrep might be suitable. Is it?
Now, this will do much better than a backtracking engine, but if you get up into the tens of thousands or hundreds of thousands of regexes, it's going to get pretty painful. Finite automata just doesn't scale that well. At that point, you really start wanting a more specialized solution. Probably the best answer to that that I know of is Hyperscan. And you're in luck; someone maintains a fork of ripgrep with support for Hyperscan: https://sr.ht/~pierrenn/ripgrep/
(A special case is tens of thousands of literal patterns. ripgrep will notice that and should use Aho-Corasick. It doesn't help so much with search time since it's just a NFA or a DFA like with regexes, but the machine itself is constructed much more quickly.)
It sounds like either plain ripgrep, or ripgrep+hyperscan, is pretty much exactly what I'm looking for. Next time I have this problem, I'll certainly be reaching for it.
It makes a lot more sense to me for something like Hyperscan to be maintained out of tree. I did work with the patch author a bit, and in particular, made some changes to ripgrep to make maintaining such a fork easier: https://github.com/BurntSushi/ripgrep/issues/1488
Bottom line is, a lot of people think that adding a dependency has nearly zero cost. But it doesn't. Not by a long shot.
You could probably do some fun optimizations by grouping the regexes which depend on the same literals into their own sets, but I never needed to.
The problem is that for a big enough NFA, you'll wind up spending most of your search doing powerset construction to build the DFA.
> One thing you'll have to watch out for is RE2's maximum DFA size
You can configure this in ripgrep with the --dfa-size-limit flag. (See also --regex-size-limit.)
I originally had it as a "phrasebook" of how to do the same thing in the different tools, but it was really ugly and took up a lot of horizontal space, and I figured it was more useful as a chart of yes/no. Also, there were cases where two tools had pretty much the same feature, but not exactly, so just listing flags didn't make sense.
I've still got a lot of the data of the switches in the JSON file that I build the chart from. https://github.com/beyondgrep/website/blob/dev/features.json If you've got ideas on how to bring back the phrasebook format, either integrated into this page, or as a separate standalone page, I'd love to hear them. Maybe the phrasebook isn't best done as a table like this, for example. Open a ticket in GitHub and let me know your thoughts.
It is really very fast.
http://users.itk.ppke.hu/~sikbo/nytech/gyak/05_morfo/xfst/bo...
Advances:
- symmetry input:output (reversable)
- readable/maintainable expressions due to _naming_ of sub-expression
Implementations:
- Xerox XRCE XFST/lexc/twolc compilers
- FOMA - https://fomafst.github.io/
>PCRE support is here to stay, but consider this option experimental when combined with the -z (--null-data) option, and note that ‘grep -P’ may warn of unimplemented features.
I did come across a few issues mentioned with -z on unix.stackexchange a few years back but they have been fixed as far as I know.
> Print lines by number
is not supported by GNU grep???
$ grep --version
grep (GNU grep) 2.25
$ cat world
hello
$ grep -Hn hello world
world:1:hello $ grep ... | head -n ...
I think this feature comparison misses how grep is supposed to be used. (See also "Pipe output through a pager or other command")However, ack and ripgrep's default unpiped output is grouped by file, and if you pipe the output, it doesn't do the grouped output.
The idea of "supposed to be used" is also different for ack than it is for grep. ack is specifically less of a general-use tool than grep. It's meant for searching source code. This is also why I have never said that ack is a replacement for grep.
I prefer tools designed with does one thing well philosophy. It lets me scale my knowledge. I can solve many problems with find and parallel not supported by these grep clones.
> I prefer tools designed with does one thing well
This is pretty unlikely. For example, you probably use a grep tool that will also do recursive directory traversal for you. It probably even has flags for defining filters on that traversal. Why use such a tool when `find` already does recursive directory traversal for you?
However, the grep -Hn feature is described in this comparison as, "Prefix the line number to matching lines"
One thing that can help people to compare this sort of tool is to pair technical descriptions like command line parameters with the natural language explanation. If tool foo has a feature you're describing as "Prevent cheesecake" then I have no idea if my tool bar can do that, whereas if you say this is -Xqm then I can read the documentation and discover that I call this "disable refrigerated dessert" and it's -VQb so yes, my tool does this too.
I spent some time recently reading the proposals to fix/ extend C++ ranges P2214 - and because this general idea is very common they often discuss Haskell, Rust or even Python. If you're experienced in a language you already know whether it would spell something FlatMap, flat_map, or flatMap but you might not guess that C++ people would call your filter_map by the name transform_maybe, or as a C++ programmer who has barely dipped their toe in Haskell you wouldn't know that Haskell doesn't use the word "transform" in this context and without being told what it's called you won't find the relevant documentation let alone be able to try it for yourself and appreciate what it's for.
If you've got suggestions on improvements, please submit an issue. I'd love to hear them.