Why GNU grep is fast (2010)
lists.freebsd.org
lists.freebsd.org
I'd love to hear people's experiences on how grep wasn't adequate and why they use ripgrep instead.
(This is not a criticism of Ripgrep: I'm glad it exists and that other people find it useful.)
And the plugins support, which enables something like ripgrep-all, which can then search PDFs, etc.
If I'm scripting, though, I try to stick to common denominator grep.
I can totally see how that would be a small, but impactful difference.
There are many common programs that I never use with their standard default options (which are very bad, IMO), e.g. cp, mv, ln, rm, rsync, date and many others, so I always define aliases for them, which include those options that I want to use by default.
So for grep, the recursive search should be included in the grep alias. There is no need for a new program in order to have this feature.
I don't buy into this. These aliases tend to come at the cost, or at least the risk, that your workflow breaks when you are at another computer or working on a shell on some server that doesn't have this alias. That's why I like additional aliases, like l for ls, but with your favorite options. But I dislike aliases that change default behavior - and often in an intransparent way.
That means that by default, it will take a lot less time and won't ruin your terminal when lines of some generated files (especially minified ones that are all on one line) match your search.
This is the main reason I use ripgrep.
If I'm in a repo, I'm using `git grep`.
That makes `rg` a mostly redundant tool for me since it's optimized for searching source code. I can't really use it as a general purpose replacement for grep since if it doesn't find anything I'm left wondering whether what I'm searching is not really there or whether `rg` just didn't bother to check. Even with `--no-ignore --all`, I'm still not sure whether it searches everything. It's one of those tools that I find is too clever for my own good.
So when `git grep` doesn't cover my use case, my fall back is `find | grep` which contains no magic and I know exactly what it's searching.
I feel like there are a billion features in git (like this) that I don't know about.
Fwiw there’s also a « .ignore » semi-standard which works with several tools, and not just greps e.g. fd also respects it by default.
I often use rg to search code across repositories so it's useful to have that regardless of grep and git
GNU grep also does binary detection by default. You have to opt into -a/--text there too. So maybe GNU grep doesn't always do what it's told either. :-)
And I totally agree that having grep installed everywhere and it's pretty fast enough. But I had a few ripgrep searches that were genuinely eyeblink fast. Like my finger hadn't fully lifted off the enter key and it was done. On 10K+ plus files, about 1 GB, with 1.5M+ LOC. And the default folder recursion and .gitignore handling is a plus.
For scripts I'll still use grep sometimes for the portability reason, naturally.
I imagine the performance difference is even more startling for anyone on a Mac who hasn't replaced BSD grep with GNU grep (install the gnu tools from homebrew and alias "grep" to "ggrep", the performance difference is huge).
I use computers with resource constraints.^1 I have grep in multicall/crunched binaries. To use ripgrep I would have to make include it as a separate binary. Is there a solution similar to crunchgen for making crunched Rust binaries.
I am not sure if the "ripgrep" name is a joke or the author is serious. Assuming the later, I am content to wait for the BSD and Linux projects I use to switch from C to Rust and from BSD/GNU grep to ripgrep, at which point I would imagine it will simply be called "grep". For portability.
1. This may be why I have less need for ripgrep. I try to keep things small. Keeping things small routinely has the deirable side effect of making things relatively fast.
What does the "rip" in ripgrep mean? Not "rest in peace." See: https://github.com/BurntSushi/ripgrep/blob/master/FAQ.md#wha...
If you only search small corpora, then ripgrep's speed benefits obviously don't matter. There's unlikely to be material differentiation among grep tools in those cases. So it's a priori not a concern for you.
ripgrep has other benefits, but they aren't quite as universally compelling as its speed benefits. For example, when searching large corpora, pretty much everyone is going to appreciate a search taking 1 second vs 10 seconds. But many fewer people are going to appreciate, say, automatic transcoding from UTF-16 in order to search data.
Other than that, ripgrep's "smart" filtering by default would be its main benefit. My first link above address that.
There are two ways to talk about "speed." On the one hand, we have the "speed" that is associated with the user experience. That is, given some problem the user cares about, how fast can you solve it? On the other hand, we have the "speed" that is associated with doing precisely the same task and measuring which tool does it faster.
ripgrep is generally faster at both of those things, but it's the former where it really shines:
$ git remote -v
origin git@github.com:nwjs/chromium.src (fetch)
origin git@github.com:nwjs/chromium.src (push)
$ git rev-parse HEAD
5d32cab40f738932eddc017980e2e409c5abef2c
$ time rg 'Xvfb and Openbox' | wc -l
1
real 0.289
user 1.526
sys 1.731
maxmem 87 MB
faults 0
$ time grep -r 'Xvfb and Openbox' ./ | wc -l
1
real 5.405
user 3.489
sys 1.890
maxmem 11 MB
faults 0
(I ran these commands multiple times each until the times stabilized. i.e., The directory tree is in cache.)We're talking about an order of magnitude improvement here to get the same results. And not just in a "ripgrep took 10ms and grep took 100ms, but both are fast enough" sense. This is the difference between "near instant results" and "this is taking annoyingly long."
Now of course, from the perspective of the second idea of speed, this isn't an apple-to-apples comparison. GNU grep is actually searching a lot more data here. We can make ripgrep search the same amount of data as GNU grep quite easily:
$ time rg -uuu 'Xvfb and Openbox' | wc -l
1
real 2.538
user 2.570
sys 3.017
maxmem 72 MB
faults 0
So, still a big improvement over GNU grep, but it's not quite as jaw dropping.The important bit here is that a lot of people care about the improvement at the UX level. You can, for example, get a fair bit of improvement with GNU grep with some extra flags:
$ time grep -r --exclude-dir='.git' 'Xvfb and Openbox' ./ | wc -l
1
real 1.630
user 0.781
sys 0.826
maxmem 11 MB
faults 0
And now you've got to shove that stuff into an alias or a wrapper script. Which... is fine. I did it for a very long time before I wrote ripgrep. I had a whole bunch of aliases and wrapper scripts, many of which were specific to certain types of projects. But once I built ripgrep, all of those aliases and wrapper scripts went away. Because ripgrep's heuristics for smart filtering by default subsumed all of them.Finally, it's worth pointing out that ripgrep isn't intended to replace grep. It literally can't. It's not POSIX compatible. So if you're writing shell scripts and care more about portability, then 'grep' is a fine choice. Indeed, I still use 'grep' for precisely that purpose. See: https://github.com/BurntSushi/ripgrep/blob/master/FAQ.md#pos...
And thanks for making an awesome, open source tool. :)
The author literally concedes that the .gitignore feature was not done for performance, and actually carries a significant overhead in large directory trees. For the sake of comparability, the study was controlled for the .gitignore overhead.
The author of rg wrote a blog post about this. According to what I recall, he did performance comparisons on same limitations and scope. So it's not like in that benchmark, the difference would be due to an obvious fact as this.
ripgrep's speed might come from ignoring files in any given use case, and it might even be the biggest reason why a search completes faster. But in my linked blog post, I control for all of that. Yes, while ripgrep might be faster in some cases because of its "smart" filtering, it's also faster in cases where "smart" filtering isn't enabled.
osx grep -r:
real 2m37.786s
user 2m27.034s
sys 0m3.958s
ggrep -r: real 0m12.842s
user 0m5.754s
sys 0m2.825s
which is the difference between something I'd avoid and something I'll use. LC_CTYPE=C
https://news.ycombinator.com/item?id=4841168That’s not entirely true these days due to things like thermal throttling, but it’s still a great way to think about performance.
$ time grep 'ZQZQZQZQZQ' OpenSubtitles2018.raw.en | wc -l
0
real 1.089
user 0.230
sys 0.858
maxmem 5 MB
faults 0
$ time grep -E 'ZQZQZQZQZQ' OpenSubtitles2018.raw.en | wc -l
0
real 1.094
user 0.210
sys 0.883
maxmem 5 MB
faults 0
$ time grep -F 'ZQZQZQZQZQ' OpenSubtitles2018.raw.en | wc -l
0
real 1.096
user 0.223
sys 0.872
maxmem 5 MB
faults 0> The key to making programs fast is to make them do practically nothing. ;-)
It is a great way to think about performance.
Is there any way that I can write code such that I avoid this work altogether...?
An older, but good article on that: https://brenocon.com/blog/2009/09/dont-mawk-awk-the-fastest-...
This is going straight to my Anki quotes collection :)