Even for multi-file scenarios, the difference is nowhere near close the difference between GNU grep and BSD grep. This means that compatibility with GNU grep takes priority for me and it's not worth switching over to ripgrep.
Even for multi-file scenarios, the difference is nowhere near close the difference between GNU grep and BSD grep. This means that compatibility with GNU grep takes priority for me and it's not worth switching over to ripgrep.
On equivalent tasks, ripgrep is not orders of magnitude faster than GNU grep, outside of pathological cases that involve Unicode support. (I can provide evidence for that if you like.)
For example, in my checkout of the Linux kernel, here's a recursive grep that searches everything:
$ time LC_ALL=C grep -ar PM_RESUME | wc -l
17
real 1.176
user 0.758
sys 0.407
maxmem 7 MB
faults 0
Now compare that with ripgrep, with a command that uses the same amount of
parallelism and searches the same amount of data: $ time rg -j1 -uuu PM_RESUME | wc -l
17
real 0.581
user 0.187
sys 0.384
maxmem 7 MB
faults 0
Which is 2x faster, but not "order of magnitude." Now compare it with how long
ripgrep takes using the default command: $ time rg PM_RESUME | wc -l
17
real 0.125
user 0.646
sys 0.654
maxmem 19 MB
faults 0
At 10x faster, this is where you start to get to "order of magnitude" faster
claims. But for someone who cares about precise claims with respect to
performance, this is uninteresting because ripgrep is 1) using parallelism and
2) skipping some files due to `.gitignore` and other such rules.You can imagine that if your directory has a lot of large binary files, or if you're searching in a directory with high latency (a network mount), then you might see even bigger differences from ripgrep without generally seeing a difference in search results because ripgrep tends to skip things you don't care about anyway.
In summary, there is an impedance mismatch when talking about performance because most people don't have a good working mental model of how these tools work internally. Many people report on their own perceived performance improvements and compare that directly to how they used to use grep. They aren't wrong in a certain light, because ultimately, the user experience is what matters. But of course, they are wrong in another light if you're interpreting it as a precise technical claim about the performance characteristics of a program.
so you bring the trick #1 from the grep author to a new level:
> #1 trick: GNU grep is fast because it AVOIDS LOOKING AT EVERY INPUT BYTE.
You stop looking at entire files, which can avoid a lot of bytes to go through.
Also parallelisms surely helps with a lot of files, or big ones really big ones which could be processed in chunks.
This is like saying "if I remove wings from a plane, then it won't go much faster than my car. So a plane is not technically faster than a car"
Of course, users should know the difference between grep and ripgrep (especially with tweaks like .gitignore, which could be confusing if you don't know).
git ls-files -z | xargs -0 grep -nHE <regexp>
can be almost as fast. TIMEFMT=$'\nreal\t%*E\nuser\t%*U\nsys\t%*S\nmaxmem\t%M MB\nfaults\t%F'Well, it's a base-2 order of magnitude!
That situation changes when you have very short patterns, very long patterns, many patterns, or small alphabets (eg DNA). As burnsushi notes, ripgrep’s performance difference for common feel usage comes more from being smarter about its input than from algorithms (its algorithms are solid, of course).
If you’re talking about the core algorithm, it’s not that much faster but the clever adjustments for common patterns and modern CPU architectures also counts for real world results. It’s just important to be precise so people know there hasn’t been some fundamental breakthrough in searching.
You are measuring the wrong thing or a case where rg can skip most files.