Ripgrep is faster than grep, ag, Git grep, ucg, pt, sift (2016)
blog.burntsushi.net
blog.burntsushi.net
function frg {
$result = rg --ignore-case --color=always --line-number --no-heading @Args |
fzf --ansi `
--color 'hl:-1:underline,hl+:-1:underline:reverse' `
--delimiter ':' `
--preview "bat --color=always {1} --theme='Solarized (light)' --highlight-line {2}" `
--preview-window 'up,60%,border-bottom,+{2}+3/3,~3'
if ($result) {
& ($env:EDITOR).Trim("`"'") $result.Split(': ')[0]
}
}
There are other ways to approach this, but for me this is a very fast way of nailing down 'I now something exists in this multi-repo project but don't know where exactly nor the exact name'edit this comes out of https://github.com/junegunn/fzf/blob/master/ADVANCED.md and even though you might not want to use most of what is in there, it's still worth glancing over it to get ideas of what you could do with it
How do you scroll the preview window with keyboard ?
alias pf="fzf --preview='less {}' --bind shift-up:preview-page-up,shift-down:preview-page-down"
That will let you run `pf` to preview files in less and lets you use shift + arrow keys to scroll the preview window. No dependencies are needed except for fzf. If you want to use ripgrep with fzf you can set FZF_DEFAULT_COMMAND to run rg such as `export FZF_DEFAULT_COMMAND="rg ..."` where ... are your preferred rg flags. This full setup is in my dotfiles at https://github.com/nickjj/dotfiles.I've made a video and blog post about it here: https://nickjanetakis.com/blog/customize-fzf-ctrl-t-binding-...
I also made https://nickjanetakis.com/blog/using-fzf-to-preview-text-fil... which covers how to modify fzf's built in CTRL+t shortcut to allow for previews too. CTRL+t is a hotkey driven way to fuzzy match a list of files.
[1] https://github.com/phiresky/ripgrep-all/wiki/fzf-Integration
Neither `glimpse` nor `qgrep`, to my knowledge, directly supports pre-processing / document conversion (like `pdftotext`), though I imagine this would be easy to add to either replicating Desktop Search. (Indirectly, at some space cost, you could always dump conversions into a shadow file hierarchy, index that, and then translate path names.)
[1] https://manpages.ubuntu.com/manpages/focal/man1/glimpse.1.ht...
> Add one or a few drops of water to your roasted coffee beans
Ah, RDT (Ross Droplet Technique)[0].
A little atomizer (“spritz” bottle) of plain water serves well here. NB: this is for single-dose grinding - e.g. measuring a small amount of beans loaded into a grinder to grind immediately. If you have a grinder with a “big” hopper on top that has (e.g.) the weeks worth of coffee (even though you grind on-demand for ea. espresso/french press/aeropress/pourover/drip/…) this isn’t for you.
[0] https://thebasicbarista.com/en-us/blogs/topics/how-rdt-broke...
And added my keyboard shortcuts.
function frg {
result=`rg --ignore-case --color=always --line-number --no-heading "$@" |
fzf --ansi \
--color 'hl:-1:underline,hl+:-1:underline:reverse' \
--delimiter ':' \
--preview "bat --color=always {1} --theme='Solarized (light)' --highlight-line {2}" \
--preview-window 'up,60%,border-bottom,+{2}+3/3,~3'`
file="${result%%:*}"
linenumber=`echo "${result}" | cut -d: -f2`
if [ ! -z "$file" ]; then
$EDITOR +"${linenumber}" "$file"
fi
}Then I tried it and I strongly dislike it. The syntax is clunky, it's really no better than popular Unix shells at being a "real" programming language, and yet it's not as good as they are at being just a shell either.
It also just doesn't feel like a quality product. On my work Windows laptop, Powershell will sometimes not quite bother flushing after it starts, so I get the banner text and then... I have to hit "enter" to get it to finish up and write a prompt. In JSON parsing the provided JSON parser has some arbitrary limits... which vary from one version to another. So code which worked fine on machine #1 just silently doesn't work on machine #2 since the JSON parsers were changed and nobody apparently thought that was worth calling out. If you told me this was the beta of Microsoft's new product I'd be excited but feel I needed to provide lots of feedback. Knowing this is the finished product I am underwhelmed.
That’s independent of the shell, and is I believe a bug in the terminal emulator. There is an open source Windows Terminal you can separately install and that is so much better.
As it turns out, the reason that "curl ..." doesn't work is because it pops up a window below all of my other windows saying that certificate revocation information is unavailable, and would I like to proceed. After that it does download my web page!
YMMV (and obviously does). I think that powershell is night and day better than bash (etc) as a programming language.
I generally don't like MS software, but their commitment to back compatibility is worth calling out.
function frg {
result=$(rg --ignore-case --color=always --line-number --no-heading "$@" |
fzf --ansi \
--color 'hl:-1:underline,hl+:-1:underline:reverse' \
--delimiter ':' \
--preview "bat --color=always {1} --theme='Solarized (light)' --highlight-line {2}" \
--preview-window 'up,60%,border-bottom,+{2}+3/3,~3')
file=${result%%:*}
linenumber=$(echo "${result}" | cut -d: -f2)
if [[ -n "$file" ]]; then
$EDITOR +"${linenumber}" "$file"
fi
} function frg --description "rg tui built with fzf and bat"
rg --ignore-case --color=always --line-number --no-heading "$argv" |
fzf --ansi \
--color 'hl:-1:underline,hl+:-1:underline:reverse' \
--delimiter ':' \
--preview "bat --color=always {1} --theme='Solarized (light)' --highlight-line {2}" \
--preview-window 'up,60%,border-bottom,+{2}+3/3,~3' \
--bind "enter:become($EDITOR +{2} {1})"
end
Still not a fan of the string-based injections based on the colon and newline characters, but all versions suffer from it. (also: nice that fzf does the right thing and prevents space and quote injection by default). code -g "$file:$linenumber"2. You nerd-sniped me into getting rid of the unnecessary `cut` process :)
file=${result%%:*}
line=${result#*:}
line=${line%%:*}https://pubs.opengroup.org/onlinepubs/9699919799/utilities/V...
Also a couple mnemonic hints:
% vs #. On US keyboards, # is shift-3 and is to the left of % which is shift-5. So # matches on the start (left) and % matches on the end (right).
# vs ## (or % vs %%). Doubling the character makes the match greedy. It's twice a wide so it needs to eat more.
Bash also supports ${parameter/pattern/string} and ${parameter//pattern/string} (and a bunch others besides) which are not POSIX:
https://www.gnu.org/software/bash/manual/html_node/Shell-Par...
fza = "!git ls-files -m -o --exclude-standard | fzf -m --print0 | xargs -0 git add"
With that in the [alias] section of a gitconfig file, running git fza brings up a list of modified and not yet added files, space toggles each entry and moves to the next entry.That alias as well as fzf+fd really speed up some parts of my workflow.
Oh and shameless plug for my guide on what to include in your zsh setup on macOS: https://gist.github.com/aclarknexient/0ffcb98aa262c585c49d4b...
git ls-files -m -o --exclude-standard | fzf -m --print0 --preview "git diff {1}" | ....
And that's just the start: it could even be that by binding a key to the fzf reload command to then display the diff in it's finder, and in turn a key to stage the selected line, you could turn that into an interactive git staging tool. GIT_WINDOWS="/mnt/c/Program Files/Git/bin/git.exe"
GIT_LINUX="/usr/bin/git"
case "$(pwd -P)" in
/mnt/?/*)
case "$@" in
# Needed to fix prompt, but it breaks things like paging, colours, etc
rev-parse*)
# running linux git for rev-parse seems faster, even without translating paths
exec "$GIT_LINUX" "$@"
;;
*)
exec "$GIT_WINDOWS" -c color.ui=always "$@"
;;
esac
;;
*)
exec "$GIT_LINUX" "$@"
;;
esac
This allows you to use `git` in your WSL shell but it'll pick whichever executable is suitable for the filesystem that the repo is in :)Yeah, I have a bit of a love-hate relationship with it. But I actually have that with all shells out there. I don't know if it's just me or the shells, or (the most likely I think): a bit of both. But PS is available out of the box and using objects vs plain text is a major win in my book, and even though I still don't know half of the syntax by heart it feels less of an endless fight than other shells. And since I use the shell itself for rather basic things and for the rest only for tools (like shown here), we get along just fine.
(add-hook 'xref-backend-functions #'dumb-jump-xref-activate)
The Xref key sequences and commands work fine with it. If I type M-. (or C-u M-.) to find definitions of an identifier in a Python project, dumb-jump runs a command like the following, processes the results, and displays the results in an Xref buffer. rg --color never --no-heading --line-number -U --pcre2 --type py '\s*\bfoo\s*=[^=\n]+|def\s*foo\b\s*\(|class\s*foo\b\s*\(?' /path/to/git/project/
The above command shows how dumb-jump automatically restricts the search to the current file type within the current project directory. If no project directory is found, it defaults to the home directory.By the way, dumb-jump supports the silver searcher tool ag too which happens to be quite fast as well. If neither ag nor rg is found, it defaults to grep which as one would expect can be quite slow while searching the whole home directory.
Ripgrep can be used quite easily with the project.el package too that comes out of the box in Emacs. So it is not really necessary to install an external package to make use of ripgrep within Emacs. We first need to configure xref-search-program to ripgrep as shown below, otherwise it defaults to grep which can be quite slow on large directories:
(setq xref-search-program 'ripgrep)
Then a project search with C-x p g foo RET ends up executing a command like the following on the current project directory: rg -i --null -nH --no-heading --no-messages -g '!*/' -e foo
The results are displayed in an Xref buffer again which in my opinion is the best thing about using external search tools within Emacs. The Xref key sequences like n (next match), p (previous match), RET (jump to source of match), C-o (show the source of the match in a split window), etc. make navigating the results a breeze!Looking at your regex---just by inspection, I haven't tried it, so I could be wrong---but I think you can drop the --pcre2 flag. I also think you can drop the second and third \b assertion. You might need the first one though.
You can find the rg binary in the VS installation (at least, I can on Windows at my place of employment).
Random other helpful flag I use often is -M if any of the matches are way too long to read through and cause a lot of terminal chaos. Just add `-M 1000` or adjust the number for your needs and the really long matches will omit the text context in the results.
Also the fact that it is a standalone portable executable can be super handy. Often when working on a new machine, I'll drop in the executable and an alias for grep that points to rg, so if muscle memory kicks in and I type grep it will still use rg.
This flag is especially convenient if you want to search e.g. .yml and .yaml in one go, or .c and .h in one go, etc.
I use ag (typically from inside Emacs) on a 900k LOC codebase and it is effectively instantaneous (on a 16 core Ryzen Threadripper 2950X). I just don't have a need to go from less than 1 second to "a bit less than less than 1 second".
Speed is not the defining attribute of the "new greps" - they need to be assessed and compared in other ways.
But as I mentioned in my comparison to qgrep elsewhere in the thread, everyone has different workloads. And for some workloads, perf differences might not matter. It really just depends. 900 KLOC isn't that big, and indeed, for simple queries pretty much any non-naive grep is going to chew through it very very quickly.
As for comparisons in other ways, at least for ag, it's on life support. I thought it was going to get removed from Debian, but it looks like someone rescued it: https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=999962
The blog post also compares Unicode support, and contextualizes its performance. ag essentially has zero Unicode support. Unicode support isn't universally applicable of course---you may not care about it---but it satisfies your non-perf comparison criteria. :-)
This suggests your corpora are small. If you have small corpora, then it should be absolutely no surprise that one tool taking 40ms and another taking 20ms will matter for standard interactive usage.
"Ripgrep – A new command line search tool" https://news.ycombinator.com/item?id=12564442 (740 points | Sept 23, 2016 | 209 comments) - there are discussions related to speed too
"Ripgrep is faster (2016)" https://news.ycombinator.com/item?id=17941319 (98 points | Sept 8, 2018 | 40 comments)
this above all true UNLESS you need multi-line matches with UTF8, where ripgrep is not so fast, because it needs to fall back to the other PCRE2 lib
Yes, qgrep uses indexing, which will always give it a leg up over other tools that don't use indexing. But of course, now you need to setup and maintain an index. The UX isn't quite as simple as "just run a search."
But there isn't much of a mystery here. Someone might neglect to use qgrep for exactly the same reason that "grep is fast enough for me" might prevent someone from using ripgrep. And indeed, "grep is fast enough" is very much true in some non-trivial fraction of cases. There are many many searches in which you won't be able to perceive the speed difference between ripgrep and grep, if any exists. And, analogously, the difference between qgrep and ripgrep. The cases I'm thinking of tend to be small haystacks. If you have only a small thing to search, then perhaps even the speed of a "naive" grep is fast enough.
So if ripgrep, say, completes a search of the Linux kernel in under 100ms, is that annoying enough to push you towards a different kind of tool that uses indexing? Maybe, depends on what you're doing. But probably not for standard interactive usage.
This is my interpretation anyway of your wonderment of (in your words) "why people forget the qgrep option." YMMV.
I have flirted with the idea of adding indexing to ripgrep: https://github.com/BurntSushi/ripgrep/issues/1497
> this above all true UNLESS you need multi-line matches with UTF8, where ripgrep is not so fast, because it needs to fall back to the other PCRE2 lib
That's not true. Multiline searches certainly do not require PCRE2. I don't know what you mean by "with UTF8," but the default regex engine has Unicode support.
PCRE2 is a fully optional dependency of ripgrep. You can build ripgrep without PCRE2 and it will still have multiline search support.
``` ./build.c ```
... and then magic happened.
ripgrep's build.rs used to do more, like build shell completions and the man page. But that's now part of ripgrep proper. e.g., `rg --generate man` writes roff to stdout.
[1]: https://github.com/BurntSushi/ripgrep/blob/2a4dba3fbfef944c5...
[2]: https://github.com/BurntSushi/ripgrep/blob/2a4dba3fbfef944c5...
The optional Google search syntax also very convenient.
So weird. I mean, it's just about some open source tool, right? :-/
There are feuds about open source tools all the time. Text editors, Linux distros, shells, programming languages, desktop environments, etc... And ugrep vs ripgrep may be a poster child for C++ vs Rust.
It is not all bad, it drives progress, and it usually stays at a technical level, I've yet to see people killing each others for their choice of command line search tool.
> ugrep is easily one of the if not most featureful grep programs in existence. And it is also fast.
which is burntsushi, ripgrep's author, defending ugrep from someone saying they only focus on performance at the cost of features.
‡ The main commercial product of the ugrep author's company at the time was the gSOAP code generator (it may still be), and that it not only works but makes a reasonably good C and C++ API from WSDL is proof that it is the product of a genius madman. It also allowed you to create both the API and WSDL from a C++-ish header, and both .NET and Java WSDL tools worked perfectly with it. We needed it to work and work it did.
At the time, the generated API was just difficult enough to use that I generated another ~1k lines of code for that project. IIRC, the generated API is sort of handle-based, which requires a slightly different approach than the strict RAII approach we were using. Generating that code was a minor adventure (generating the gSOAP code from the header-ish file, generating doxygen XML from the generated gSOAP code, then generating the wrapper C++ from the doxygen XML).
Ugrep has that. In my case, I’m working with zipped corpora of millions of small text files, so I can skip unpacking the whole thing to the filesystem (certain filesystems have trouble at this scale).
I’m grateful for both tools. Thanks to the respective authors!
I think the killer feature is compatibility with existing grep command line switches. Not needing to learn a whole new set of options is quite nice.
Sounds like that could introduce a ton of breakage, for little value. People who want a faster grep will use a different thing, while people who use grep can continue to use it. Sounds like an ideal situation already.
My entirely anecdotal and unscientific impression is that rg and grep perform similarly on Linux (though rg has nicer defaults for searching through source code). The old version of grep that Apple preinstalls on the Mac was slower last time I checked though.
With respect to compatibility, see my FAQ on the topic: https://github.com/BurntSushi/ripgrep/blob/master/FAQ.md#pos...
Like automatic encoding detection and transparently searching UTF-16?
Or simple ways for composing character classes, e.g., `[\pL&&\p{Greek}]` for all codepoints in the Greek script that are letters. Another favorite of mine is `\P{ascii}`, which will search for any codepoint that isn't in the ASCII subset.
Or more sophisticated filtering features that let you automatically respect things like gitignore rules.
Those are all things that ripgrep does that grep does not. So I do not favor this explanation personally.
ripgrep has just about all of the functionality that GNU grep does. I would say the two biggest missing pieces at this point are:
* POSIX locale support. (But this might be a feature[1].)
* Support for "basic" regexes or some equivalent that flips the escaping rules around. i.e., You need to write `\+` to match 1 or more things, where as `+` will just match `+ literally.
Otherwise, ripgrep has unfortunately grown just about as many flags as GNU grep.
[1]: https://github.com/mpv-player/mpv/commit/1e70e82baa9193f6f02...
ripgrep is a specialist, opinionated tool, designed primarily to search through source code repositories.
There's not much you can add to general purpose text search to make it faster; you can make it use mmap() at the risk of it crashing on truncated files, you can reduce the expressiveness of regular expressions so they can be computed faster. You could throw out general support for all locales and charsets and hardcode support for only UTF-8 / UTF-16, but you shouldn't.
I was under the impression that grep removed mmap() support because it was slower than normal file i/o
$ ls -l full.txt
-rw-rw-r-- 1 andrew users 13113340782 Sep 29 12:30 full.txt
$ time rg -c --no-mmap Clipton full.txt
294
real 1.337
user 0.470
sys 0.866
maxmem 15 MB
faults 0
$ time rg -c --mmap Clipton full.txt
294
real 1.045
user 0.722
sys 0.323
maxmem 12511 MB
faults 0
But in recursive search, especially when used for lots of little files, they end up provoking substantial overhead that slows everything down.And this might change depending on the platform.
Oh I beg to differ! The blog post goes into this. Here's a simple demonstration using ripgrep 14:
$ ls -l full.txt
-rw-rw-r-- 1 andrew users 13113340782 Sep 29 12:30 full.txt
$ time rg -c --no-mmap 'Clipton' full.txt
294
real 1.419
user 0.539
sys 0.879
maxmem 15 MB
faults 0
$ time LC_ALL=C grep -c 'Clipton' full.txt
294
real 6.911
user 6.078
sys 0.829
maxmem 15 MB
faults 0
$ time rg -c --no-mmap 'DMZ|Clipton' full.txt
1070
real 1.643
user 0.747
sys 0.894
maxmem 15 MB
faults 0
$ time LC_ALL=C grep -E -c 'DMZ|Clipton' full.txt
1070
real 8.317
user 7.384
sys 0.930
maxmem 15 MB
faults 0
No memory maps. No multi-threading. No filtering. No fancy regex engine features or reducing expressiveness. No locales. No UTF-8. No UTF-16. Just a simple literal and a simple alternation of literals. It's just better algorithms.Also, you can disable ripgrep's opinions with `-uuu`. It's not designed to just be for code searching. You can use it for normal grepping too. It will even automatically revert to the standard grep line format in shell pipelines.
If you want to innovate in this space, why sign up for all that? Invent a better wheel, and if people like it, they'll migrate over time.
I remember using ag in the old days, and I use rg now. But there's things rg does by default that I don't like at times... so I go back to old fashioned grep.
rg is at the point where many programmers use it. I think it is on its way to becoming one of those "standard tools". It needs... another 5 years?
When POSIX has a rg standard... we'll know ripgrep "succeeded" and teargrep will soon come into existence ;)
> so I go back to old fashioned grep
If you do `rg -uuu` then it should search the same stuff grep will. Not sure if that's what you meant though.
[1]: https://github.com/minad/consult#grep-and-find [2]: https://lambdaland.org/posts/2023-05-31_warp_factor_refactor...
Very often, I don't want to look for files that aren't tracked under Git VC, and I'm not looking for matches in binary files, so by default, ripgrep does that, which can cut time by 99%. I used to grep in small dirs, now I can I can ripgrep in my whole home, not that I do it, but I can. That + Sourcegraph on master branch, and it makes searching for any other thing than plain text feel sooo slow (Atlassian Confluence and Jira, Google docs, etc.).
Thank you so much Burntsushi and contributors!
The standard way to run "AND" queries is through shell pipelines. That is, `rg foo | rg bar` will only print lines containing both. But composition usually comes with costs. The output reverts to the standard grep format and it doesn't interact nicely with contextual options like -C/--context.
Otherwise it would look like this:
# rg nokogiri | rg linux
<stdin>:11:Gemfile.lock:647: nokogiri (1.15.5-x86_64-linux)
But that is a me problem.The workaround is of course just to pipe into grep instead.
Still losing the coloring but you can't have everything.
rg nokogiri.*linux rg -P '(?=.*pat1)(?=.*pat2)(?=.*pat3)'
You could create a shell function shortcut if you need to use it often. But yeah, having it as a feature of the tool itself would be nice.I wanted and boolean syntax mixed with fzf instant search. It’s not as fast as ripgrep of course but it’s not solving the same problem.
https://github.com/BurntSushi/aho-corasick/blob/f227162f7c56...
But no `_mm256_sad_epu8`. What an oddly specific question..?
It computes `sum(|x[i] - y[i]|)` for consecutive `i` at different offsets, so it should be zero at substring matches.
For context: https://epubs.siam.org/doi/pdf/10.1137/1.9781611972931.10
I was slightly mistaken, the instruction of interest is _mm256_mpsadbw_epu8
I mentioned exactly that paper (I believe) in my write-up on Teddy: https://github.com/BurntSushi/aho-corasick/tree/master/src/p...
If you need drop-in compatibility with grep, then use grep. :-)
For those just using it to search through a codebase, don't forget -F for string literals.
From my perspective it's a no brainer. I don't HAVE a grep (because I don't have a Unix) so when I install a grep, any grep, reaching for rg is natural. It's modern and maintained. I have no scripts anywhere that might expect grep to be called "grep".
Of course if you already have a grep (e.g. you run Unix/Linux) then the story is different. Your system probably already has a grep. Replacing it takes effort and that effort needs to have some return.
sn is 'git status --untracked-files=no'.
I imagine a lot of devs have grep preinstalled. In fact, where is grep not installed, now that WSL exists?
I never worked out how that could be happening.
Another possibility is that your grep search was hogging up your system's memory. That might make it swap. On my systems which do not have swap enabled but do have overcommit enabled, I experience out-of-memory conditions as my system essentially freezing for some period of time until Linux's OOM-killer kicks in and kills the offending process.
I would say the first is more likely than the second. In order for grep to hog up memory, you need to be searching some pretty specific kinds of files. A simple log file probably won't do it. But... a big binary file? Sure:
grep -a burntsushi /proc/self/pagemap
Don't try that one at home kids. You've been warned. (ripgrep should suffer the same fate.)(There are other reasons for a system to lock up, but the above two are the ones that are pretty common for me. Well, in the past anyway. Now my machines have oodles of RAM and lots of I/O bandwidth.)
Source? Ripgrep's benchmarks show it significantly faster.
It really just depends. The way I like to characterize `git grep` (at present) is that it has sharp performance cliffs. ripgrep has them too, to be sure, but I think it has fewer of them.
If you're just searching for a simple literal, `git grep` is decently fast:
$ git remote -v
origin git@github.com:torvalds/linux (fetch)
origin git@github.com:torvalds/linux (push)
$ git rev-parse HEAD
f1fcbaa18b28dec10281551dfe6ed3a3ed80e3d6
$ time LC_ALL=en_US.UTF-8 git grep -c -E 'PM_RESUME'
Documentation/dev-tools/sparse.rst:3
Documentation/translations/zh_CN/dev-tools/sparse.rst:3
Documentation/translations/zh_TW/sparse.txt:3
arch/arm/mach-omap2/omap-secure.h:1
arch/arm/mach-omap2/pm33xx-core.c:1
arch/x86/kernel/apm_32.c:1
drivers/input/mouse/cyapa.h:1
drivers/mtd/maps/pcmciamtd.c:1
drivers/net/wireless/intersil/hostap/hostap_cs.c:1
drivers/net/wwan/t7xx/t7xx_pci.c:15
drivers/net/wwan/t7xx/t7xx_reg.h:7
drivers/usb/mtu3/mtu3_hw_regs.h:1
include/uapi/linux/apm_bios.h:1
real 0.215
user 0.421
sys 1.226
maxmem 161 MB
faults 0
$ time rg -c 'PM_RESUME'
drivers/mtd/maps/pcmciamtd.c:1
drivers/net/wwan/t7xx/t7xx_reg.h:7
drivers/net/wwan/t7xx/t7xx_pci.c:15
drivers/net/wireless/intersil/hostap/hostap_cs.c:1
drivers/usb/mtu3/mtu3_hw_regs.h:1
drivers/input/mouse/cyapa.h:1
arch/x86/kernel/apm_32.c:1
Documentation/translations/zh_CN/dev-tools/sparse.rst:3
Documentation/translations/zh_TW/sparse.txt:3
Documentation/dev-tools/sparse.rst:3
arch/arm/mach-omap2/pm33xx-core.c:1
arch/arm/mach-omap2/omap-secure.h:1
include/uapi/linux/apm_bios.h:1
real 0.078
user 0.259
sys 0.577
maxmem 15 MB
faults 0
But if you switch it up and start adding regex things to your pattern, there can be substantial slowdowns: $ time LC_ALL=C git grep -c -E '\w{5,}\s+PM_RESUME'
Documentation/dev-tools/sparse.rst:1
Documentation/translations/zh_CN/dev-tools/sparse.rst:1
Documentation/translations/zh_TW/sparse.txt:1
real 5.704
user 55.671
sys 0.585
maxmem 207 MB
faults 0
$ time LC_ALL=en_US.UTF-8 git grep -c -E '\w{5,}\s+PM_RESUME'
Documentation/dev-tools/sparse.rst:1
Documentation/translations/zh_CN/dev-tools/sparse.rst:1
Documentation/translations/zh_TW/sparse.txt:1
real 24.529
user 4:34.42
sys 0.753
maxmem 211 MB
faults 0
$ time LC_ALL=en_US.UTF-8 git grep -c -P '\w{5,}\s+PM_RESUME'
Documentation/dev-tools/sparse.rst:1
Documentation/translations/zh_CN/dev-tools/sparse.rst:1
Documentation/translations/zh_TW/sparse.txt:1
real 1.372
user 16.980
sys 0.647
maxmem 211 MB
faults 1
$ time rg -c '\w{5,}\s+PM_RESUME'
Documentation/translations/zh_CN/dev-tools/sparse.rst:1
Documentation/dev-tools/sparse.rst:1
Documentation/translations/zh_TW/sparse.txt:1
real 0.082
user 0.226
sys 0.612
maxmem 18 MB
faults 0
In the above cases, ripgrep has Unicode enabled. (It's enabled by default irrespective of locale settings. ripgrep doesn't interact with POSIX locales at all.)In some trees, git grep will be a lot faster because it searches a smaller part of it.
Once again mercurial has/had more useful defaults, `hg grep` searches through the history by default, that’s it’s job.
One practical result of this is that it will mean `ack` will be quite slow when searching typical checkouts of Node.js or Rust projects, because it won't automatically ignore the `node_modules` or `target` directories. In both cases, those directories can become enormous.
`ack` will ignore things like `.git` by default though.
I believe `ag` was the first widely used grep-like tool that attempted to respect your .gitignore files automatically. (Besides, of course, `git grep`. But `git grep` behaves a little differently. It only searches what is tracked in the repo, and that may or may not be in sync with the rules in your gitignores.)