Ripgrep 11 Released
github.com
github.com
In addition to being really fast, it "just works". By that I mean it automatically excludes the files I want to be excluded, like `.gitignore`d files and binary files. I know I can configure ack-grep (ag) and other tools to do that, but not needing to configure it is nice.
BTW, if anyone hasn't read https://blog.burntsushi.net/ripgrep/, highly recommended. It's about how ripgrep is so fast. (Edit: Just saw another comment mentioned this, too. Goes to show that making a single high-quality blog post has a big impact.)
Now, I will do my obligatory "burntsushi needs to clag the rest of the Teddy code" post which I must do on every discussion of ripgrep. :-)
The "subtitles_alternate_casei" examples would be a good benchmark - these 48 strings should not overwhelm Teddy as they could be sensibly merged into Teddy's 8 "buckets" (in fact, they could be merged into 5 buckets) using the simple greedy merging strategy in the original Teddy implementation.
This would probably be a quite good project for someone who wants to contribute to ripgrep and could likely get some nice performance wins...
I'll get to it myself eventually if someone else doesn't, but it will likely be a while.
Teddy is very much 'of its time' (SSSE3) and there are a lot of new approaches that seem interesting (AVX512 of Skylake generation, VBMI, Sunny Cove's even bigger slate of instructions, ARM NEON, SVE).
I also have better ideas about followup confirm than I used to. There are also some prospect to pick a 'fragment' out of the whole string within Teddy or equivalent at a position not strictly at its suffix - this can even be done with ordering preserved if you are careful not to make fragment choices that allow o-o-o matches (only possible if strings overlap).
I might do a bit of work on this, but I'm a bit jaded on string matching and regex matching after 13 years.
I am hoping to move on, but I admit I do have a "few more ideas" in that area - possibly even slightly less 'meatheaded' than previous outings. Maybe (although everything looks better on paper).
> I will introduce a new command line search tool, ripgrep, that combines the usability of The Silver Searcher [ag] …
(Ack also supports the full PCRE syntax but that's less of an issue now rg has -P).
If your work takes you to minority platforms ack is certainly worth keeping in your toolbox.
I'm using ag inside my source directories and cannot remember it not doing the job.
Only I've never been able to make ag (or pt, another similar too) respect .gitignored and other such settings as good as rg does out of the box. Plus it's slower.
I basically never want this, so for me, ripgrep's one flaw is that it never "just works".
alias rg="rg -uuu"Separately, ripgrep is a great advertisement of what Rust is capable of. First, it shows it’s possible to write highly performant and reliable applications without resorting to non readable code. Second, it shows how simple it is to separate applications into a library + binary thanks to Rust’s package management. Ripgrep is merely a command line interface to libripgrep, which is reused in other applications like VS Code. The regex and directory walking code that was once part of the main ripgrep code base are now crates that are reused to great effect throughout the ecosystem.
libripgrep is nice for building your own Rust programs for searching stuff. But there is a not insignificant amount if code that translates your argv into uses of libripgrep.
The author breaks down all of these details in this 17,000 word blog post, full of detailed benchmarks. https://blog.burntsushi.net/ripgrep/ (It's from 2016, but it's still mostly in good shape; ripgrep has only gotten faster since then!)
Do you know what they are? I haven't noticed them, although it has been a while since I've rigorously benchmarked it. From looking at the source code, I don't think much has changed. e.g., it appears to still be using memory maps on Linux when searching large directories.
Also, your README describes "modularity" as a difference between fgrep and ripgrep, but ripgrep appears _significantly_ more modular. Its components are split into several libraries, each with their own good API documentation. You might consider mentioning that, as your README kind of makes it sound like ripgrep isn't available as a library where as fgrep is.
BTW, I see a 20% performance jump from rg-0.10 to rg-11.0 for a single file benchmark. What are the key differences between these two versions?
I don't know. It would be easier to explain if I could more easily see what actual commands are being run. Your README just has you running `./all_tests`, but I want to see the actual command line invocations so that I can reproduce them. I'm also not sure which benchmark in particular you're referring to, so I don't know which corpus to use. Look at ripgrep's README for an example of what I mean. All the inputs are carefully specified and the commands being run are clear. You can even see the raw commands for the full benchmark suite: https://github.com/BurntSushi/ripgrep/blob/master/benchsuite...
I realize doing benchmarks right is a lot of work. So if you just have a particular command for me to try and compare performance, then I'd be happy to just do that.
> Note that the matched lines may not the same for a search pattern
Indeed. ;-) That's exactly why I asked. That's a really important UX concern IMO.
I mostly use Vscode because of its fast search, decent auto complete and simple editing experience.
Whoever made ripgrep, deserves massive kudos.
I would use OSX's slow grep every single day to find things and I never had any of the non-problems problems which ripgrep is trying to solve. About recursive, case insensitive and search and file number just use grep with the flags "-Rni", so I don't get in which aspect is superior. I never ever had half a thought about grep's performance at the systems I use it (already faster than my cognitive bottle neck). If I ever would need something better than grep I would try to move to the next level or deal with lots of unstructured data I would perhaps move to a local search engine using indexes.
And personally I would consider an antipattern to skip the gitignore contents per default. What kind of machines or use cases are hitting the people in the comments?
Do you refer maybe to deep interdependent javascript projects in which a small library requires hundreds of MB in dependencies? How many Gb have your projects and how many lines?
Please include deps & builds (my current project is a small python one, 1.8Gb with deps, and I would usually grep only over the same folders all the time):
$ time find ./ -type f -print|wc -l
140197
real 0m0.626s
user 0m0.165s
sys 0m0.473s
Anyway, if it is about grepping over several gigabytes of unstructured code and there is no chance to have a better tool to index everything that could be a use case.Vanilla grep, meanwhile takes 30 seconds -- too long to search while maintaining flow.
For me, the whole thing _is_ that it is faster than grep. AFAIK, a lot of its speed is due to breaking compatibility with grep to allow for more convenient defaults (skip binary files by default, skip .gitignore by default, etc) which also happen to be faster. For me this is a clear win-win.
Also, I have an 8 core machine. To be waiting for a 30s search knowing 7 cores are doing nothing makes me sad.
That's a good use case for a new tool. I might try the next time my greps take longer than 2 seconds.
P.D. May I ask which kind of project/tech has this ratio of LOC and size?
The project is a fork of LLVM. We're implementing compiler support for a new design for threading on x86.
More "technical" reasons here: https://github.com/BurntSushi/ripgrep#why-should-i-use-ripgr....
For me, the performance is best in class, ux is great, and has defaults that make sense.
Perhaps this helps add context. Every day there are probably thousands of new people learning to write software and learning and command line search tools. They would compare grep, and perhaps something like ripgrep on how they stack up today. They probably don't want to search their javascript dependencies by default (from gitignore, eg node_modules). They probably don't want to search their build (binary/gitignore). They would like to have an easy way to say "only search javascript files" or "only search go files". They would like the fastest tool for the job, why would you opt for the slower one? None of greps historical clout/muscle memory has an impact on them. They would probably appreciate coloured, formatted for easier human consumption output by default.
Sorry for the rant!
I just tried ripgrep and like it even more because it automatically follows .gitignore, which is huge. To me, that's the killer feature.
That's OK. Not everyone has the same set of problems, and not every tool that is built must be used by everybody. I document pretty clearly what ripgrep does in the README, so if it doesn't target your use cases, then obviously you just don't need to use it. But also, just as obviously, it should be clear that plenty of other people do fit into the use cases that ripgrep targets. It turns out that a lot of people work on big code bases, and as a result, running the default grep command can be quite slow compared to something that is just a little bit smarter. It gets super annoying to always remember to type out the right `--include` commands to grep. I know, because I did it for ten years before I built ripgrep. I mean, this is the same premise that drives other tools like `git grep`, ack and ag. ripgrep isn't alone here.
See also: https://github.com/BurntSushi/ripgrep/blob/master/FAQ.md#pos...
> And personally I would consider an antipattern to skip the gitignore contents per default.
Then use
alias rg="rg -uuu"
and you'll never have to worry about that specific anti-pattern ever again. ;-)At the moment first impressions are good, seems to be killing it at the most common grep use case I am having, so I will continue using it and trying for some weeks until I can have a better formed opinion.
The only unknowns for me are are the edgy use cases such as piping to other processes where I might end up falling back to system's grep or maybe learn better the rg possibilities.
In my usecase ripgrep is killer, I use it both to search through something around 10GB of a monorepo, and to search through 400GB of textfiles (custom format, similar to json) every day and not having rg performance would make my life considerably worse.
When you have very large sets of data to search through the difference from rp to other solutions is like night and day.
Tangentially, it is because of great tools like these written in Rust that I have major respect for the language.
alias rgpre='rg -i --max-columns=1500 --pre rgpre'
*edit Adding my script here, and would be grateful for any improvements. Goal was to catch doc,docx,ppt,pptx,xls,xlsx, epub, mobi, and pdf and still work on termux with default packages there only.
https://gist.github.com/ColonolBuendia/314826e37ec35c616d705...
One bummer is that it can slow ripgrep down quite a bit: one process per file is not good. As a tip, if you add
--pre-glob '*.{pdf,xlsx,docx,doc,pptx,...}'
and whatever else you need, then ripgrep will only invoke an extra process for those files specifically, which can significantly speed up ripgrep if you're searching a mixed directory (or just want to have `--pre` on by default).A more prominent mention, and a sample script, would be very useful imho. Just googling for a way to do this for even just one of these file types yields inferior commercial solutions at best, often just nothing.
And while I'm not sure how to make it much faster from a search perspective, I was able to find a 40% improvement in invocation by switching from rgpre to rgg as my alais.
I ended here:
alias rgg='rg -i --pre-glob '*.{pdf,xlsx,xls,docx,doc,pptx,ppt,html,epub,mobi}' --pre rgpre'
Even a cursory mention of the ability at the end of the readme would likely help many. With a verbatim script and verbatim command using it helping many more. I only stumbled upon this feature with enough context to really get it in the Ubuntu man pages and I don't know if I ever would have gotten to the pre-glob feature to be honest.
I do believe that technically oriented non-programmers would be the biggest winners from more exposure around it.
I've lately been doing some Windows programming. Ripgrep from a Git Bash command-line can generally find ALL instances of some word before Visual Studio's Find In Solution can even return one result!
The only downside is that rg does not understand namespaces when searching, so a commonly used member variable will show up all over the place. Can still usually find the right location by visual inspection before VS2017 though!
I just tried again and same thing. If I run with "-u" I get results.
I've just spent some time hunting and figured out that in my "projects" directory I had a ".gitignore" that had "", which I did a long time ago to stop "git status" from spamming me if I ran it in the wrong directory (base projects, a hg project which I have a lot of).
For some reason "ack" recognizes that it is in a hg repo when I run it, and "rg" also recognizes that, but it ALSO* searches up the tree looking for a ".gitignore" which ack does not. Good to know.
If ripgrep doesn't search any files at all, then it should print a warning to stdout.
ripgrep ascends your directory hierarchy for .gitignore files because that's how git works; .gitignore files in parent directories apply. ripgrep doesn't know anything about Mercurial, so it doesn't know to stop ascending your directories.
ack is different since it doesn't look at your .gitignore files, so it doesn't have the same behavior here.
Python wouldn't be a good choice for cli usage
Perl is awesome to use from cli, and it is not just simple search and replace, see my tutorial[1] if you want to see examples
sed and awk are awesome on their own for cli usage, sed is meant for line oriented tasks and awk for field oriented ones (there is overlap too) - one main difference compared to perl is that their regex is BRE/ERE which is generally faster but lacks in many features like lookarounds, non-greedy, named capture groups, etc
you could check out sd[2] for a Rust implementation of sed like search and replacement (small subset of sed features)
[1] https://github.com/learnbyexample/Command-line-text-processi...
rg -n pattern | sed 's/:/(/' | sed 's/:/)/'
That won't work in every case (e.g., in the case of a file path containining a `:`), but it might get you pretty far. And it could probably be made more precise.Otherwise, you could always use ripgrep's --json output mode to craft its output into whatever format you like, but that's probably more work. :-) The JSON format is documented here: https://docs.rs/grep-printer/0.1.2/grep_printer/struct.JSON....
Parallelism? Some Rust-specific features?
It's even faster and comes with its own built in evaluation. There's also no more reliance on an external database.