Ripgrep 13.0
github.com
github.com
He knows how to keep a good changelog https://github.com/BurntSushi/ripgrep/blob/master/CHANGELOG.... it’s good to read through and see the progress.
He also cares about ergonomics as well as performance which I like. I find it much more easier to use than existing Grep tools.
Interestingly, I believe Ripgrep started off as a test harness to benchmark the Rust Regex crate (he was the author of that). It was originally not intended to be a consumer cli. I guess he dogfooded it that well it became a successful tool
That's right. IIRC, I wrote it as a simple harness because I wanted to check whether someone could feasibly write a grep tool using the regex crate and get similar performance. Once I got the harness up and running though, I realized it was not only competitive with GNU grep in a larger number of cases than I thought it would be (after some perf bug fixes, IIRC), but was substantially faster than ag. So that's when I decided to polish it up into a tool for others to use.
[1] - https://github.com/BurntSushi/ripgrep/blob/master/RELEASE-CH...
If not, one cannot convincingly tell what the changes are actually for to other people and the changelogs get vague or omitted of details.
On macOS: "brew install ripgrep" - then use "rg searchterm" to search all nested files in the current directory that contain searchterm.
You may be using it without realizing already: the code search feature in VS Code uses ripgrep under the hood.
Usually I'm searching code for things like variable names, so actually case-sensitive searches are a better default for me.
On Unix, you can just do `alias rg="rg -S"`, and you get smart-case by default. :-) Or if you truly always want case insensitive by default, then `alias rg="rg -i"`.
"make ripgrep similar to grep" is certainly one goal, but there are other goals that compete with that and whether or not ripgrep matches grep's behavior with respect to case sensitivity is only one small piece of that.
It was proposed to enable --smart-case by default: https://github.com/BurntSushi/ripgrep/issues/178
I usually know when I want case insensitive match - and I frequently want case-sensitivity with all lowercase search terms. Eg: doing a search for methods or variables in ruby - where "model" is right, but "Model" is a class or module (and vice-versa - but then smart case would work).
While I do allow for breaking changes between major versions, I don't really see myself ever doing any kind of major breaking change. Especially one that subtly changes match semantics in a really major way.
By default, ripgrep does not search a bunch of files, in particular anything gitignored and anything with a leading dot (which the docs call "hidden"):
https://github.com/BurntSushi/ripgrep/blob/master/GUIDE.md#a...
These flags override that.
I wouldn't want to miss anything on my home directory or repository-specific dotfiles.
Edit: oh, this is different from --no-ignore.
The result is instantaneous and only found out recently it uses ripgrep from my co-worker.
[0]: https://github.com/sharkdp/bat
[1]: https://github.com/junegunn/fzf[0]: https://github.com/sharkdp/fd
[1]: https://tldr.sh/
For instance: 'fd .apk' to search for android builds.
As to find, its flexibility certainly comes handy on occasion, but...
find: paths must precede expression
Aaaargh. $ curl cheat.sh/tarI also have a folder of a bunch of curl commands that I can search and apply fuzzy finding on the results that helps me explore to find.
Contrived example, search for "t" in all these files and then pipe to fzf: rg t . | fzf
But with fzf[1], that whole workflow of searching through your history is probably two orders of magnitude faster. Now you hit CTRL-R and you start typing any random part of the command you're trying to remember. If there was some other part of the command you remember, hit space and type that search term after the first search term. FZF will then show you the last 10-ish matches for all the search params you just typed, AND it will have done all this with no UI lag, no hitching, and lightning fast.
I don't know what other people use FZF for, as this is the SINGLE feature that's so good I can't live without it anymore.
[1] - https://github.com/junegunn/fzf#key-bindings-for-command-lin...
git for-each-ref refs/heads/ --format='%(refname:short)' --sort='-authordate' | fzf +s --query "$*" | xargs git checkout [alias]
cof = !git for-each-ref --format='%(refname:short)' refs/heads | sort | uniq | fzf | xargs git checkout
cor = !git branch --list --remotes | sed 's@origin/@@' | sort | uniq | fzf | xargs git checkout
cord = !git branch --list --remotes | sort | uniq | fzf | xargs git checkout alias f = fd | fzyhttps://darrenburns.net/posts/tools/
On one hand, it’s awesome. It’s both excellent from a programmer’s standpoint and indispensable from a user’s standpoint. It’s just so neat it’s hard not to love it. A remote interviewer thought I just had a custom alias for grep -r and it was hard not to laugh and share the joy and relief that I didn’t.
On the other hand, it’s decidedly anti-Unix, arguably even more so than GNU utilities. The man page goes on for ages, and not because it’s written badly (it’s not), but because the tool does so much. I really wish you didn’t have to bundle all of that functionality into a monolithic binary to get this kind of performance.
For reference, walking directory trees was specifically cut out of the Plan 9 userland and left for du(1) because it was stupid how many implementations of it there were. Yes, Rust’s is in a library, but then you don’t have Unix as your OS, you have Rust instead. “Every language wants to be an OS”[1].
The plan9 comparison is kind of meh to be honest, because it never achieved widespread adoption. Popularity tends to beget complexity.
But otherwise this isn't so surprising to me. One of the hardest parts about making code fast is figuring out where to fall on the complexity vs perf trade off. It's pretty common that faster code is more complicated code. Directory traversal is a classic example. There are naive approaches that work in most cases, and then there are approaches that work well even when you throw a directory tree that contains hundreds of thousands of entries. And then of course, once you add gitignore filtering, you really want that coupled with tree traversal for perf reasons.
When you specialize, things get more complex. Then we build abstractions to simplify it and the cycle repeats. This to me seems more like a fundamental aspect of reality than even something specific to programming.
(Next up, an attempt at self-justification, followed by even more philosophizing.)
I don’t think it’s a totally accurate description, either, for what it’s worth :) The use of the GNU tools (and not, I don’t know, Active Directory) as a reference point was in part to signal that the comparison is intentionally horribly biased. In any case it’s more of a provocative synopsis of my sentiments than it is a criticism.
I hope you take the grandparent comment as the idle musings that it is and not as any kind of disapproval. My mixed feelings are actually more about what ripgrep tells us about the limits of pipes-and-streams organization than they are about ripgrep itself.
The tty prettyprinting problem (and the related machine-readable-output problem) is (of course) not limited to ripgrep; I don’t even know how to satisfyingly do this for ls. What people consider ridiculous can differ: I’d certainly say that egrep fork/execing fgrep in order to filter its input through it is ridiculous, but apparently people used to do this without batting an eye. Having a separate program whose only purpose in life is to colorize grep output (outside of some sort of system-wide convention) would be silly, though, and that’s is a much stronger standard. So no scorn from me here, even if I’d have preferred the man were a couple screenfuls shorter.
I’m not sure how much stock one should put in adoption. (If AT&T’s antitrust settlement expired a couple of years later, would we be using today?) Popularity begets complexity in various ways: handling of edge cases, gradual expansion of scope, fossilization of poorly factored interfaces. Not all of those are worthy of equal respect, except they’re not discrete categories. But I’m not evangelizing, either; only saying that Plan 9 has an unquestionable (implementation) simplicity aesthetic, so it seems useful to keep it in view when answering the question “how complex does it need to be?” Even though its original blessed directory walking ... construct, du -a | awk '{print $2}', is a bit kooky.
The case of traversal with VCS exclusions is fascinating to me, by the way. It looks like it begs to be two programs connected by a unidirectional channel, until you introduce subtree exclusion, at which point it starts begging to be two programs connected by a bidirectional channel instead, and the Unix approach breaks down. I’m very interested in a clean solution to this, because it’s also the essential difference between typesetting with roff and TeX: roff (for the most part) tries to expand complex commands first, then stuff the resulting primitives into the typesetter; TeX basically does the same, except there’s a feedback loop where some primitives affect the expansion state (and METAFONT, which is a much more pleasant programming environment in the same vein, doubles down on that). It seems that some important things that TeX can do and roff can’t are inevitably tied to this distinction. And it’s an important distinction, because it’s the distinction between having your document production be a relatively loosely coupled pipeline of programs (that you can inspect at any point) and having it be a big honkin’ programming environment centered around and accepting only its own language (or plugin interface, but that’s hardly better). I would very much to have this lack of modularity eliminated, and the walk-with-VCS-exclusion issue appears to be a much less convoluted version of the same even if it’s not that valuable in isolation.
Yup, all is well. :-)
My favorite example of trading performance for simplicity is `memchr`. I have a small little write-up on it here: https://docs.rs/memchr/2.4.0/memchr/#why-use-this-crate
The essence of it is that if you want to find a byte in a slice in Rust, then that's really easy:
fn memchr(needle: u8, haystack: &[u8]) -> Option<usize> {
haystack.iter().position(|&b| b == needle)
}
So why bother building a whole crate for this? Well, because, to make it fast, it takes low-thousands (to cover the variety of memr?chr{2,3} variants) lines of code to make use of platform specific SIMD instructions to do it.This is a good example of something the Plan9 folks would probably never ever do. They'd write the obvious code and (probably) demand that you accept it as "good enough." (Or at least, this is my guess based on what I've seen Rob Pike say about a variety of things.)
I have a lot of sympathy for this view to be honest. I'd rather have simpler code. And I feel really strongly about that. But in my experience, people will always look for something faster. And things change. Back in the old days, the amount of crap you had to search wasn't that big. But now that multi-GB repos are totally normal, the techniques we used on smaller corpora start to become noticeably slow. So if you don't give the people fast code, then, well, someone else will.
Anyway, none of this is even necessarily a response to you specifically. I'd say they are also just kind of idle musings too. (And I mean that sincerely, not trying to throw your words back in your face!)
> “how complex does it need to be?”
Yeah, I think "need" is the operative word here. And this is kinda what I meant by popularity breeding complexity I think. How many Plan9 users were trying to search multi-GB source code repos or walk directory trees with hundreds of thousands of entries? When you get to those scales---and you accept that avoiding those scales is practically impossible---and Plan9's "simplicity above everything else" means dealing with lots of data is painfully slow, what do you do? I think you either jump ship, or the platform adapts. (To be clear, I'm not so certain of myself as to say that this is what Plan9 never achieved widespread adoption.)
> The case of traversal with VCS exclusions is fascinating to me, by the way. It looks like it begs to be two programs connected by a unidirectional channel, until you introduce subtree exclusion, at which point it starts begging to be two programs connected by a bidirectional channel instead, and the Unix approach breaks down.
Yeah, I think this (and many other things) are why I consider the Unix philosophy as merely a guideline or a means to an end. It's a nice guardrail, and where possible, hugging that guardrail will probably be a good heuristic that will rarely lead you astray. That's valuable. It's like Newton's laws of motion or believing that the world is flat. Both are fine models. You just gotta know not only when to abandon them, but that it's okay to do so!
But yes, this has been my struggle for the past several years in my open source work: trying to find that balance between performance and complexity. In some cases, you can get faster code without having to pay much complexity, but it's somewhat rare (albeit beautiful). The next best case is pushing the complexity down and outside of the interface (like memchr). But then you get into cases where performance bleeds right up and into APIs, coupling and so on.
When you have a design guideline – in this case, "Do one thing and do it well" – then you must never forget that it's just a guideline, a means to an end, a tool to achieve your goal. If your guideline conflicts with or contradicts your goal, then the guideline is wrong – or should be ignored, at least. Otherwise, the guideline becomes an ideology, and ideologies can be harmful.
I see this too often with various guidelines ("do one thing and do it well", "never use gotos", "linked lists are bad"), which are zealously repeated without understanding whether they make sense in that particular context. (I'm not saying this applies to you, this is a general observation)
Of course, what's "new wave" to one person might be "mild evolutionary step" to another. (And perhaps even "evolutionary" implies too much, because not everyone sees smart filtering by default as a good thing.)
Kudos for your ripgrep achievements, though.
To be honest, I don't really know what you're after here. It is fine to say that the "new wave" is not rooted in Unix, but that doesn't mean its inaccurate to call ripgrep a Unix tool.
Now that would be new wave UNIX
https://www.oxfordlearnersdictionaries.com/definition/englis...
When you start quoting the dictionary to prove a point, maybe it's time to take a step back.
Seriously: the amount of time I’ve spent trying to figure out what something does after clicking an interesting HN headline just boggles me.
I feel like I want to start using ripgrep just because of that line.
BREAKING CHANGES: Binary detection output has changed slightly.
also noteable, an underlying lib got some vectorization speed improvements.
Going to install ripgrep ASAP
Can't believe I wasn't following BurntSushi on GitHub, what a track record.
[1] https://github.com/lotabout/skim [2] https://github.com/zenofile/fish-skim
I haven't looked into all the tools in that list, but I would not for example recommend exa over ls as it is simply not reliable enough: if a filename ends in a space, you won't see that in exa, and that bug has been reported for years. To me that is a clear blocker, and if it is still there, I simply cannot trust the file listing from exa, no matter how pretty it may look.
Anecdote: I once had to recover a system with a corrupted libpcre.so. This will break almost every standard gnutil. The easiest way to do it without a recovery OS was to use a few alternatives written in Rust, which don't have this problem because they statically link their dependencies (and cargo still worked, so it was easy to install them).
Either way, thanks for a quality tool, more, if it inspired others to come up with good rust tools.
With that said, Go tools were doing this before Rust even hit 1.0 as far as I remember. There was 'sift' and 'pt' for "smarter" greps for example, although I don't think either of them ever fully matured.
alias rg="rg --max-columns 200 --max-columns-preview"
So now if you hit a minified css or js file, it will truncate any match longer than 200 characters instead of flooding your screen with a million chars line.
This has reminded and motivated me to submit a donation to the project.
------------------------------------------
How can I donate to ripgrep or its maintainers?
As of now, you can't. While I believe the various efforts that are being undertaken to help fund FOSS are extremely important, they aren't a good fit for me. ripgrep is and I hope will remain a project of love that I develop in my free time. As such, involving money---even in the form of donations given without expectations---would severely change that dynamic for me personally.
Instead, I'd recommend donating to something else that is doing work that you find meaningful. If you would like suggestions, then my favorites are:
The Internet Archive
Rails Girls
Wikipedia
* [0] https://github.com/BurntSushi/ripgrep/blob/64ac2ebe0f2fe1c89...The -r option is handy as well for some search and replace problems (faster than GNU sed, literal search with -F, can use PCRE2 when needed, etc): https://learnbyexample.github.io/substitution-with-ripgrep/
---
I'm currently checking out another tool written in Rust - frawk (https://github.com/ezrosent/frawk) as an awk alternative
Yes it's an impressive achievement to make it that fast, but you could get the same or better performance by just using an index to search. I've been using GNU id-utils [2] for a long time, that is using an index, and it gets comparable performance for a fraction of the source code and brain power, and likely energy use too.
gcc -DHAVE_CONFIG_H -I. -D_FORTIFY_SOURCE=2 -march=x86-64 -mtune=generic -O2 -pipe -fno-plt -MT fseterr.o -MD -MP -MF .deps/fseterr.Tpo -c -o fseterr.o fseterr.c
fseterr.c: In function 'fseterr':
fseterr.c:74:3: error: #error "Please port gnulib fseterr.c to your platform! Look at the definitions of ferror and clearerr on your system, then report this to bug-gnulib."
74 | #error "Please port gnulib fseterr.c to your platform! Look at the definitions of ferror and clearerr on your system, then report this to bug-gnulib."
| ^~~~~
make[3]: *** [Makefile:1590: fseterr.o] Error 1
So I can't even easily get idutils to try it, but suspicion is that it misses some key use cases. And from what I can tell, its index doesn't auto-update, so you now have to use your brain power to remember to update the index. (Or figure out how to automate it.)ripgrep is commonly used in source code repos, where you might be editing code or checking out different branches all the time. An indexing tool in that scenario without excellent automatic updates is a non-starter.
[1] - https://aur.archlinux.org/cgit/aur.git/tree/PKGBUILD?h=iduti...
Parallelism applies both to directory traversal, gitignore filtering and search. But that's where its granularity ends. It does not use parallelism at the level of single-file search.
Longer story: Ripgrep uses Rust's regex library, which uses the Aho-Corasick library. That does not just provide the algorithm it is named after, but also "packed" ones using SIMD, including a Rust rewrite of the [Teddy algorithm][1] from the Hyperscan project.
[1]: https://github.com/BurntSushi/aho-corasick/tree/4499d7fdb41c...
(I should really add a link to that in the README.)
It's not a direct comparison with Rust's regex crate or with ripgrep, but this article shows the comparison with re2. Notice that on average it takes 140K worth of scanning to "catch up" with RE2::Set with 10 patterns - the situation would be even more marked with 1 pattern.
https://www.hyperscan.io/2017/06/20/regex-set-scanning-hyper...
Personally, I'm dissatisfied with the approach to regex scanning in Hyperscan (too heavyweight at construction and too complex) but not much more pleased by the Rust regex crate or RE2 (frankly, the whole compile-a-giant-DFA-as-you-go isn't that great either). I feel on the verge of taking another crack at the problem. Lord knows the world needs another regex library...
With that said, I have longed for simple ways of composing regexes better. I've definitely fallen short of that, and I think it's causing me to miss a whole host of optimizations. I hope to devote some head space to that in the next year or so.
What I'm thinking about lately is sticking a lot closer to the original regular expression parse tree when implementing things. Yes, that leaves performance on the table relative to Hyperscan, but I suspect the compile time could be extremely good. Also, it would be better suited to stuff like capturing and back-references.
Like I said, I suspect I'll be building "Ultimate Engine the Third" sometime in the not-too-distant future (Hyperscan's internal name was "Ultimate Engine the Second", an Iain M. Banks reference).
I'm guessing rg can be faster in general - because of less memory allocation/copying by sorting before outputting?
That --sort-files disables parallelism is not really a theoretical limitation.
Perhaps the difference is then explained by my choice of search term. The term I tried after upgrading must happen to appear early in the sorted corpus.
I just now tried it with a very rare term, and it does indeed take longer overall to complete the search.
I guess ripgrep might be worth trying. Anyone used both of these who can compare? I'd like to know what I'm missing out on.
https://github.com/BurntSushi/ripgrep/issues/875#issuecommen...
# Move deleted files to macOS user trash (safer)
trash () {
command mv "$@" ~/.Trash ;
}https://github.com/BurntSushi/ripgrep#quick-examples-compari...
Emacs is unusable in a modern dev environment with its defaults. You have to write lisp.