https://darrenburns.net/posts/tools/
On one hand, it’s awesome. It’s both excellent from a programmer’s standpoint and indispensable from a user’s standpoint. It’s just so neat it’s hard not to love it. A remote interviewer thought I just had a custom alias for grep -r and it was hard not to laugh and share the joy and relief that I didn’t.
On the other hand, it’s decidedly anti-Unix, arguably even more so than GNU utilities. The man page goes on for ages, and not because it’s written badly (it’s not), but because the tool does so much. I really wish you didn’t have to bundle all of that functionality into a monolithic binary to get this kind of performance.
For reference, walking directory trees was specifically cut out of the Plan 9 userland and left for du(1) because it was stupid how many implementations of it there were. Yes, Rust’s is in a library, but then you don’t have Unix as your OS, you have Rust instead. “Every language wants to be an OS”[1].
The plan9 comparison is kind of meh to be honest, because it never achieved widespread adoption. Popularity tends to beget complexity.
But otherwise this isn't so surprising to me. One of the hardest parts about making code fast is figuring out where to fall on the complexity vs perf trade off. It's pretty common that faster code is more complicated code. Directory traversal is a classic example. There are naive approaches that work in most cases, and then there are approaches that work well even when you throw a directory tree that contains hundreds of thousands of entries. And then of course, once you add gitignore filtering, you really want that coupled with tree traversal for perf reasons.
When you specialize, things get more complex. Then we build abstractions to simplify it and the cycle repeats. This to me seems more like a fundamental aspect of reality than even something specific to programming.
(Next up, an attempt at self-justification, followed by even more philosophizing.)
I don’t think it’s a totally accurate description, either, for what it’s worth :) The use of the GNU tools (and not, I don’t know, Active Directory) as a reference point was in part to signal that the comparison is intentionally horribly biased. In any case it’s more of a provocative synopsis of my sentiments than it is a criticism.
I hope you take the grandparent comment as the idle musings that it is and not as any kind of disapproval. My mixed feelings are actually more about what ripgrep tells us about the limits of pipes-and-streams organization than they are about ripgrep itself.
The tty prettyprinting problem (and the related machine-readable-output problem) is (of course) not limited to ripgrep; I don’t even know how to satisfyingly do this for ls. What people consider ridiculous can differ: I’d certainly say that egrep fork/execing fgrep in order to filter its input through it is ridiculous, but apparently people used to do this without batting an eye. Having a separate program whose only purpose in life is to colorize grep output (outside of some sort of system-wide convention) would be silly, though, and that’s is a much stronger standard. So no scorn from me here, even if I’d have preferred the man were a couple screenfuls shorter.
I’m not sure how much stock one should put in adoption. (If AT&T’s antitrust settlement expired a couple of years later, would we be using today?) Popularity begets complexity in various ways: handling of edge cases, gradual expansion of scope, fossilization of poorly factored interfaces. Not all of those are worthy of equal respect, except they’re not discrete categories. But I’m not evangelizing, either; only saying that Plan 9 has an unquestionable (implementation) simplicity aesthetic, so it seems useful to keep it in view when answering the question “how complex does it need to be?” Even though its original blessed directory walking ... construct, du -a | awk '{print $2}', is a bit kooky.
The case of traversal with VCS exclusions is fascinating to me, by the way. It looks like it begs to be two programs connected by a unidirectional channel, until you introduce subtree exclusion, at which point it starts begging to be two programs connected by a bidirectional channel instead, and the Unix approach breaks down. I’m very interested in a clean solution to this, because it’s also the essential difference between typesetting with roff and TeX: roff (for the most part) tries to expand complex commands first, then stuff the resulting primitives into the typesetter; TeX basically does the same, except there’s a feedback loop where some primitives affect the expansion state (and METAFONT, which is a much more pleasant programming environment in the same vein, doubles down on that). It seems that some important things that TeX can do and roff can’t are inevitably tied to this distinction. And it’s an important distinction, because it’s the distinction between having your document production be a relatively loosely coupled pipeline of programs (that you can inspect at any point) and having it be a big honkin’ programming environment centered around and accepting only its own language (or plugin interface, but that’s hardly better). I would very much to have this lack of modularity eliminated, and the walk-with-VCS-exclusion issue appears to be a much less convoluted version of the same even if it’s not that valuable in isolation.
Yup, all is well. :-)
My favorite example of trading performance for simplicity is `memchr`. I have a small little write-up on it here: https://docs.rs/memchr/2.4.0/memchr/#why-use-this-crate
The essence of it is that if you want to find a byte in a slice in Rust, then that's really easy:
fn memchr(needle: u8, haystack: &[u8]) -> Option<usize> {
haystack.iter().position(|&b| b == needle)
}
So why bother building a whole crate for this? Well, because, to make it fast, it takes low-thousands (to cover the variety of memr?chr{2,3} variants) lines of code to make use of platform specific SIMD instructions to do it.This is a good example of something the Plan9 folks would probably never ever do. They'd write the obvious code and (probably) demand that you accept it as "good enough." (Or at least, this is my guess based on what I've seen Rob Pike say about a variety of things.)
I have a lot of sympathy for this view to be honest. I'd rather have simpler code. And I feel really strongly about that. But in my experience, people will always look for something faster. And things change. Back in the old days, the amount of crap you had to search wasn't that big. But now that multi-GB repos are totally normal, the techniques we used on smaller corpora start to become noticeably slow. So if you don't give the people fast code, then, well, someone else will.
Anyway, none of this is even necessarily a response to you specifically. I'd say they are also just kind of idle musings too. (And I mean that sincerely, not trying to throw your words back in your face!)
> “how complex does it need to be?”
Yeah, I think "need" is the operative word here. And this is kinda what I meant by popularity breeding complexity I think. How many Plan9 users were trying to search multi-GB source code repos or walk directory trees with hundreds of thousands of entries? When you get to those scales---and you accept that avoiding those scales is practically impossible---and Plan9's "simplicity above everything else" means dealing with lots of data is painfully slow, what do you do? I think you either jump ship, or the platform adapts. (To be clear, I'm not so certain of myself as to say that this is what Plan9 never achieved widespread adoption.)
> The case of traversal with VCS exclusions is fascinating to me, by the way. It looks like it begs to be two programs connected by a unidirectional channel, until you introduce subtree exclusion, at which point it starts begging to be two programs connected by a bidirectional channel instead, and the Unix approach breaks down.
Yeah, I think this (and many other things) are why I consider the Unix philosophy as merely a guideline or a means to an end. It's a nice guardrail, and where possible, hugging that guardrail will probably be a good heuristic that will rarely lead you astray. That's valuable. It's like Newton's laws of motion or believing that the world is flat. Both are fine models. You just gotta know not only when to abandon them, but that it's okay to do so!
But yes, this has been my struggle for the past several years in my open source work: trying to find that balance between performance and complexity. In some cases, you can get faster code without having to pay much complexity, but it's somewhat rare (albeit beautiful). The next best case is pushing the complexity down and outside of the interface (like memchr). But then you get into cases where performance bleeds right up and into APIs, coupling and so on.
When you have a design guideline – in this case, "Do one thing and do it well" – then you must never forget that it's just a guideline, a means to an end, a tool to achieve your goal. If your guideline conflicts with or contradicts your goal, then the guideline is wrong – or should be ignored, at least. Otherwise, the guideline becomes an ideology, and ideologies can be harmful.
I see this too often with various guidelines ("do one thing and do it well", "never use gotos", "linked lists are bad"), which are zealously repeated without understanding whether they make sense in that particular context. (I'm not saying this applies to you, this is a general observation)
[0]: https://github.com/sharkdp/bat
[1]: https://github.com/junegunn/fzf[0]: https://github.com/sharkdp/fd
[1]: https://tldr.sh/
For instance: 'fd .apk' to search for android builds.
As to find, its flexibility certainly comes handy on occasion, but...
find: paths must precede expression
Aaaargh. $ curl cheat.sh/tarI also have a folder of a bunch of curl commands that I can search and apply fuzzy finding on the results that helps me explore to find.
Contrived example, search for "t" in all these files and then pipe to fzf: rg t . | fzf
But with fzf[1], that whole workflow of searching through your history is probably two orders of magnitude faster. Now you hit CTRL-R and you start typing any random part of the command you're trying to remember. If there was some other part of the command you remember, hit space and type that search term after the first search term. FZF will then show you the last 10-ish matches for all the search params you just typed, AND it will have done all this with no UI lag, no hitching, and lightning fast.
I don't know what other people use FZF for, as this is the SINGLE feature that's so good I can't live without it anymore.
[1] - https://github.com/junegunn/fzf#key-bindings-for-command-lin...
git for-each-ref refs/heads/ --format='%(refname:short)' --sort='-authordate' | fzf +s --query "$*" | xargs git checkout [alias]
cof = !git for-each-ref --format='%(refname:short)' refs/heads | sort | uniq | fzf | xargs git checkout
cor = !git branch --list --remotes | sed 's@origin/@@' | sort | uniq | fzf | xargs git checkout
cord = !git branch --list --remotes | sort | uniq | fzf | xargs git checkout alias f = fd | fzyOf course, what's "new wave" to one person might be "mild evolutionary step" to another. (And perhaps even "evolutionary" implies too much, because not everyone sees smart filtering by default as a good thing.)
Kudos for your ripgrep achievements, though.
To be honest, I don't really know what you're after here. It is fine to say that the "new wave" is not rooted in Unix, but that doesn't mean its inaccurate to call ripgrep a Unix tool.
Now that would be new wave UNIX
https://www.oxfordlearnersdictionaries.com/definition/englis...
When you start quoting the dictionary to prove a point, maybe it's time to take a step back.