Rga: Ripgrep, but also search in PDFs, E-Books, Office documents, zip, tar.gz
phiresky.github.io
phiresky.github.io
Currently the main branch is undergoing a refactor to add support for having custom extractors (calling out to other tools), and more flexible chains of extractors.
Ripgrep itself has functionality integrated to call custom extractors with the `--pre` flag, but by adding it here we can retain the benefits of the rga wrapper (more accurate file type matchers, caching, recursion into archives, adapter chaining, no slow shell scripts in between, etc).
Sadly, during rewriting it to allow this, I kind of got hung up and couldn't manage to figure out how to cleanly design that in Rust. I'd be really glad if a Rust expert could help me out here:
In the currently stable version, the main interface of each "adapter" is `fn(Read, Write) -> ()`. To allow custom adapter chaining I have to change it to be `fn(Read) -> Read` where each chained adapter wraps the read stream and converts it while reading. But then I get issues with how to handle threading etc, as well as a random deadlock that I haven't figured out how to solve so far :/
I don't quite grok the problem here. If you file an issue against ripgrep proper with code links and some more details, I can try to assist.
Taken literally, ripgrep uses that exact same approach. There are potentially multiple adapters being used. Each adapter is just defined to wrap a `std::io::Read` implementation, and the adapter in turn implements `std::io::Read` so that it can be composed with others. The part that I'm missing is why this has anything to do with threading or deadlocks. I/O adapters shouldn't be having anything to do with synchronization. So I'm probably misunderstanding your problem.
Sorry, I don't think I explained my issue very well. In general it has nothing to do with the interaction with ripgrep, that works fine.
It's that each adapter (e.g. zip -> list of file streams) needs to have an interface of fn(Read) -> Iter<ReadWithMeta>
But then if there's a PDF within the zip, I have to give the returned ReadWithMeta to the PDF adapter - but it can't take ownership, because the Archive file iterators only give borrowed reads. I maybe worked around this by creating a wrapper type [3] and adding an unsafe here [2], but something deadlocks when adapting zip files currently.
Also, for external programs, I have to copy the data from the Read into a Write (stdin of the program) - which needs to happen in a separate thread, otherwise the stdout is never read [1], but some Reads I have aren't Send since they come from e.g. zip-rs, so they can't be passed to a thread.
[1] https://github.com/phiresky/ripgrep-all/blob/baca166fdab3d24...
[2] https://github.com/phiresky/ripgrep-all/blob/baca166fdab3d24...
[3] https://github.com/phiresky/ripgrep-all/blob/baca166fdab3d24...
Every name in the shadow is built from the name in the origin but maybe with ".txt" added (or .txt.gz if you want to keep the compressed with whatever is the fastest decompressor builtin to ripgrep as a library not called as a program). Untranslated names can be just symbolic/hard links back to the origin. Build rules become as flexible as your build system.
This also scales to deployments that have more disk space than memory. Admittedly, in that case, the whole procedure probably becomes disk-IO bound, but maybe not. Maybe some translations cannot even keep up with disk IO - NVMe storage is pretty fast, for example. Or available memory may vary dynamically a lot, sometimes allowing the shadow to be fully in the buffer cache, other times not. It strikes me as less presumptuous to assume you can find disk space vs. having that much memory available. (EDIT2: though I may be confused about how `rga` operates - your doc says "memory cache", though.)
On the pro-side, but for updating the shadows based on origins, the user could even just `rg` from within the shadow and translate filenames "in their head", although stripping an always present string is obviously trivial. Indeed, you won't need `rg --pre` at all and the grep itself could become pluggable. I doubt any of your other `fzf`/etc. integrations would be made more complicated by this design, either.
This all strikes me as simple/nice enough that someone has probably already done it...EDIT1: Oh, I see from thumbs ups and other comments over at [1] and [2] that @phiresky is probably already aware of this design idea, but maybe some HN person knows of an existing solution along these lines.
[1] https://github.com/BurntSushi/ripgrep/issues/978 [2] https://github.com/BurntSushi/ripgrep/pull/981
Any plans to integrate with skim, a Rust implementation of fzf?
The -bin suffix is an AUR convention to let you know that it's downloading a precompiled binary rather than building from source.
https://aur.archlinux.org/packages/?O=0&SeB=nd&K=-git&outdat...
https://aur.archlinux.org/packages/?O=0&SeB=nd&K=-hg&outdate...
For fun, I pointed a 12-core/32GB RAM 2018 MBP at a 9GB network share full of PDFs, while still using the laptop for other things (so not a benchmark, just an anecdote).
Initial cold/uncached run:
rga -j 12 testword share 1140.70s user 77.58s system 31% cpu 1:03:55.85 total
Cached:
rga -j 12 testword share 8.09s user 4.88s system 92% cpu 14.048 total
Cache after the run is 77M.
that way I can open a browser tab, wait 5 seconds for it to load, locate the new screen location of the search bar, click it, wait for javascript to finish loading so I can click the search bar, click it for real this time, mistype because there's some kind of contenteditable event jank, wait 5 seconds for my results to come up, fix the typo, and just have my results waiting for me
I'm not going to learn a new tool when web is fine
More !operators here - https://duckduckgo.com/bang
Firefox keyword search has one little known killer feature: You can combine it with data URIs and JavaScript to run small "command line snippets" stored in your bookmarks from your browser bar.
To get started, create a keyword search from any form (like the search bar on duckduckgo.com) and edit the URL of the entry in the bookmark manager to point to
data:text/html,<script>alert("%s")</script>
instead.What you can do with this is (fortunately) limited by cross-origin restrictions but there are some useful applications. For example, I use this snippet
data:text/html,<script>i="%s";firstSep=i.indexOf(" ");if(firstSep==-1)firstSep=i.length;subreddit=i.substr(0,firstSep);searchTerm=i.substr(firstSep+1);location=`https://old.reddit.com/r/${encodeURIComponent(subreddit)}${(searchTerm!=null&&searchTerm.replace(/\s/,"").length>0)?`/search?q=${encodeURIComponent(searchTerm)}&restrict_sr=1`:""}`</script>
as a nice shortcut to Reddit ("<keyword> <subreddit>" to jump to a Subreddit, "<keyword> <subreddit> <search string>" to search within a Subreddit).You can also insert content directly into the document which opens the possibility for instant marquee
data:text/html,<marquee>%s</marquee>[0] https://bugzilla.mozilla.org/show_bug.cgi?id=1625901
EDIT: That error occurred when going Right Click -> Add Keyword from the website. If setting the bookmark manually completely, it works.
And there are thousands more at duckduckgo.com/bang .
The best sarcasm lies on a ridge, you cannot tell if it's sarcasm or not.
This is in fact what's happened with Schrodinger's Cat: it was meant as an argument from absurdity against the Copenhagen Interpretation of quantum mechanics, but it's presented seriously and so people take it in that way.
The purpose of sarcasm is not to make an argument, it's to have fun. The best fun is had when exactly half of the audience does not get the joke (as the other half makes fun of them).
if god wanted me to access my files in less than 15 seconds, they wouldn't have commanded google to package the search bar as a separate JS bundle that only gets downloaded when you focus the search bar
I'm no frontend dev but I know a thing or two about HTML + there's no built-in way to input text into a box -- this is the best we can do and we'll just have to wait for 5G + moore's law to solve this
Death to SPA (Angular, React) Long live SPA (Mithril, Vue)
Because every obersable event you trigger has to go into an add machinery
hahaha, nice one (continued)
For GSuite/Workspace this needs to be enabled by an admin: https://support.google.com/a/answer/9121487?hl=en
Stop using cloud, usb is fine.
I’ve recently been playing with Recoll for full-text-search on content. Since it indexes content up front, the search is pretty fast. It can also easily accommodate tag metadata on files.
It would be interesting to consider how ripgrep based tools can fit into generically broad “search your database of content” workflows (as opposed to remember or go through your file system paths).
ls -sh ~/.cache/rga/
total 336M
336M data.mdb 4.0K lock.mdbI use the same system in Vim to browse source code. It's very powerful, very fast, works with any language and requires zero configuration.
Also, I have a specific pattern to write some tags inside files that I can parse with ripgrep.
You'll need to set NOTES_DIR in your environment to wherever you want your notes to be stored. Then you can write `note something` to create or open $NOTES_DIR/something.md with your $EDITOR.
If you type "note" without parameter you'll start a search on all the note names, ordered by last use. If you type "note -f" it starts a full text search.
For best results you should have the fzf.vim's preview.sh somewhere in your fs, otherwise it'll use "cat" but it won't be as good looking (see FZF_PREVIEW in the script).
Hopefully despite being shell it should be readable enough to tweak to your liking.
Note that it was written and used exclusively on Linux, but I did try to avoid GNU-isms so hopefully it should work on BSDs and maybe even on MacOS with a bit of luck.
From my `~/.gitconfig`:
[alias]
brt = "!git for-each-ref refs/heads --color=always --sort -committerdate --format='%(HEAD)%(color:reset);%(color:yellow)%(refname:short)%(color:reset);%(contents:subject);%(color:green)(%(committerdate:relative))%(color:blue);<%(authorname)>' | column -t -s ';'"
I always spent a lot of time being confused about branches, and never realised how easy the solution was. brt = "!git for-each-ref refs/heads --color=always --sort -committerdate --format='%(HEAD)%(color:reset) %(color:yellow)%(refname:short)%(color:reset) %(contents:subject) %(color:green)(%(committerdate:relative))%(color:blue) <%(authorname)>'"The closest I can find is mlocate but it does not have a GUI but more importantly it does not index my Windows or NTFS drives.
Would appreciate any suggestions if someone knows something like 'everything' for Ubuntu.
I just learned how to mount all my Windows drive under /mnt using (using the `disks` software), so hopefully this should index those files too.
From the linux command-line, I like fzf ( https://github.com/junegunn/fzf ), that you can instruct to use the faster fd ( https://github.com/junegunn/fzf#environment-variables ). Fzf even offers keybindings for your shell. For example, it binds Alt+C to fuzzy-finding a directory, and Enter cds to it ( https://github.com/junegunn/fzf#key-bindings-for-command-lin... ).
Fzf is great for other things too; here is a fish function to bing Alt+G to fuzzy-pick a Git branch and jump to it:
function fish_user_key_bindings
bind \eg 'test -d .git; or git rev-parse --git-dir > /dev/null 2>&1; and git checkout (string trim -- (git branch | fzf)); and commandline -f repaint'
bind \eG 'test -d .git; or git rev-parse --git-dir > /dev/null 2>&1; and git checkout (string trim -- (git branch --all | fzf)); and commandline -f repaint'
endAlso for gnome there is tracker which does search and indexing built into the system. I think by default its set for minimal use but it can be configured by the settings/search panel to index many locations. I haven't played with is much recently though.
Just wondering since Linux knows about these drives but doesn't mount them automatically at startup.. so if this is out of a reason or just convention?
Hmm.. I seem to remember creating an excel file for this client a while back.. open Everything -> filter client.xlsx .. boom. Or maybe I didn't name it properly, at all? Well still just a simple '*.xlsx' and sort by date, I can generally find anything this way. As long as you let Everything open on windows startup, it will be instant after use.
I'm the author of plocate. Since version 1.1.0, plocate no longer depends on mlocate for building its database, but is a full replacement.
For general file search at ludicious speeds like Everything does on windows its pretty good :)
The search in PDF viewers is an anti-feature in terms of UI and performance. Their advantage is that they allow to scroll to and highlight the found phrase back in the document.
$ sudo dnf install -y ripgrep-all
[...]
No match for argument: ripgrep-all
Error: Unable to find a match: ripgrep-all
Rust's package manager fails: $ cargo install ripgrep_all
[...]
failed to select a version for the requirement `cachedir = "^0.1.1"`
candidate versions found which didn't match: 0.2.0
location searched: crates.io index
required by package `ripgrep_all v0.9.6`
Quick search on the web shows that more people have problems with cachedir version.There is a github issue to make this the default behaviour of cargo, but you miss out on updates which might fix security bugs so the cargo team is unwilling to change the default.
https://github.com/phiresky/ripgrep-all#integration-with-fzf
Do you have a link for that? That's news to me.
1. https://poppler.freedesktop.org/releases.html
pdftotext -layout file.pdf | grep -E ...
for PDFs, good to see a Swiss Army knife utility for all sorts of file though!ripgrep-all can do the same regexes as rg on any filetypes it supports. So you can could do something like --multiline and foo(\w+[\s\n]+){,20}bar
It won't work exactly like this, but something similar should do it:
--multiline enables multiline matching
* foo searches for foo
* \w+ searches for at least one word character
* [\W]+ searches for at least one space/nonword character like sentence marks
* {,20} searches for at most 20 iterations of the word-space combination bar searches for bar
Once that's done, you have all the options available to perform that search. But I don't know of a search tool that does the OCR for you. I did read a blog post of someone uploading PDFs to google drive (they OCR them on upload) as an easy way to do this.
I answered this a while back: https://old.reddit.com/r/rust/comments/c1bjw4/rga_ripgrep_bu...
Something perhaps more helpful but so far unmentioned (and somewhat OS-specific) is that statically linked executables usually fork & exec (especially exec) much faster than dynamically linked ones. This difference is usually only like 50..150 us vs 500..3000 us but can multiply up over thousands of files.
This only matters on the first run of `rga`, of course. While the dispatched-to decoder is likely mostly out of one's linking control, this overhead can be saved for the dispatcher, at least. So, I would suggest `rga-preproc` should have a static linking option/suggestion, at least on Linux.
Of course, this overhead may also fall below the noise of PDF/ebook/etc. parsing, but maybe not the decompression of small files in some dark horse format. :-)
I've also re-run my original set of benchmarks[1] with ugrep included: https://github.com/BurntSushi/ripgrep/blob/master/benchsuite...
I'm currently not using any of ugrep or rga, although I have used pdfgrep in the past. It'd be nice for casual users like me to know more about why I should use rga over ugrep (or vice-versa).
dir *.c* | rg somesubstringinfilenameIt looks like rga can handle SQLite out of the box, so just making sure your history .db file is visible to rga may be all you need.
You can also use my Datasette tool to get a web UI against your history, see https://docs.datasette.io/en/stable/getting_started.html#usi...
https://www.lesbonscomptes.com/recoll/
A PDF in a Zip file, in an email attachment. recoll can index it and do OCR if you like
I can understand it might be nice to have a personal library of PDF books and searching in them. I can't think of a time I've ever wished I could search my bookshelf in that way, but you never know.
Obviously I use tools like ripgrep for searching codebases and the like.
But the extreme flexibility of this one in particular (and others like MacOs Spotlight) makes it seem more like a data recovery tool for me. If my directory structures and databases ever completely failed for some reason I might need to search through everything to find the data again. It's good to know such tools exist, I suppose.
But my fear is that tools like this teach people to not worry about organisation of data and to just fill up their disks with no structure at all. I think that unless something goes terribly wrong nobody should ever need a tool like this. Once you rely on it, you're out of luck it if it ever fails you. What if you just can't remember a single searchable phrase from some document, but you just know it must exist somewhere?
It's similar to what Google has done to the web. When I was growing up it used to be a skill to use the web. People used tools like bookmarks and followed links from one place to another. Now it's just type it into Google and if Google doesn't know, it doesn't exist.
But I don't think this tool deserves the same sort of mixed feelings. I don't think this replaces structure -- there's still value to having a conceptual mapping of where documents are stored, and for grouping sets of documents together. It's just that having a structure doesn't help if you don't know where in the structure something is stored. This sort of tool is a bottom-up approach for the times when the top-down approach doesn't work very well.
Do you have similarly mixed feelings if sometimes, even with my carefully-crafted set of bookmarks with all their nested folders, I use the search tool to find the bookmark I'm looking for? It's the same idea. Sometimes a top-down structure is beneficial. But sometimes things get misclassified, or you forget about some piece of the structure, or you aren't familiar with some new structure, and in those cases, having bottom-up tools are immensely useful. There's no risk of vendor lock-in here. It's just a difference of approach in information retrieval.
It's more intuitive to simple search for something in the something you are looking for and clicking it.
I haven't used a folder organization structure in many many years. Other than the defaults for my cloud folders and a separation between Personal + Work.