Recoll – Full-text search for your desktop
lesbonscomptes.com
lesbonscomptes.com
The recoll UI could be improved and in general the integration e.g. with python scipting or other tools could be made easier but this is a much appreciated project and it is good to see it keeps being developed.
In an ideal universe projects such as this should converge with other desktop apps to create truly empowering information tools (repatriating the agency that has been relinquished to the "cloud")
what is quite handy usability-wise: incremental index updates (after inserting new files) are fast and can be done on the fly while fully using the desktop
Another Xapian-based tool I use is notmuch (email) and that one is very snappy too.
This last option makes Searx/SearxNG useable for all types of searches, both local as well as remote. I've been using this exclusively for many years now over a large collection of documents (about 600.000 entries) with good results.
[1] https://github.com/Yetangitu/recoll-webui
> It seems that Recoll will sometimes find data that Spotlight misses (especially inside pdfs apparently, which is probably more to the credit of poppler than recoll itself).
But TIL that some software already exists - right on the iPhone’s Camera app: https://support.apple.com/guide/iphone/use-voiceover-for-ima... Use VoiceOver for images and videos on iPhone. We are almost there!
https://www.voidtools.com/faq/#does_everything_search_file_c...
If I recall correctly, Firefox keeps your browser history in sqlite, seems trivial to add your own "doctype" in order to support it, as it's open source (https://framagit.org/medoc92/recoll) and written in a modular way (check the pdf handler as an example: https://www.lesbonscomptes.com/recoll/usermanual/usermanual....)
If you have something doing this to every page you visit, and Recoll can see it, then Recoll can index it.
Regarding automatically saving every page you visit, there's multiple tools that do this. One I played with and liked - but I can't remember the name - it's 5 numbers and refers to a port you can type with localhost to search through all recorded pages. That or something like it would work really well with Recoll.
https://github.com/jeremy-compostella/pdfgrep/pull/8#issueco...
It doesn't index contents just filenames so it is fast.
Searching within file contents seems to lack good options right now. Historically you had X1 desktop search, Google had a desktop search product, I think copernic. But most seem to be out of date.