Recoll: A desktop full-text search tool
lesbonscomptes.com
lesbonscomptes.com
I wonder about e.g. Baloo (Baloo is the file indexing and file search framework for KDE Plasma, https://community.kde.org/Baloo). Is this similar? Or different? How is it different?
Just yesterday my dad complained about a couple of Baloo error messages, Baloo crash reports, disk full, Xsession log full of Baloo messages. And he was not really aware about what Baloo actually is.
The problem is similar to the one here: https://askubuntu.com/questions/1214572/how-do-i-stop-and-re...
So I disabled the content indexing, as it was suggested there. The stack trace of the Baloo crashes also suggested that it was related to indexing the content of some documents.
I think this is a problem for many such file indexing services. Even Spotlight on MacOSX frequently crashes when it tries to index some documents. Does Recoll handle this differently?
If Recoll is in general better than Baloo, maybe it should replace Baloo? Or Baloo can be extended to use Recoll under the hood?
Baloo used to use Xapian under the hood, like Recoll, when initially launched, but in 2014 (I think?) I moved away from it to a custom DB built on top of LMDB. It resulted in a massive performance increase, one of the main criteria was for results to always take less than 'x milliseconds' (I can't remember exactly how much) so that KRunner (similar to OSX Spotlight) would always feel snappy. Another criteria was for Baloo to operate well with minimal memory and CPU, which wasn't the case with Recoll (Things might have changed since then)
I do remember ensuring that Baloo always did content indexing in another process, though I can't remember at what point it blacklists a file.
Sorry, this isn't the best answer to your question, but most of the details are now quite fuzzy to me, and I haven't kept up with how the project has changed after I stopped contributing to KDE.
In retrospect though, I definitely should have spent far far more time ensuring that Baloo doesn't end up in situations like the one you described, and putting far more limits on the IO and CPU usage it's allowed to use.
It's very fast and ~4 lines of code. It's surprising how often I rediscover old blog posts & papers that are much better than what Google yields me.
From my experience Recoll isn't very good at searching for aliases sadly.
https://thesephist.com/posts/monocle/ written in personal language, takes multiple doc sources
https://apse.io interesting because it takes periodic screenshots and performs OCR
https://keminglabs.com/finda/ rust, fast, seemingly abandonded
I use Xapian on the desktop but I only store text documents so I have no need for Recoll. It is refreshing to use search that is not tuned for "popularity" and selling online ad services. If search is too easy, then IMO it is not really search. Search should require some skill. It should not be a game where a commercial entity Hoovering up user data and selling ad services runs the search and tries to guess what the searcher is searching for.
I never tried it, but this looks somewhat similar to Recoll:
Windows File Explorer can do this to some degree. I am not sure if File Explorer looks at the actual content of the documents or just titles. Either way, Recoll is magnitudes better and is absolutely worth it for my case.
https://www.devontechnologies.com/apps/devonthink
https://download.devontechnologies.com/download/devonthink/3...
It's not free, but one of the utils I wouldn't be without.
it is incredibly useful in itself, has a CLI/scripting interface etc, but its homemade GUI is quirky and using its output for next steps in workflows can be cumbersome (may need coding). better integration with file managers and other desktop apps would help it really takeoff
It seems to depend on pdftotext on Linux. Can I get that for Windows? If so, is an executable somewhere in my path enough?
However if you want to OCR a PDF which contains image based 'pages' then you'll need something like Tesseract OCR.