Xapian: Open source search engine library
xapian.org
xapian.org
I dream about a similar thing that can do OCR on scanned docs and extract text from my also sprawling library of epub and mobi files. If someone builds something like this, with maybe a LOCAL LLM to extract text descriptions from photos and movies as well as indexing metadata for everything, subtitles from movies and lyrics for songs, and add that to a NAS appliance, it’d be a killer.
I wrote a scraper to download all of the California EdCode from the governments site, convert them all to txt docs and I can ask questions about California EdCode in plain English.
I work in a shared governance capacity that requires us to refer to the Ed code for contractual negotiations and it’s been extremely helpful.
The big AI players are probably already scraping the bottom of this "barrel" in their search for training data, I am sure ...
Is there a way how to run curiosity.ai fully offline, without an account on your servers?
It's only part of what you want, but ocrmypdf will add a OCRed text layer to PDF files, making the text selectable and indexable
Linux is GPL too, didn't hinder companies making trillions on top of it.
Maybe not exactly the same, its a server that you can store documents and then retrieve their ID using a search string.
Xapian is more like sqlite while elastic would be mariadb
https://xapian.org/history https://sigir.org/files/forum/S2000/MUSCAT_note.pdf
Cybernetics OR steering filename:Heidegger ext:pdf
It's an absolute power tool.
Title Year Author Name.pdf
Same benefits as you mentioned. You can also filter by time that way.
year__author1_author2_author-n~~title-of-book~subtitle##tag1#tag-2#tagn.pdf
This means that files are automatically organised by year of publication,that I can search by tag name, and that I dont have to escape chars in the terminal. One day I hope to get round to building an Emacs mode to filter by the different elements.
I'd say my favourite thing about Xapian is that it's just a simple library you can embed in any app, no need for a separate database and JVM tuning. For simple usecases and small-to-medium datasets, it just works.
It's been trouble free and very performant, a real workhorse.
Imo people should use cross-platform alternatives.