Hister – A private, full content search index that you control
hister.org
hister.org
My first free software search project was Searx, a privacy respecting metasearch engine, but because of the limitations of the metasearch concept, I've decided to take a different approach.
Hister builds a personal search index from pages you visit, bookmarks, browser history, local files, and crawled websites. It stores extracted content with offline result previews, so information remains searchable even when the original page changes or disappears. It supports full text and semantic search, can run entirely on your own machine, and includes a web interface, command line tools, and an MCP endpoint for assistant integrations.
Project page: https://github.com/asciimoo/hister
Tiny read-only demo: https://demo.hister.org/
would've been great with a more liberal license
Not in the near term. Right now I am focused on developing Hister rather than operating a hosted service. There are already plenty of centralized hosted search engines, so my longer term interest is in federation and distributed search. I want to make the core system mature first.
> would've been great with a more liberal license
It depends on how do you define liberal. =] I chose AGPLv3+ because I want Hister to remain free software and available to their users.
i still think it could've benefited by Apache 2.0 which more or less gets you to your goals
Immediately interested and will check it out, thank you! I've wanted a "search stuff you've seen online" tool for a long time, but everything seems to be research-oriented or "archive but don't search" or some weird combination that means it's nigh useless to me. I've got decades of bookmarks and archives and I've kinda been stuck grepping them at best (it's rare but I do sometimes want a page I saw once three years ago and I love having that option), while hoping someone would build something better.
One question if ya don't mind, while I explore: any chance of singlefile support? Content-extraction is useful in lots of situations (e.g. wallabag) and it's a great default, but sometimes it fails and sometimes you really do want the page, relatively close to how it actually was. Singlefile does that much better than most, and it does so well enough (and manually-handle-able enough if needed) that I don't feel any desire to switch to WARCs or similar.
Though specifically I'm probably looking for something like "content-extract everything" + "key combo to save singlefile version too" + "upload singlefile archives to backfill / recover". Like 99% of the time content extraction is preferred, and I'm glad to see it... it's just not always enough, and having to go elsewhere for exceptions breaks a lot of the utility.
> One question if ya don't mind, while I explore: any chance of singlefile support?
Yes, partially. Hister can already import HTML files created by SingleFile, but there is no direct integration yet. In the longer term, I would like the SingleFile extension to be able to send snapshots directly to Hister.
Overall I really like what I'm seeing, it ticks a lot of important boxes for me and it's pleasantly straightforward. Hopefully I'll find time to contribute!
Hister always stores the original material.
> That way you could also switch from an extracted view to a "full" view in the UI.
It isn't even needed, we just need a SingleFile specific extractor (an interface in Hister to parse specific page content and provide custom previews) that provides the full original HTML for the preview panel.
> Hopefully I'll find time to contribute!
I'd appreciate it. <3
Thank you again!
Would Hister support this basic workflow? I'd love to retire my own software.
The next phase was going to move to a recording proxy.
[0] https://github.com/canvas-ui/canvas/tree/main/apps/browser-e...
I often find myself irritated because I read an article on my phone 6 months ago and the history is gone.
Automatic page capture on mobile currently requires Firefox, since mobile Chrome does not support browser extensions.
Then joining the mobile device to tailscale network where Hister is hosted on a server on a host in the network … probably this will work at least that is my plan to try.
Thank you for an amazing project. I’m a heavy msgvault user and adding Hister to my browsing will be hugely complementary.
Happy to contribute the feature if needed
Would Hister be suitable for this? Can it index mbox files? Would it handle this amount of data? Does it have a search API so I can build an MCP server?
The indexer it uses (Bleve) can handle millions of records according to their docs, but sure that Hister would be the best choice for this task. I'd probably use Meilisearch (https://www.meilisearch.com/).
> Can it index mbox files?
Not yet.
> Does it have a search API so I can build an MCP server?
It has both search API and MCP server endpoints.
One thing at first look, like a little feature request ;).
Personally I think it would be very nice if it was possible to index different sites to different search indexs. So you could separate different stuff, like job, different projects you have, and other stuff and then last everything else etc to there on search index db.
So it would be possible to have like "profiles" you could easy choose from. So in the addon, when you click on the hister icon on the toolbar there would be a list of profiles that you could choose where to index the site to. But also a setting to choose which domains and sites that always should index to one search index profils. And all sites that dont is added to a profile is added to the default index instead.
For me this would make it much more easy to find stuff and sort out what Im looking for (for me things get very "thing"/project based what I want to find), so having this..
1 - The "profiles" would become like an important filter. When you search you could easy mark one or more "profiles" which you would search from, and you remove alot of unwanted data automatically, specially when you index every site you visit by automatically.
2 - It would also make it easy if the db gets big over time to remove data that is less important later on and that you dont need to be index anymore.
3 - It would also make it easier to backup only the most important search index DB:s if the DB for the differnt search index:s where in different folders/files. I guess the "default index" that index all sites could easy get big, but there it would be alot of not important data. So would be nice to be able to easy just use a backup program and only backup the only important search databases, to save space.
Thanks
Academics usually take it to mean the Danube river instead, but in conspiracy/New Age contexts, the association with Hitler is prevalent, and mentioned in several pop culture works, including at least one feature film.
I'm hacking on a in-process DB on top of LMDB+Lance(for now, hilbert space kung-fu with a custom matryoshka embedding setup with separate spatial + temporal + internal and content derived anchors will replace that ~last-century~ last-year tech) + roaring bitmaps as the primary indexing engine with the same or similar goal[0] and will definitely deep-dive into yours.
I'm also trying to index users unstructured documents and workflows(tabs, emails, files, notes, identities etc)
- Organize them into semantically meaningful user or agent created context or directory-like virtual trees (the same photo of a nice kitchen may be surfaced under `/travel/barcelona` and `/arch/interieour/kitches`)
- ..where tree nodes are mapped to bitmaps - `/travel/barcelona` does a fast and cheap `travel` AND `barcelona`, want to "zoom-out" you just go one directory up to `/travel` and see all documents tagged with travel)
- You can use multiple timelines - extract that fancy md-converted en-wiki hf dataset into a wikipedia db dataset + timeline, tag your personal timeline as "personal" - wanna know the zeitgeist of your grandmothers birth date - search for it with timelines personal + wikipedia in layered mode and you'll get everything that happened or was happening during that time.
- You can have long-running stateful query sessions and refine your searches dynamically - search for "winter" and get all documents with a winter scenery or mentioning winter - refine with "nice view" then "laptop" - citing a recent example[1]
- Documents have relations that would be cumbersome to map in a virtual tree structure(worth an experiment due to the zoom-in/out you get with context bitmap trees though) - hence on top of the initial structure you can use graph edges(also powered by bitmaps - as most indexes are)
- All vector queries always run on top of a candidate set you get by the bitmap/bitmap-based filter algebra hence searching through 100k+ docs is usually pretty fast
Anyhow, let me stop here, thank you once again!
[0] https://github.com/canvas-ui/canvas-synapsd (sorry for the sloppy ai readme, no time to resurrect my old one with the updated APIs)
I scraped and imported posts from the blogs I regularly reference for award travel, then hooked it up to OpenCode/Codex as an MCP server and used that corpus for research on those topics. So I can ask things like "has anyone ever mentioned running into this problem before?" [1]
If you have a hobby or working situation that requires you to regularly reference a core set of websites or reference materials, Hister provides almost all the tools out of the box to start a search engine against it. The default datasets they promote include the Python Stlib, MDN and RFC corpus, as an example. [2]
What tools or features would Hister need to support your complete search workflow?
The only thing that I think would be interesting to see is native support for crawling via a sitemap.xml instead of recursively. I worked around this by implementing a basic scraper that fetched pages exclusively from the sitemap.xml to add into Hister.
I think you're already aware of this, but I also experienced some data loss during the import because I was running a concurrent reindex. I clocked it pretty quickly so I didn't think too much of it. [1]
[1] "TODO store new documents in both indexes while running reindex to guarantee not losing any data." @ https://github.com/asciimoo/hister/blob/master/server/indexe...
It also has a "public mode" where anyone can search the indexed content, but only authenticated users can add or modify it.
Also indexed data is persisted on a per-user basis, so you got this isolation and certainty that your searches will not be polluted by your family's
https://github.com/rumca-js/Internet-Places-Database
Also I maintain android app that can be used to search places.
https://f-droid.org/en/packages/io.github.rumcajs.offlineweb...
I believe hister you have to fill in with your data, right?
- let us say I want to index every blog ever listed on HN
- should be a small subset of the 400 billion pages out there on the internet no?
- First I need to gather data, what do you use to load so many webpages rapidly? asyncio with aiohttp in python? are there better options?
- how do you handle proxies? rotation? are there libraries you recommend for this?
- what about pages that use cloudflare? or block your request or present a captcha or a challenge of some kind?
- what are the filetypes you collect? only html or media as well?
- where and in what format do you store all these collected files? flat file storage? duckdb? postgres? hstore? something else?
- what is the frequency at which you refresh each page? once a day? once a week? something else?
- what kind of pre-processing do you use on the collected data? remove extra spaces? special characters? some kind of complex regex pipeline? LLM?
- how do you match the incoming query with processed data? simple text matching? regex? vector embedding match? something else?
Others could be answered by examining the source code, as it is open source.
It's good etiquette to check existing documentation and resources before badgering OSS project teams with a lengthy list of questions.
Webtm.io
All open source and small enough to deploy. I deploy to cf webworkers so it’s the only place it’s tested.
One cool thing is we work on iOS, chrome and friends, Firefox and pretty much everywhere. We do require you bring your own LLM though.
I suppose the one thing I need would be an easy way to share links into Hister. Hopefully we will see both Android/iOS mobile apps to send links into Hister and better support for Safari. I wonder what could be possible with the extensions support available on mobile browsers too (Safari, MS Edge, Firefox).
For the semantic search is there any chunking/processing that happens with the content or do you need to be diligent about having a large embedding context (and/or small content)?
Planning to contribute significant improvements to the vector search side here.
Got enough Hitler in my life with Google...
I feel this is a weird spin on the "Clicks to Hitler" game
Where it says that “Hister” is a variant of its Latin name.
I think the name is lovely with no reason to change. Surely there are few people who would connect this with silly medieval superstition. Perhaps you can take part in wiping out that association.
It's a terrible choice of name
Unlike say Hipster with the descender on the p.