Show HN: WorldBrain – full text, local search of your browsing history
worldbrain.io
worldbrain.io
Chrome is just horrible when it comes to this, and I can never get back to previous pages when searching via the addressbar.
The sad part is, that option just opens the built-in History page, just like pressing Ctrl + H, which I believe that's not what we wanted at all.
I think Chrome just redefined "full" while it regularly deletes history entries older than three months.
And yeah, ^H isn't fun(/useful).
Almost can't imagine and in my test just now, that did not work.
On the other hand, Chromium (the open-source version of Chrome) does this too so maybe it’s simply a UX decision rather than an evil plot to send more data to Google, though I personally believe that it’s the latter…
Say there's two websites I visit regularly with similar names? If I usually load one after typing three characters, but load the other one after typing four characters, Firefox will present the former first in the first case, and the latter first in the second case.
I don't know why it was removed (after my time) but I do know it had a lot of problems.
1) Most users were unaware of the feature, and it's hard to find a way for them to discover it.
2) The index costs a lot of disk space, which makes #1 worse. There's a whole bunch of tradeoffs around how much storage to use vs how much history to keep. Pages that self-reload can cause the index to bloat endlessly (a real bug we had).
3) Having an index of pages locally is not sufficient to make a useful search engine. There's a lot of ranking involved in making google search good. Similarly we tried to show "snippets" of page text that showed why we showed the results we did and that itself requires a lot of effort to be useful.
#2 and #3 are just bugs that can be fixed with more effort, but it's hard to motivate that effort in the presence of #1.
(FWIW, the sibling comment about how it was removed to favor google.com searches isn't plausible to me, that's not how the Chrome team works.)
On the one hand, when I have a bookmark for a site like HN, I want it to give me a quick way to get to the front page and see the latest posts so I can read and discuss them. In this case, the bookmark is serving like a pointer to an address whose contents are subject to change.
On the other hand, I bookmark pages with specific information I want to keep so I can refer back to it later. This is where bookmarks often fail me. Websites are transient by nature and bookmarks are extremely vulnerable to link rot. What I think I'd really want is a bookmark that saves an archive of the page so that I'll always have access to the information in its original form, even if the page changes or the site goes down. In this case, the bookmark is functioning more like a constant than a variable.
I'm aware that browsers such as Safari have the ability to save a page as a web archive file which includes all the data needed to render the page in its original form. The problem is that this feature completely punts on the issue of managing a set of these web archives, delegating that task to the Finder. I want a UI that is more integrated into the browser. If I bookmark a site it should save and manage the archive automatically. Full text search on all my bookmark archived pages should be built into the address bar. Perhaps in order to avoid the conflict between the two uses of bookmarks described above there could be a difference between favourites and bookmarks.
I don't know, does anyone else have any thoughts about this stuff?
thanks for speaking my mind here :)
We built Memex for that combination of use cases in mind. So right now you can already full-text search your bookmarks, and filter by time, domain and tags. Already on the mid-term roadmap we plan to enable full-html/text snapshots of visited pages, both locally and on-demand. For the latter we are potentially working with the Internet Archive.
I still don't understand why firefox refuses to slightly gray out their background by default. Chrome puts less strain on my eyes and that's far more important than anything firefox does well.
Every time I try to stick with firefox, my eyes get strained and eventually I go back to chrome.
Would this extension be interesting to anyone? It would be very simple, open source, and have no middle man. It would send links directly to a google CSE via their API.
The only solution I'd accept is one where all data is stored and indexed locally and there are good guarantees that deleted entries are actually deleted and/or wiped from the indices.
My issue with these services is (and maybe this one is different) is that I have to run 3rd party software and extensions and I never know how long these companies/products will stay around.
Yeah as mentioned the tool allows you to do bookmark full-text search. It's open-source, so it will stay no matter what. We build for resilience and for us as the WorldBrain.io company not needing to stick around in order for the service to survive. We see ourselves as the stewards of this tool, not the sole benefactor or proprietor.
Hope it will be helpful to you :)
https://softwarerecs.stackexchange.com/q/46270/16751
While I've been waiting for this, I've been using 'Export History (2.2)' browser extension in Chrome. This saves your history to CSV or JSON.
This extension looks like it could finally make it convenient enough to be more commonly useful, and privacy focused enough that I'm willing to try it. Great work! Are you planning to charge something for this in the future, or for extra features? I would definitely be willing to pay for it.
Oli here, from the team developing Memex. Memex is open-source, so the browser extension will always stay free to use.
What we will charge for are some of the services that require us to host stuff. Like backups, multi-device syncing, API calls etc. We will run it as a completely modular pricing model, where you can upgrade on only those features you need. We don't like those usual 3 tier model, where you have to upgrade to the 'monster mega plan' in order to just get one feature :) But you can also completely self-host that, as we will make the server software open-source as well.
Hope Memex can be useful to you. We are running a crowdfund to support its development, where we offer some good discount on the future features in return: worldbrain.io/pricing
Let me know if I can be of more help.
[0] http://www.future-perfect.co.uk/grammar-tip/is-it-focussed-o...
I prototyped something roughly like this several years ago. I wrote a simple Firefox plugin that communicated with a locally running server written in Closure with a Clojurescript web app for browsing that used the same server backend. I stopped working on the because services like Evernote do a better job, at the loss of some privacy.
Edit: I didn’t intend to imply that Evernote reads or uses user data.
This seems like it's local search. Are you saying you wish it was open source? The only cloud feature seems to be something about highlighting.
We custom built a search technology on IndexedDB and Dexie.js, which is capable of indexing around 5 years of your personal web-research locally in the browser.
Otherwise, I think this is super!
It starts to get icky when you notice the devs have a business plan to outsource the indexing to the cloud, but there's a commitment to keep the servers parts open source too, so that you can self host.
Yeah indeed, having stuff in the cloud is not ideal when it comes to privacy and centralisation.
One of our core values is privacy and data ownership. So we do our best to make our business not dependent on (analysing and selling) your data, and instead provide you with service value you're willing to pay for. We are built with interoperability in mind, that will allow you to switch providers of Memex and Memex Cloud without frictions, in case there are breaches of trust, or simply better service.
We follow these values by currently building for offline first usage, where your data is locally indexed and searchable primarily. With our search technology you'll be able to get up to 5 years of your research done in the browser.
For the cloud part, it is unfortunately not yet possible to do performant search on encrypted data, otherwise it would not be such a big issue to have your index in the cloud. Equally unfortunate is that it comes with a lot of drawbacks to replicate all your data on all nodes, as opposed to have a central point to query. Especially when we are looking at phone usages. There it is really not practical, so there is a need to have some sort of cloud - UX is still very important. Most people can't be bothered with the drawbacks of decentralised and distributed systems (yet). We hope to get that switch in multiple smaller steps that guide (non-technical) users through a smooth transition to a Memex system that is as distributed as possible. (Check out Dat https://datproject.org/, a technology we likely use to make that first step possible)
And as you already noted, this stuff will be self-hostable. We see ourselves as a service provider first and want to serve people who can't/don't want to run their own server. A bit like the Wordpress model.
You can read more about our approaches to running this business in our vision post: worldbrain.io/vision
I was wondering about mobile usage too, and agree that having a server somewhere available for queries is the best solution. Since this deals with such private data, having a possibility for self-hosting is the correct solution.
That will most likely return a lot of hits for sites I didn't like and would like to forget about.
Yeah indeed, just full-text search can let you end up with a lot of garbage. This is why it is so important, that you can search for various other "vague memories" to narrow down your search.
What you often remember about an article is stuff like: Did I bookmark it, when did I visit it, did I like/share/cite it on social media You can already filter by time, tags, domains, bookmarks, and soon also if you liked/shared or even seen it in your newsfeed, or on a friends wall, on Twitter and Facebook.
We gradually expand it so you can search with as much of your associative memories as possible.
Any research on that?
Info on RAM/CPU/Storage, but no info on energy pressure.
Good news: that's the plan! worldbrain.io/vision_deck