Building an Open Source Decentralized E-Book Search Engine
github.com
github.com
https://github.com/JakeKalstad/IPFSPytorchDataset https://github.com/JakeKalstad/load_ipfs_pytorch_model
It is open source and they're always looking for contributors. I think they'd especially welcome help improving search!
Is there any project working on this?
I believe Google documented some of this in its early days, noting that a search index returns the relevant metadata matching a specific query. The query space itself is largely based on both raw keywords and tuples (2- or 3-word ngrams if memory serves, though I'm hazy on this), the latter meeting some minimum frequency requirement. Longer search terms can be constructed from shorter ngrams.
A typical advanced native-tongue English vocabulary is about 40,000 words. An expansive dictionary might contain fewer than 250,000 words, including obsolete ones.
Mapping a vocabulary to works citing those words is relatively straightforward. Ngrams experience combinatorial expansion, but are still a reasonably constrained space. And we now have well over a quarter-century's experience indexing written content at Web scale.
A laptop could probably make a decent cut at providing a useful index of many millions of books, though you'd probably want a somewhat larger system for a more comprehensive index, in particular to rank-index the search space, which is probably the more considerable challenge.
I've been doing some local LLM stuff at work recently, and even with the amazing advances in quantization lately, doing that kind of stuff on a ThinkPad is feasible, but still strongly inferior to just renting out a VPS with a couple 4090/H100s for several hours.
The biggest thing with summarizing stuff is that most local LLM models often don't have very big context-windows, so they have trouble with larger texts like even a short Vonnegut novel (I was just testing em' with summarizing GitHub issues, and even with a 16k token context window they still sometimes struggle if there are a lot of comments).
There are probably smarter people than I who could get this working on a Raspberry Pi though... ;)
I will build an open sourced version too!
There is an open sourced version for torrent searching here, using the same tech.
"I was recommended ... Liber3 ..., which uses ENS domain names ... running on ENS and IPFS ... they appear to be using Glitter ... a ... service built with Tendermint."
This sounds like a signal from outer space to me. In a language used in a different galaxy.
I tried that Liber3 thing, but whatever I do, I get "Oops! Something went wrong. Please refresh or try again later".
What is this all about?
The blockchain ecosystems really are their own little world unto themselves. It's all pretty cliquey, not in an exclusionary way but if you're not actively seeking it out then there's very little chance of you hearing about any of it.
Side note: IPFS is well worth checking out if you're interested in databases or decentralized zero-trust systems, and even if you're a blockchain skeptic. They're doing some really interesting work under the hood. The team hasn't latched on to the gold rush mentality the way so nearly all blockchain projects have.
The bulk of the article is implementation details, helpfully hyperlinked.
I'm trying to encourage publishers and authors to offer legitimate sales of DRM-free ebooks, so would prefer we try not to have the term "ebook" associated with piracy.
I actually understand your point well but I think it's even more important not to group in any legitimate use of technology with illegitimate use of it. Especially considering recent events (lawsuits over Yuzu and Dolphin emulators).
What I was commenting upon is that I think this particular thing will actually appear to many people to be for piracy of ebooks, and I don't want "ebook" to become synonymous with "pirated book" in the mind of the public (and especially not in the minds of publishers and authors who I want to encourage and support making non-DRM ebooks).
As a compromise, I'd like people whose efforts are intended for pirating books to distinguish that from legitimate ebooks, not call it simply "ebooks". Maybe then it'll be easier for our respective "freedom and access" goals to coexist.
(FWIW, I'm actually sympathetic to some of the uses of pirated books, such as sharing the wealth of the world's information with people who just can't afford it, or who have it officially denied to them. I'm less sympathetic to people who could pay for something but choose to take it instead, but I'm not trying to combat that here. I just want to mitigate some of the bad effects of piracy, such as legitimate buyers only being able to get DRM'd books, and only for consumer-hostile locked-down devices. Please don't inadvertently sabotage that orthogonal effort by appropriating language.)