Is there any project working on this?
Is there any project working on this?
I believe Google documented some of this in its early days, noting that a search index returns the relevant metadata matching a specific query. The query space itself is largely based on both raw keywords and tuples (2- or 3-word ngrams if memory serves, though I'm hazy on this), the latter meeting some minimum frequency requirement. Longer search terms can be constructed from shorter ngrams.
A typical advanced native-tongue English vocabulary is about 40,000 words. An expansive dictionary might contain fewer than 250,000 words, including obsolete ones.
Mapping a vocabulary to works citing those words is relatively straightforward. Ngrams experience combinatorial expansion, but are still a reasonably constrained space. And we now have well over a quarter-century's experience indexing written content at Web scale.
A laptop could probably make a decent cut at providing a useful index of many millions of books, though you'd probably want a somewhat larger system for a more comprehensive index, in particular to rank-index the search space, which is probably the more considerable challenge.
I've been doing some local LLM stuff at work recently, and even with the amazing advances in quantization lately, doing that kind of stuff on a ThinkPad is feasible, but still strongly inferior to just renting out a VPS with a couple 4090/H100s for several hours.
The biggest thing with summarizing stuff is that most local LLM models often don't have very big context-windows, so they have trouble with larger texts like even a short Vonnegut novel (I was just testing em' with summarizing GitHub issues, and even with a 16k token context window they still sometimes struggle if there are a lot of comments).
There are probably smarter people than I who could get this working on a Raspberry Pi though... ;)