Lunr.js - Simple full-text search in your browser
github.com
github.com
In fact, if you had poked around, you probably could have snagged about 90% of this code from various projects ... too bad I didn't put it together like you did.
Ah well ... internet fame points to you I guess.
For example, the ability to search for a paragraph that contain two non-contiguous words would be very useful, but no browser (that I know of) is able to return elements that contain a set of tokens.
All browsers do is return exact matches from a string, with no concept of words.
It would be interesting to know if in this solution the index can be persisted to file or if it has to be rebuilt every time?
I haven't read through all of Lunr's docs and source, but based on my Solr/Elasticsearch experience, I'd expect to see (in time)…
Tokenization and (presumably) term normalization/analysis; a faster and smarter query language, for term order independence and boolean combinations of clauses; relevance scores and maybe even score boosting per field.
Better queryability really shouldn't be understated here. Just having term order independence focused on a specific set of JSON is going to be way better than naively matching any substring on the entire rendered page.
In the future I want to add even more powerful querying, restricting search to specific fields, taking into account the distance between terms, and adding faceted search to reduce the total documents being searched over.
One of the original goals of the project was specifically to provide a better alternative to just using the browsers built in find-in-page functionality
Why? This is a red flag.
I'm wondering how efficient this would be given that indexing a lot of data via javascript might really not be a good idea..
edit: looking at the docs it's unclear if it's possible. I guess index should be a JS object, so it's pretty simple to save it to a disc and then fetch it from client.
If indexing performance starts to become an issue the whole search index can be moved into a web-worker, which prevents indexing from blocking the rest of the page.