Show HN: Instantly search 28M books from OpenLibrary
books-search.typesense.org
books-search.typesense.org
Thank you @mekarpeles for helping me get access to the data quickly and giving me pointers about the schema.
I should add - I built this in about 12 hours as a weekend project, so there might be some lurking issues.
> Details about the Tech Stack:
The dataset has ~28.6 million books and is indexed on Typesense [1], an open source alternative to Algolia/ElasticSearch that a friend and I are working on.
The UI was built using the Typesense adapter for InstantSearch.js [2] and is a static site bundled using ParcelJS.
The app is hosted on S3, with CloudFront for a CDN.
The search backend is powered by a geo-distributed 3-node Typesense cluster running on Typesense Cloud [3], with nodes in Oregon, Frankfurt and Mumbai.
Here's the source code: https://github.com/typesense/showcase-books-search
[1] https://github.com/typesense/typesense
[2] https://github.com/typesense/typesense-instantsearch-adapter
I'm going to try the free tier of typesense right now, perfect for my current use-case of site-search.
How large is the dataset for the books? How large nodes are needed?
The OpenLibrary dataset has ~28M records, takes up 6.8GB on disk and 14.3GB in RAM when indexed in Typesense. Each node has 4vCPUs. Took me ~3 hours to index these 28M records. Could have probably been done in ~1.5hrs - most of the indexing time was because of cross-region latency between the 3 geo-distributed nodes in Oregon, Mumbai and Frankfurt.
Is the indexed size 6.8GB or 14.3GB?
Well done!
There are some issues with ranking and duplicates that I see. I typed "Neal Stephenson" and got lots of duplicates for his latest novels. Cleaning up the data might be necessary to fix this. It seems biased towards newer/recent titles.
The ranking is also a bit tricky. If you type "Orwell" you get top results about George Orwell, rather than the book that made him famous (1984). I suspect this will be an issue with many literature works that will have a lot of books talking about these books.
Spelling correction is kind of working but surfaces a lot of noise as well. E.g. Neal Stevenson just surfaces a lot of titles and authors with either Neal or Stevenson in them. But impressively it did find one Neal Stephenson title.
Building good search and ranking is hard. So, not bad for a 12 hour effort.
I didn’t seem to find a popularity metric in the dataset, which would have solved the issue you pointed out about literary work having books written about them.
I’ve left the typo correction settings to be a little loose, so need to reduce its aggressiveness. I’ll fix that.
I love to see this, more access for the book community -- especially how quickly it came together. Nice work Jabo and excited to see Typesense in action.
Open Library gladly & thankfully tweeted its support of your hard work:
https://twitter.com/openlibrary/status/1338531822536310789
In every intention of being additive to Jabo's work, thought I'd share a few resources folks may not be aware of:
1. Open Library's Full-text Search
Major hat tip to the Internet Archive's Giovanni Damiola for this one: Folks may also appreciate the ability to full-text search across 4M of the Internet Archive's books on Open Library: https://blog.openlibrary.org/2018/07/14/search-full-text-wit.... You can try it directly here: http://openlibrary.org/search/inside?q=thanks%20for%20all%20...
As per usual, nearly all Open Library urls are themselves APIs, e.g.: http://openlibrary.org/search/inside.json?q=thanks%20for%20a...
2. Open Library's search API
By adding .json to our search endpoint, you can use Open Library's json search endpoint (which also contains bookcover links). You can also use this to fetch books by ISBN:
http://openlibrary.org/search.json?q=isbn:9781400067824
3. Book Cover Service (has issues)
We need to improve our cover service, it's definitely been neglected, but the cover IDs from our data dump, solr search index, and book APIs can be plugged in to our cover store to point to or fetch book covers. We don't yet have a current data bulk dump of covers available (it's on the roadmap) but if you need "all" the covers, please email us openlibrary@archive.org instead of hitting covers 1-by-1. https://openlibrary.org/dev/docs/api/covers
4. The new (unannounced) Open Library Explorer
I'm honestly kind of hoping no one reads this far into the comments because we haven't announced this one yet and I'm sure there are major performance issues (it's beta), but cdrini recently added dewey + library of congress classifications to our search index... Making this masterpiece possible:
https://openlibrary.org/explore
The ability to browse Open Library digitally, with all the glory of a physical library.
Please do contribute to OpenLibrary in any way you can. Here are some ways how: https://news.ycombinator.com/item?id=25408744
I'm always sort of surprised by this stuff. Like, isn't 99% of the effort to get access to the database, how to search it efficiently, how to display useful snippets? Why go to all that effort and not take a useful search syntax off the shelf?
> I built this in about 12 hours as a weekend project, so there might be some lurking issues.
> Yup, exact matching using quotes is on the horizon. Should be available in a few weeks.
> Looks like I’m filtering a little too aggressively which is affecting the results, I’ll take a closer look it in a few hours.
Working on support for exact matches using quotes, which will also help here.
Suggestion: Link the red book title to the OpenLibrary page.
I just pushed out an update to link the title to the OpenLibrary page.
This "Instant Search" doesn't improve on that.
Also, some of the Amazon links are broken for me. They look like https://www.amazon.com/s?k=9798654289605
Do you mean this option doesn't work through the search.json API? If so, can you please make sure we have an issue open https://github.com/internetarchive/openlibrary/issues/new/ch...?
You should be able to add has_fulltext:true to your query (which at least returns a set of books for which there is some full-text in existence:
e.g. http://openlibrary.org/search.json?q=has_fulltext:true%20AND...
@rayrag, there are researchers, vendors, libraries, partners like wikipedia, authors, and hundreds of thousands of people who use Open Library to track when they're reading, to promote books and reading, and to create curated lists.
Also, organizations like Internet Archive use such catalogs to manage what books to acquire (based on interest/demand from the community).
Having a home for bookcovers, citations, quotes, ISBNs, publishers, years, subjects, related editions are all valuable even if the book (source) is not yet available.
Hope this helps!
Google Books does this, not sure how many books though.
https://openlibrary.org/works/OL7116092W/The_church_of_Chris...
And https://openlibrary.org/works/OL2525391W/Holiness?edition=
I notice many titles did not have authors. I spot checked this one: https://openlibrary.org/books/OL25434821M/Still_Star-Crossed
This title has no author in the search results, but does have one on the linked page. Perhaps an import issue.
Edit: Forgot to say "This is awesome. Particularly for 12 hours work!"
Searching from Perth, Australia.
Just added a little debounce in any case. Hopefully that helps.
And I don't think there's any "average" you could easily advertise for without misinforming.
That said, I’m planning to add some examples and pre-configured configs to that page so it’s easier to get a good picture.
https://openlibrary.org/works/OL2964952W/Unintended_conseque...
https://books-search.typesense.org/?b[query]=Unintended%20Co...
Typesense does support searching by phrases on the server-side.
If you click through to the OpenLibrary dataset, you can see that the data model understands that the variations are editions of the same underlying text.
There's another dataset called "works" which has one record for all editions of a book, but it didn't have all the data points I wanted. So I decided to keep it simple for a weekend project and use the editions dataset.
Has anybody built it? No :(.
That would totally be the dream though. No reindexing. Pg_search is def fast enough for a large swath of use cases.
What I do is I’ve built a real barebones search controller that connects to Postgres. My controller only lets you search on one attribute, and I use it for autocompletes. Anything more involved, I use typesense.
What I really want is an integration between Instantsearch.js, and Postgres.
Typesense has written a custom adapter for Instantsearch & Typesense. I would love to just hook Instantsearch up to directly to my rails all. You could use pg_search to get 90% of the value, with next to no complexity on the backend.
You can "poll" for new records for indexing by using a query, but handling deletes/updates is the real deal breaker. You will then have to use both the binlog (atleast for mysql) and querying, at which point it becomes quite complicated to reason about.
For eg: you could hook into your ORM and send change events to Typesense like another commenter pointed out; if you have an event bus, you could hook into those messages and push to Typesense from a listener; you could use something like Maxwell to get a JSON stream from the DB binlogs and write the data you need into Typesense; or you could write a batch script which runs a set of SQL queries on a schedule and batch export + import into Typesense.
Also, typically you'd want to curate the fields that you put in a search index, vs what you store in your transactional database, so there is some level of data transformation that's needed anyway.
So long story short, it depends on your architecture!