Thanks!
The metadata system is a nightmare.
To start, we did everything manually.
We source the book cover from the publisher, create author entries, etc.
Eventually, we started paying Nielsen a lot of money to use their Book metadata API. It is "ok," and they don't update it often. But it helps us automatically pull in an author's name, book title, genres, and age-group.
We still manually source the book covers as they only have super small and blurry covers. And we screen every book we add to make sure the data is correct.
What is especially frustrating?
- Author names are text and not linked in any way. So we have to decide what is a slightly different name but the same author and what is a different author.
- The BISAC genre standard is full of errors and abuse by publishers. For example, they might tag Dune as "AI" which marks it also as being nonfiction because they don't know how the BISAC standard they created works :).
- It lists book editions and has no concept that all book editions belong to one book.
- No real concept of a book series.
- Terrible book descriptions where publishers put in all kinds of reviews and nonsense that we need to figure out how to scrub eventually. They also abuse weird symbols to make it stand out.
It requires a lot of work to fix and manage all of this.
I am about to redo the entire topic/genre system due to some of these problems (this winter).
I am hoping to build a database of all books to use with new features in 2025. I don't know what we are going to do here. We might license a full DB of books from Ingrams (expensive) or Bowker. I liked Ingrams, but Bowker didn't email me back for months and gave me a lot of worry about working with them in the future. I might just do the best I can or break down and use Amazon's API (lots of stipulations in using it).
Some thoughts here of what we've built so far to manage this:
https://build.shepherd.com/p/a-big-focus-for-2024-improving-...
Hit me up any time to chat (ben@shepherd.com).