I really enjoy the service. Mostly go through all the top books and see if anything calls out to me, and so far the site has done really well at keeping the informal contract of "I give you my email, you don't spam me".
The exception will be items from sellers of products not available on Amazon. And there may be some communities in Reddit that are more loyal to another e-commerce platform or site you would miss (crafters to Etsy, antique collectors to eBay, etc.)
1. Get a few thousands book titles that are "long enough" so that the probability they appear in a sentence without meaning the book is low
2. Search these spans in HN comments and use this corpus to train Stanford's CoreNLP NER
3. Run the NER on all comments
4. Check on Openlibrary or another book DB that the extracted spans are real books titles
I'm not an expert, but I think that a boostraping technique doesn't imply continually improvement of the model.
> Semi-supervised learning may refer to either transductive learning or inductive learning. The goal of transductive learning is to infer the correct labels for the given unlabeled data.
Agreed on bootstrap, but in the proposed approach you're not artificially expanding your sample size by sampling with replacement.
I'm in the process of adding Goodreads and other book websites to get better suggestions.(Using Amazon API is limiting on scale)
As for your question: I've been researching python natural language processing to get probable named entities and check them against Amazon API. I suggest https://spacy.io that has reasonable named entities extraction. However doing it at a scale might produce lot of books that are named as a common phrases.
I would look at the OpenLibrary dataset - you might be able to match titles in there to comments, or use it to validate the NER output, if you don't want to go the Amazon link regex route.
The entire dataset is available for download, or you can build a prototype with their API - I did this to map speakers to books with https://www.findlectures.com (you can see it if you hover over a name - e.g. https://www.findlectures.com/?p=1&speaker=-Barack%20Obama).
"name of book" by "author_first_name author_last_name"
If you could determine the most common words written before the title begins (read, liked, loved, recommend, etc.) you could probably parse out a lot of titles plus their authors.
NER will give you a lot of entities that you need to resolve against something to check if they're actually books anyway, depending on volume you may not find a free API tier.
Ex. If I write a comment saying "The number of our clients went from zero to one instantly", in this sentence, "Zero to One" will match Peter Theil's book.
If you are willing to put up with such "noise" in the output, then you don't need to train a thing or even use ML. Just chunk the given piece of text and look up all the tokens permutations of given length (from 1 up to N) in your SQL Entities Database.
You can generally find this kind of Entities Databases by aggregating a lot of datasets from all over the web or even use something like google books datasets: https://books.google.co.uk/
https://storage.googleapis.com/books/ngrams/books/datasetsv2... - This is an ngram set but there should be one with only book titles somewhere. Can't find it atm but you can search for it yourself!
You could even (if you are brave enough) try to use wikipedia's dumps to mine book titles from articles