Show HN: 40k HN comments mentioning books, extracted using deep learning
hacker-recommended-books.vercel.app
hacker-recommended-books.vercel.app
I built this small app in my spare time to aggregate books recommended on Hacker News. I personally find books recommended on HN to be super helpful, so I think this is the way that I can contribute back.
This book aggregation idea is not new. A bunch of sites have done similar things [1, 2, 3].
Yet one common limitation of those sites is that they have limited recall (i.e. not able to get a comprehensive set of book mentions), and thus don't paint an accurate picture of what the top books are. They're all based on insufficient rules, e.g., looking for Amazon Links. As you can see from my app, people often do not include Amazon links when recommending a book.
I wondered, why can't we just match book names? Well, not so easy. Some books have pretty short names, e.g. Meditations [4], or Steve Jobs [5]. Some book name might as well be the name of a movie, e.g. Ready Player One [6]. Simply matching the names of the books would produce a whole lot of irrelevant results.
This is where Deep Learning comes into play. Recent advances in large NLP models (transformers and BERT in particular) have made machine language understanding unprecedentedly accurate. It enables me to fine-tune a BERT model on a couple thousand labeled HN comments and predict accurately whether each word in a comment is part of a book or not - a task commonly termed as Named Entity Recognition (NER).
As a result, my app is able to present a whole lot more results while maintaining desirable accuracy. For example, NER works pretty well on the tough examples I mentioned ([4, 5, 6]). Compared to prior sites, my app captures 9-50X more mentions and thus presents a much more complete picture of what books are recommended on HN.
Furthermore, I've made sure that the comments are presented well in the UI because the recommendations are just as useful as the books. I highlighted the mentioned book name, and used a custom NLP-based ranking function to sort the comments. These are non-trivial improvements over prior sites, which I hope you can find useful.
Nevertheless, this app is not without limitations: 1) matching book names would fail when two books have the same or similar names; 2) although not often, this approach would wrongly classify some short stop-word names [7] and 3) sometimes NER fails to see that the commenter actually hates the book. These problems can be alleviated with more Deep Learning. For 1), one can use BERT to learn the authors mentioned which can be used as a filtering criteria. 2) and 3) should be fixable with more training data (currently there are only ~4,000 hand-labeled HN comments).
Lastly, I'd like to especially thank my gf who helped me label ~1,000 comments, which boosted the model accuracy by 5 percent! I also want to thank the people who create and maintain the HackerNews big query dataset [8]. And of course, thank everyone on HN who recommends books to others.
Hope you enjoy this app! Feedback and suggestions are welcome :)
[1] https://news.ycombinator.com/item?id=15169611
[2] https://news.ycombinator.com/item?id=10924741
[3] https://news.ycombinator.com/item?id=12365693
[4] https://hacker-recommended-books.vercel.app/category/0/all-t...
[5] https://hacker-recommended-books.vercel.app/category/1/all-t...
[6] https://hacker-recommended-books.vercel.app/category/0/all-t...
[7] https://hacker-recommended-books.vercel.app/category/12/past...
[8] https://news.ycombinator.com/item?id=19304326
P.s. The amazon links are NOT sponsored. This app is free of monetization.
Anyone technically minded used FTP back in the days. So why is there a need for Dropbox (at that time).
That is what makes it hard to invest if you think too much or "know" too much. You get blinkers that prevent you seeing what become obvious successes because of UX improvements for the non-technical crowd.
Knowing the subject well may put blinders in some situations… but what’s the alternative? Know the subject poorly, flail about wildly and hope you land something by pure chance? Obviously knowledge isn’t the problem here; you need it to qualify if it’s a good decision or not. It’s the over-specialization, combined with the lack of empathy for the average user that derived the dropbox event.
I'm going to take extreme issue with your use of the word "often" here.
IMO it's not a "bad surprise" to a user that an amazon link is an affiliate one, it's just annoying when the only way you can get information on a book is through affiliate links.
I'd go even further. It's a loud minority that even cares at all about this and they make it feel like everybody agrees with them, but most people don't care and you're totally within your rights to do it. You should go for it!
Clicking through an affiliate link associates future orders for 24 hours, and once an order is placed no further orders are associated. (If an item is added to the cart while associated with an affiliate link it stays associated for 90 days, but it's not clear to me what information is provided to the associate regarding other items in the order if that item is eventually purchased.)
https://affiliate-program.amazon.com/help/node/topic/GPTZ495...
You mentioned transformers and BERT for large NLP models. I've been playing around with this too and it's a really powerful approach. Have you used spacy-transformers? [0]
The approach is pretty cool and can be used with BERT, GPT-2/Hugging Face etc.
I'm just starting to experiment with GPT-J and thinking of trying this approach also [1].
Anyway, totally awesome project and the results are really good. This stuff really is almost unreasonably effective!
Replies to your pinned top comment seems disabled, so allow me to ask here (sorry for the hijack):
> used a custom NLP-based ranking function to sort the comments
Can you expand on this function? (I'm familiar with NLP and most SOTA models)
Though if you wanted to try other things just for fun, maybe:
- Count matching NER, comments with a lot of book recommendations tend not to detail why they like the specific one currently filtered
- maybe down weight comments that are too short (after the matched NER is subtracted) as they seem to just have a title+link?
Not ML but while I'm here, on Android the comment section appear fine but has an horizontal scroll that seems spurious, with lots of blank space to the right (Galaxy S10, chrome & firefox)
On labeling, if you have a method statement or some go-by referance I am sure you would get some support here - I know I would help ! Maybe package a few blocks of 100 unlabeled comments with a readme & see what happens ?
Edit: another one that is tough is Open by Agassi - seems most of these comments do not actually have anything to do with the book. I would guess most one-word titles will have similar issues.
What is the level of sentiment analysis in natural language processing? Would it be easy to add the feature, to recognize whether the book was mentioned in a positive or negative light?
i.e "Guards Guards by Sir Terry Pratchett is a great book" vs. "I've never read anything as slow and uninteresting as The Two Towers by J.R.R. Tolkein" or "I thought Seveneves by Neal Stephenson was good - but it probably should've been two separate books with the second half actually having some meat to it."
https://hacker-recommended-books.vercel.app/category/15/all-...
Not a criticism! Sentiment analysis seems to remain an unsolved problem.
See also
https://news.ycombinator.com/item?id=28598341
https://news.ycombinator.com/item?id=28596882
Edit: oops, my title edit ("Show HN: 40k books on HN extracted using deep learning") was inaccurate. It's 40k comments, not 40k books. Fixed now.
> I have not yet read the good book Atlas Shrugged but be sure to check it out based on your recommendation.
You're delusional. Where did I ever recommend reading
Atlas Shrugged? Ayn Rand is nuts.I am amazed that the app can find any relevant comments mentioning "It" by Steven King, for example. That seems like magic. A few IT and It's are in there too, but still impressive.
I was intrigued to see "Twilight" quite high on the list of fiction, doesn't seem like something the HN demographic would read. Turns out most comments were about the board game / video game Twilight Struggle, or Zelda Twilight Princess, and those comments mentioning the book were not doing so in a particularly positive way.
"One Hundred Years Of Solitude" is often commented as "100 Years Of Solitude" so the system can miss a few comments there.
For an improvement, it'd be handy if the comments that mention other books had a way to link you to that book, so you could see other comments. For example, if you click Dune, and see someone saying "yeah, I liked Dune, and also x, y and z", then those other books should be clickable too, so you could see what other people say about them. Seems like the information could already be there in the database as a list-like comment could potentially appear in many book recommendations, but who knows.
I did something similar with RoBERTa and my own Kindle library to graph (with D3.js) all mentions/citations between my books (which books cites another books I have). I sorted the final graph by publication date to see some cool historical patterns of books citing another older books [1]
I also manually annotated ~1000 book mentions, but I combine RoBERTa with string search (I list all titles I want to search a priori) to reduce the number of false positives. I also augumented the dataset with thousands of books titles and metadata from goodreads.
I explain all the process on a blog post[2]
[1] https://thiagolira.blot.im/_projects/book_graph/main.html [2] https://medium.com/mlearning-ai/graphing-citations-between-b...
The medium post is amazingly written! I basically did the same thing - and you beat me with the data augmentation piece. I tried using nlpaug [0] but it didn't improve the model performance. I'll definitely try swapping book titles around.
A few years ago I found an article that was something like '100 short books everyone should read before they're 40'. It was a mix of fiction and non-fiction. I've never been able to find it again! But I really liked the list because these are books you can consume in a few hours and may be life changing.
I remember a few of the titles: Games People Play, Meditations, The Prince, The Art of War. (I suppose it may have been non-fiction only, although I think The Awakening may have been on there.)
Wish I could find the link again.
I've learned a lot from HN. But it wouldn't be good to fool myself into thinking that an employer wants to fund my personal development in this regard. Otherwise, they'd pay me to HN all day.
The crux of the issue is that it's impossible to work 8 hours every day. We all invent lies to fill the downtime.
Would your employer pay you to HN all day? If not, precisely how much of your day are they comfortable with you HN’ing? Are you sure it’s officially approved?
Waiting for Compiles is the usual, there's a lot of waiting in software - waiting for compiles, scripts to run, someone else to do something.
And the problem with Stephenson is that's rarely succint so a bad book from him turns into a huge loss of time.
Most comments I’ve read here indicate the last part of the book is the most unpopular, so perhaps I should go back someday and read the middle.
Of course some might argue those are the best things about his books, but while Chryptonomicon is perhaps my favorite sci-fi novel I've read, I think it takes a certaim type of person to enjoy a book where the plot gets interupted for 5 page descriptions on how to eat Captain Crunch or a who fucked who of the Greek pantheon.
1. I regret you earned $0 for helping me spending so much on books. Have you considered setting up affiliate links or a donation button? Maybe affiliate links as a service will be your next project.
2. The Amazon links are for Amazon.com, but I'm in Canada. Maybe easy internationalized Amazon affiliate links will be your next project.
Considering I already buy books on Amazon, if there's anyway I can just find an affiliate (any affiliate), Amazon gets 5.5% less revenue.
For tracyhenry, they would get ~$8.25 CAD straight out of Amazon's pocket for my $150 purchase.
https://associates.amazon.ca/help/node/topic/GRXPHT8U84RAYDX...
(edit) just scroll down that page: https://pi-hole.net/donate/#sponsorship
People should get paid for work.
Whether that work is having a job. Or making a website.
I don't see the difference...
Someone who does useful work deserves wages.
People should get paid for work.
Whether that work is having a job. Or making a website.
Someone who does useful work deserves wages.
Even 2,000 years ago the Bible said: "the worker deserves his wages."
Most of the people who are again monetization are perfectly happy to get paid by their employer.
Is direct employment the only morally upright way to receive payment for hard work?
You helped me spend $150 on books!
Check your local Library. Depending on where you are, it could be a fantastic resource for books.Note also that addall.com used book search doesn’t include thriftbooks in the results, so I just always go straight to thriftbooks and don’t bother searching addall.com anymore..
No affiliation, just a longtime happy customer.
Step 1: make elites addicted to drugs
Step 2: monopolize drug trade
Step 3: install a religious fundamentalistic regime with yourself at its head
(All very logical until this point, but next step might be a problem, can anyone offer advice)
Step 4: transform into a worm
??!
Looking forward to the criticism that results from mapping the co-ordinates of this ontology, as one could weave a narrative around most of the books that aggregates them into types and categories themselves, then transmit the criticism without the substance to some believers, which could codify into an "anti-HN" ideology (which is just a peculiar form of fan club.) Calling "hacker-critical" as a pseudo academic backlash trend now, and a "hacker studies" course designed to encircle these ideas with criticism as levers to manage people who have them. Really, if you aren't using AI to create predictive levers about people's beliefs and behaviors to manage and extract value from them, what are we doing with it. :)
Super cool to create this though, as it would be really interesting for other comminities, potentially subreddits.
Edit: I just did an HN search for the few books I couldn't find with this app and was able to find comments recommending all of them. Not sure if this means I need to branch out or just that HN reads a lot..
Sentiment analysis is hard. In fact I've never seen it work yet.
Regardless, fantastic work already!
Recommendations would include comments like, "This novel is really the one that ties the previous 37 books together." or "You might want to skip the next dozen books if you're squeamish about things that ooze."
Then there is there is the unashamed embrace of over-the-top in so many different ways.
A while ago I created something adjacent to this that looks for hacker news review of books on goodreads (https://github.com/spookyuser/hacker-reads)
So I'm very curious how you managed to find book titles, I ran into a lot of issues trying to figure out, for example, with "Clean Code" whether to search for "Clean Code" or "Clean Code: A Handbook of Agile Software Craftsmanship" since people mentioning the book used both instances. And of course someone mentioning just "Clean Code" might be referring to the concept not the book. I ended up settling on `${titleMinusColon} - ${author}` but I'd love to know what your approach was given that you used deep learning to search.
EDIT: Just read your comment below on your approach, very interesting!
I'd assume that Arxiv links are often there. So it's a problem that can be addressed with an easier solution (just looking for Arxiv links).
Edit: another little nit: it looks like quite a few books list audiobook narrators as coauthors
Animal, Vegetable, Junk: A History of Food, from Sustainable to Suicidal by Mark Bittman
A couple of thoughts:
* It would be great if each book were to have its own URL (for sharing).
* Consider allowing the search to allow author input, e.g. if I want to find the book 'Who' by Geoff Smart, the single-word title isn't specific enough to show that book at the top of the search results.
In the case of the book I searched 'Who', showing it in 4th position seemed about right.
What data source are you using for the books, authors and covers? I looked at OpenLibrary [1] but the covers are not the same, so I suppose it is something else? Maybe Amazon directly somehow?
[1] https://openlibrary.org/search?q=zero+to+one&mode=everything
Does it take into account negative reviews/comments? I have seen that Why we sleep is being recommended in the 6 months tab, but, while it was received with a lot of praise, it was soon critizised by others researchers in the field and I would expect that the HN crowd would have followed that trend.
To summarise: organisms evolving to change the environment around them to their benefit. I went to Foyle’s one day with butlying a book on Termite mounds in mind, that is one chapter in the book.
I found out too late that UCL had hosted a talk by Dr Turner a year too late.
("Cryptonomican" was a good story, but I really hated all the jumping around every five pages. "Diaspora" has sort of so-so writing, but very hard-math sci-fi and quite interesting ideas. I think it's the "hardest" sci-fi book I've read, which includes "The Martian" and the red-green-blue mars books.)
I just bought Cryptonomican. Looking forward to reading it.
The book I am reading right now is also chosen based on HN recommendations (The Talent Code). And I am about to read the GTD book by Allen. (I hate self-help generally with very particular exceptions)
After packing/archiving my library of physical books around 2010-11, I went all digital. However, I came back and try to stay roughly at 1:4 (physical:digital) book ratio when my daughter complained that me reading on the Kindle, "Are you really reading a book."
I re-started buying physical books in 2018. I have a knack of buying books recommended by Hackernews comments and the curated list of people I follow on Twitter. I re-started with less than 10 physical books around mid-2018. Between me and my daughter, we might have crossed 200+ physical books. I need to figure out a better way to deal with this.
95% of every book I have ever read or owned is in the first 20 pages.
Its almost just as fun to read the comment chain about each book.
You must be independently wealthy because I know no one cares if their is an affiliate link. I believe affiliates are always paid to the last cookie you have.
Interesting that that's one of the recommendations.
But, what's wrong with using Amazon affiiliate links? If anything, monetizing would be great since it would give you more incentive to maintain this wonderful application? And it doesn't cost us users anything.
There was a fantastic HN comment[0] which actually spurred me to buy and read it. Do your queries go back far enough to pick this one up? It's an interesting example where one sentence mentioning the author alludes to an association with the title of the book in a sentence that is explicitly about the OS.
Do you think comments like these also could be picked up by your algorithm?
I was curious how many times some common textbooks were mentioned but didn't find them via the search, which could be user error. But to give a specific example. None of the books in this comment thread were found:
https://news.ycombinator.com/item?id=19893447
Comment text like this: "CMOS VLSI Design: A Circuits and Systems Perspective (4th Edition)" by Weste and Harris
should've been caught, right?
Is seems simple and easily justifiable reward. I didn’t click the links, but hopefully you used smile.amazon for charity.
This is novel and useful. Thank you.
- told by friend
- someone I admire read book and commented on it
- I'm working on some problem, during that exploration i came across books.
- random people mentioning book on platform like HN on a topic/post of my interest.
Same! Some of my favorite book recommendations have especially come from this one, I don't know why but a one line comment on a HN thread of "what book changed your life" has become my favorite way for discovering books.
I find that I get sold on a lot more when it is just a random single comment on some thread somewhere that focuses on a single aspect of a book.
If you can find a hyper specific subreddit/forum/etc. for a sub-genre you like, then you will spend more time reading books than reviews...
I think a book list is more useful if it has some books in it that some love and some hate, rather than only books that no one minds very much. Maybe some of them will turn out to be ones I love.
(I happen not to be a Rand fan myself.)
For example, The Design of Everyday Things has some interesting content, but I found it almost unreadable. It's such a poorly organized and written book that its design seems to go against almost every rule discussed in the actual content. Keeps surprising me that this is such a highly praised book.
Do you think it would work on podcast transcripts? The formatting of the titles wouldn't be nearly as regularized so it might have quite a few false positives/negatives, but there's a lot of book conversations out there.
You could have a form to notify you of a post which seems to be not processed, e.g. "https://news.ycombinator.com/item?id=28549134", or "id=28591398" etc.
(BTW: very great work, and thank you for your invaluable service)
I just bought 5 Children's Books and I don't mind if you can benefit from the affiliate link that goes from your account.
The longer extracts are more useful than the shorter extracts.
For Brave New World, I noticed the first 100 - 200 comments are short, and not useful as reviews so much as indicators of preference. Then after that, the comments are longer, and hence, more useful because they explain something.
It would be useful to be able to filter word length so as to be able to distinguish between Opinion Mode vs Review Mode.
What are your future plans?
Of course, taste is subjective, and it should perhaps be expected that much of the list is in line with what is read by the general public, but many of the books are either presenting fact or attempting to convince the reader of the veracity of a certain viewpoint. I'd like to read more open-ended works that ask for interpretation on the part of the reader or, at the least, don't explicitly spell out what they want the reader to walk away with. (certainly some books here fit the bill, e.g. Infinite Jest, Pride & Prejudice, etc.). Again, interests are subjective.
In light of this, book recommendations?
Not good or bad, just a function of where they're coming from.
And as for book recommendations, Children of Time by Adrian Tchaikovsky.
Fantastic work, and a terrific interface to boot.
Thus the problem with existing solutions is NOT "limited recall" or "insufficient rules" or "no Amazon link".
And the problem with this "solution" is that there is no justification for why a book is great and applicable to my circumstances, and people have to trust your black box. Otherwise I'm likely to waste my time, just like reading books from any other crappy recommendation engine.
With a deep learning model reducing all the reviews to "book names" you've successfully removed the value of the book discussions themselves. Therefore, for me this engine and all similar engines are strictly worse than simply going through the actual big threads themselves, i.e. https://news.ycombinator.com/item?id=21900498
Edit: I've just seen the embedded comments by switching to a desktop browser. It's a nice addition. However, for me to make sure I'm not wasting my time going through arbitrary books and comments, I would need to know why a book is ranked highly compared to other books. And I want to be sure that ranking is tailored to me, at a very, very high accuracy.
It literally shows each comment in full that it extracted the book name from. It also includes a link to the comment in the original thread. What more could you possibly want?
Also occasionally break the rule for a book you want to read. It isn’t like that would kill you.
Novels, philosophies, and histories are works that can stand the test of time if they’re good enough
The comments panel show the actual recommendations. And the books are ranked by number of recommendations. Is this not enough?
One thing that I've noticed is that when I select another book, the scrollbar in the comment section doesn't automatically scroll up.
Is it because it is a number?
https://news.ycombinator.com/item?id=20285306
On the topic of Brave New World, the site categorizes it as a Reference.
The app may need to special case it.
That said it's an interesting project that clearly took some effort to put together.
Edit: Saw that the author commented on 'books with similar names. Many of the comments I saw had the books authors name in them as well, next iteration should perhaps look for and match on those as well.
https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu...
Unrelated to the above, but it would also be nice if the site could search by author (I don't seem to get any hits when putting in author names) or even year of publication.
How often are you updating your results ? Can I download the recommendations dataset for offline queries ?
So many interesting books I will never have time to read. Could somebody please find a solution for the problem of mortality? If there is time.
Is it ok to call these type of applications crowdsourcing?
Also, if it makes sense, have a monthly list.
Wonder how “Focus” went under “Christian Books & Bibles”, though.
In some cases there are books that some user recommended and then the book listed is not that one. As an example, lots of people are recommending Spivak's "Calculus", or "Calculus Made Easy", from some other author, but the one listed is "Calculus" from James Stewart. Same happens for the book "Calculus: Early Transcendentals".
Sun Tzu's Art of War is also repeated with two different editions.
I also recommend “Islamic imperialism” from Yale, “the bomb in my garden” by mahdi obeidi and “nothing to envy.”
(Plus the fact that it's a good book on its own terms. At least, it is so far as I can tell; I am not an architect and maybe some of the advice in it is actually terrible. But it seems almost always reasonable and frequently insightful, and it's well written, and the "pattern language" idea that software engineering borrowed from it is a nice one. (Though the software-engineering borrowings don't generally amount to actual pattern languages as opposed to miscellaneous grab-bags of alleged patterns.)