Extracting Hacker News book recommendations with the ChatGPT API
blog.reyem.dev
blog.reyem.dev
The website looks like a clever way to generate a lot of clicks to Amazon using his affiliate link.
If you are into books, I highly recommend searching HN with the simple keyword "books", and filter using "Ask HN" tags [1], or simply by "books". This is how I choose almost all of my English books now (I am trilingual and can read more languages)- even non-technical books. I have been doing this for more than two years, and I really like HN for books recommendations.
Many years' worth of high quality reads can be found on HN threads related to books. They are goldmines.
EDIT: There is also Hacker News Books [2]. Check out their Top Books of All Time section [3].
[0a]: https://hacker-recommended-books.vercel.app/category/0/all-t...
[0b]: https://news.ycombinator.com/item?id=28595967
[1]: https://hn.algolia.com/?q=Ask+HN+books
I understand and know the pain of maintaining a fully online model, and that is asking too much from an individual. But, say, a manual update every six months or even a year would be okay.
So I went to your link at [0a], selected last 6 months. There are 58 pages of recommendations and 15 recommendations per page. So I rolled a 15 sided dice to choose the page. Then another roll to choose which book on that page.
I will be reading The Very Hungry Caterpillar
15 sided die? Why didn't you also roll a 58 sided die then?
> I will be reading The Very Hungry Caterpillar
Great luck! I loved that book as a child. It may have been what led me to HN...
I thought “Hacker Recommended Books” sounded immensely useful, so I decided to save the bookmark to my pinboard account.
“Previously saved 2 years ago.”
Oops.
(Not a critique of the site! More of myself and forgetting services like these exist despite my best intentions.)
Have to admit I struggled with the latter half, mainly because I didn't quite memorize the first half even though "I get it".
I do like the "teach it to me like I am 5" approach. People try to gloss over the basics nowadays like it's not complex enough intrinsically.
It's really great.
It gives you the false sense of you getting the full picture. But not really.
Among other HN darlings in books- The Little Schemer, SICP, The Black Swan were totally worth it.
It’s Meditations (by Marcus Aurelius).
[1] https://en.wikipedia.org/wiki/Meditations_on_First_Philosoph...
I own several of his books (admittedly they were gifts), and have never read them. So them not showing up in the top five doesn't surprise me much.
https://corecursive.com/066-sqlite-with-richard-hipp/
Richard: You just pick things up. People tell you these things, and that happened to some with Bloomberg. They’d come to us and say, “Hey, why aren’t you doing this optimization,” and I said “Never occurred to me.” “Well, can you do it?” “Let’s see what we can do,” and then it would go in, so, yeah, kind of figure it out as you went along. I had to invent a lot of this myself. Nobody ever taught me about a B tree. I had heard of it. When I went to write my own B tree, on the bookshelf behind me, I’ve got Don Knuth’s The Art of Computer Programming, so I just pulled that down, I flipped to the chapter on searching and looked up B trees and he described the algorithm. That’s what I did.
Funny thing, Don gives us details on the algorithm for searching a B tree and for inserting into a B tree. He does not provide an algorithm for deleting from the B tree. That’s an exercise at the end of the chapter, so before I wrote my own B tree I had to solve the exercise at the end. Thanks, Don. I really appreciate it.
Adam: That’s awesome. Did you pull anything else from that book?
Richard: Well, it’s an amazing volume. I can’t give you a specific example, but from my era, everybody has to have read or at least skimmed through, at least browsed through The Art of Computer Programming, and know that algorithms that are there, maybe not Don’s exact implementation. I mean, I never took the time to learn MIX, which is his assembly language, but it’s useful to flip through and look at all the algorithms he talks about. I think that just a year or two ago I needed a pseudorandom number generator, and I was, “Let’s see what Don recommends.” You pull it off. You see what he does.
The post appended the raw data provided by GPT, allowing you to verify the integrity of the data. This makes the post trustworthy from a methodological pov.
> I don’t trust anything published on the Internet after Llms went mainstream.
You always had to verify the integrity of the data and methods used in any publication, regardless of the medium. The responsibility of both authors and readers hasn't changed. If you took things for granted before LLMs, you shouldn't have, and if you don't trust trustworthy authors post LLMs, you should.
But now with LLMs everywhere, people will realize it is necessary to verify.
A post is not trustworthy if it’s reposting trash, even if it shows the source.
> you took things for granted before LLMs, you shouldn't have, and if you don't trust trustworthy authors post LLMs, you should.
The nature of how LLMs hallucinate is different from how garbage used to appear on the internet. Before LLMs there was a relatively good inverse correlation between quality and blatant bullshit. Not enough to pass the verification rigor required for an academic publication by any means, but enough that you didn’t have to second guess every single statement on every web page listing something as simple as book authors.
When it comes to what LLM's write, I find that LLM hallucinations are like self-driving car crashes. We are hyper-aware of the machine doing something that we ourselves do every single day and consider a normal defect of biological conscious.
I can't believe that. LLMs always talk in the same confident tone, entirely regardless of what they're saying is true or not. What is true in the real world literally doesn't come into the equation.
Whereas at least some of the time, humans will say that they're not sure and might be wrong, or otherwise sound less confident. And that's related to how true the thing they're saying is.
Is that so? https://innocenceproject.org/dna-exonerations-in-the-united-... These people were convicted by people who were 100% convinced their memory was correct. The DNA evidence, which is "harder" evidence, said otherwise, and in these cases, was exonerating. (There are hundreds, possibly thousands, of other cases like this by the way, where the imprisoned innocent is NOT yet exonerated, all based on overconfident eyewitness testimony that yet managed to convince a judge/jury.)
There is also the well-known Dunning-Kruger effect, the cognitive bias where individuals with limited knowledge or expertise in a particular area tend to overestimate their competence and confidently assert their opinions. We've literally seen this countless times just since the 2016 US election, just watch literally any Jordan Klepper interview https://www.youtube.com/watch?v=LoZ2Lt_aCo8 (honestly, this is a little too political for me to use as an example, but I ran out of time seeking out unbiased examples... Mandela Effect? Misplaced keys being common?)
I'm afraid you're a little off, here, on your faith in humans not hallucinating memories and knowledge.
How many times have innocent people been wrongly convicted? The innocence project found 375 instances in a 31 year period.
How often do LLMs give false info? Hope it never gets used to write software for avionics, criminology, agriculture, or any other setting that could impact huge amounts of people…
Luckily I only said humans add some doubt to what they say some of the time :-)
I think this is a valid area of improvement, and I think we'll get there.
Also experts tend to be much more accurate at evaluating how knowledgable they are (this is also part of the D-K effect). So Id much prefer to have a 130 IQ expert answer my question than an LLM
They do, but they are better at verifying what they've already said. So a simple prompt asking them to verify the facts they've presented often improves the accuracy. There are also other techniques like chain of thought and tree of thought that further improves accuracy.
> at least some of the time
YMMV
oh here we go. You're one of those people conveniently restricting this accusation to a machine that scores a 130 IQ (https://www.reddit.com/r/singularity/comments/11t5bhh/i_just...), instead of also including humans, who notably will send someone to prison 1000% sure that they witnessed that person doing the thing, when in fact, later DNA evidence exonerates them (https://innocenceproject.org/dna-exonerations-in-the-united-...). Fucking LOL. Get out of here, doomer, the rest of us have AI-enhanced work to do.
Those humans who recalled incorrectly could have a 130 IQ, proving my point above and making your ad hominem reddit speak luddite insult fall flat.
It is invalid if you take what an LLM says as simply what another human (who happens to have a broad knowledge reach) would say.
Of course humans also hallucinate, but we didn’t have to take that into account every single time we read a piece of information on the internet. Humans have well-documented cognitive biases. Also, usually, a human’s attempt to deceive has some motivation. With LLMs, the most basic of information they provide could be totally false.
That said, overall utility of anything plummets drastically as reliability goes below 100%. If a particular texting app or service only successfully sent 90% of your messages, or a person you depended on only answered 90% of your calls, you'd probably stop relying on those things.
(I wish I could edit out my vitriol.)
Blaming LLMs for everything is becoming the preferred excuse for people who like to reject what they read and substitute their own beliefs instead.
It’s true that LLMs hallucinate and are definitely not correct all the time, but the way people are using that as an opening to reject everything on the internet and elevate their own prior beliefs to the top is strange.
Maybe it will be right, but the best part? We'll continue to be reminded of this fact.
Thanks for reminding us all.
Sometimes, it is that simple: "Rand bad".
But, yes, Rand’s turds do appeal to those poor enough financially that they didn’t receive a rich and broad education and poor enough intellectually that they couldn’t give themselves one.
To me, it is the story of someone trying to create cool stuff and the world making that hard (felt like a celebration of human creativity). She is brutalistic in her messaging but it is an interesting story. I weirdly like her writing style, its like someone pounding a hammer against my skull. Haven't read her books in 10+ years but in my teens and twenties, I found them thought provoking and inspiring.
But I believe it might be toxic for the HN crew.
Let's be frank, people don't read much on average, so they take whatever book or author you read as your own personality, when in reality people should read many books with different ideas, so you can form your own.
One can learn a lot from reading from Friedman to Marx.
This is why I also think that one should read all the classics instead of chasing book recommendations, any community you ask will give them recommendation for non-technical subjects that are from their own worldview, which in turn, will make you not smarter, but likely more stupid than you initially started before you've read those books.
She had a really unique view of the world, I appreciated it, and it definitely helped me develop my own view of the world. I am about to read the Upanishads along a similar line of thought.
It could do with being recommended less, imo. There are probably good summaries of it around.
Popularity is more a question of marketability, especially in the dire non-fiction category. Reading is a secondary use for most of these books, and it makes it even harder to discern which lifestyle books are resume-padding filler or truly generalizable advice.
Newton’s Principia is totally and absolutely impossible to understand to the average curious person raised on Cauchy and Weierstrass. There was a course at university aimed at people with advanced math knowledge to prepare you to read it.
Possibly Mein Kampf could compete in short passages by a "reading difficulty" score, but it's really tedious and much longer (by a factor of 5!). Darwin could actually write. Origin of Species was intended for the educated public as well as for scientists.
Hold on, hold on.
I started to read Origin of Species.
One time. But I did read it. Not all. Unconvincing.
And I did read all of Gabriel Marcel's Metaphysical Journal. So ... not a matter of Slack. If one in 100 reads Origin that's a rough approximation to the number of masochists
If you dig into any "recommend me books" post, you'll find some truly great recommendations near the bottom.
Like if you already have a group of friends and no worries about keeping that, you probably already know everything in the book. As someone who grew up as the perpetual outcast, the advice in that book was really useful for me. It was essentially cliff-notes for social skills I should have learned at the age of 10 but didn't.
The only "red flag" I see when someone mentions the book is that they were probably very socially inept at some point in the past like I was. They still might be, but they're working on it.
Ultimately the core of the book is "think about the other person and their perspective" and gives advice on how to apply that in cases of making friends, keeping friends and having good business relationships. Or to frame it a more selfish way, "the best way to get what you want from other people is to give them what they want first."
Like I said, it's nothing groundbreaking, but the stories and examples were helpful to me when I was socially clueless. The average person might still find value in it because even though these things are easy to know, it takes conscious effort to apply. But I also think writing "Pause and think about the other person" on your hand with a marker will have exactly the same benefit unless you really need some of these things spelled out for you.
It was just a spot check -- not exhaustive. Also many who refer to Descartes in the raw data, do so regarding other works.
I think this is an interesting mistake.
That doesn't look trivial at all to me... but... if you did that, please share with us a Github repo.
Like I mentioned it only looks for books that are linked (or videos, arxiv papers, Wikipedia etc.). I then use the link to get information from the site itself.
I calculate scores for the given link weighted by a number of metrics.
I'm sorry if I gave the impression this was fancier than it actually is.
Learning how to use a new tool with older type of work can be useful and enlightening.
https://replicationindex.com/2020/12/30/a-meta-scientific-pe...
Edit: here's the article with the comment from Kahneman:
https://replicationindex.com/2017/02/02/reconstruction-of-a-...
Your reply states that for you even the word "libertarian" is offensive.
[1] http://rationallyspeaking.blogspot.com/2010/10/about-objecti...
https://old.reddit.com/r/AskReddit/comments/16yaui6/what_boo...
For example: take a look at some of the comments here https://hacker-recommended-books.vercel.app/category/0/all-t... while there are some recommendations (I expect this) there are also quite a lot of people using the book as a way to deride purveyors of its ideology.
I think what we're both experiencing here is an expectation that this list will be GOOD recommendations, but the methodology is popular mentions, which happen to contain some people's (subjective) GOOD recommendations but also peoples (subjective) BAD popular mentions.
That's a good idea. It wouldn't be too hard to get chatgpt to also return the sentiment of the book mentions. Negative sentiment comments could be ignored for the purposes of a top 50 list.
In a similar way Fooled By Randomness, The Black Swan, and Antifragile are all a part of the Incerto series and could have been condensed to one line item.
In any case, great work! I'd like to see a post like this with the top 50-100 authors and the comments full of discussions about which of the books are the author's best / most accessible.
It would be nice if there was some further sentiment analysis.
[1] (PDF) - https://monoskop.org/images/d/dc/Barbrook_Richard_Cameron_An...
So like the number of times a book is mentioned changes?
Wouldn’t you fail a cs student whose program did that?
It's easily not only the better work but Dawkin's best and one of the best books I've personally read.
I always cringe when I see The Selfish Gene, The Blind Watchmaker, etc listed because you're not getting his absolute best stuff.
I say four! four last books.
For instance, I began reading Susan Nieman's "Evil in Modern Thought" but found the writing style a bit tedious after just two chapters. I turned to ChatGPT for a chapter-by-chapter summary, which it provided brilliantly. I shared this with a friend who earned her PhD in the 2000s, and she was astounded. It's hard to overstate the time saved.
[EDIT]: Here's the summary from ChatGPT: https://chat.openai.com/share/2f3d0aca-46d3-490b-8426-5e8813...
Interested to see what that looks like. Also would be great to see if users were self promoting content...
Would love to see aggregations of stuff like this from hacker news.
Would it be better to read in the order #50 --#> 1 or #1 --> #50?
1. "The Pragmatic Programmer" by Andrew Hunt and David Thomas 2. "Sapiens: A Brief History of Humankind" by Yuval Noah Harari 3. "Clean Code: A Handbook of Agile Software Craftsmanship" by Robert C. Martin 4. "The Design of Everyday Things" by Donald A. Norman 5. "How to Win Friends & Influence People" by Dale Carnegie 6. "The Innovator's Dilemma" by Clayton M. Christensen 7. "Gödel, Escher, Bach: An Eternal Golden Braid" by Douglas R. Hofstadter 8. "Thinking, Fast and Slow" by Daniel Kahneman 9. "Deep Work: Rules for Focused Success in a Distracted World" by Cal Newport 10. "Surely You're Joking, Mr. Feynman!" by Richard P. Feynman
> A week later, just after 1:00 a.m. on the morning of November 8, the FBI stated that Ivins was observed throwing away "a copy of a book entitled Gödel, Escher, Bach: An Eternal Golden Braid, published by Douglas Hofstadter in 1979" and "a 1992 issue of American Scientist Journal which contained an article entitled 'The Linguistics of DNA,' and discussed, among other things, codons and hidden messages".
Background: I run hackernewsbooks.com, and the only thing affiliate revenue covers is the hosting and mailing list costs (~$50 a month). It is a fun project I bought from the original owner to learn (and because I loved the idea). I want to add more visuals and unique navigation to it. I also run Shepherd.com, which is a much larger book website and we are not covering our costs yet, it is a struggle.
There is a tendency on H.N. to think everything is a scam; I'd love to see that moderated a bit so that people realize it's people like them building cool stuff that excites them.
I think most people who recommend a book do it to help the world and spread knowledge and would be glad to have their recommendation reshared.
All Google does is hijack other people's hard-written content and slap ads around it.
I am half-kidding, as I love Google, but I think your comment lacks some self-awareness about the reality of curation and what it offers to people. And how the world is funded. I mean how much of your house is paid for from ad revenue while you worked at Google?
Google didn’t make money on the curation, they made money from the separate ads on the side. I would have no problem if the author jammed ads off on the side.
The fact that you think ad revenue is the same as adding and replacing referral links from things that people already curated is pretty disturbing.
The creator of this project threw in some referral links, but he didn't replace them. And maybe he made $20 to $50 bucks for what was a huge chunk of his time to create this fun data analysis project.
Seems a weird thing to complain about.
That seems sufficiently transformative to make his referral links okay.
web 3.0 looking good
But as it is, I agree.