Hacker News Official API
github.com
github.com
Instead, try considering the (more unknown) Algolia API[0]. It supports fetching the entire tree in one request. Unfortunately the comments aren’t sorted, so you’ll still have to use the official one (or parse the HN HTML) if you need that.
For my HN client[1], I parse the HN HTML (as it’s required for voting), and sort the comments from the Algolia API. It’s much faster than recursively requesting the official API.
Van Puffelen answers some of this here, but he doesn't get into specific queries: https://www.youtube.com/watch?v=66lDSYtyils
He also mentions here that Firestore offers a bit more for WHERE type clauses over Firebase:
https://stackoverflow.com/questions/26700924/query-based-on-...
I want to build a HN web client but it's not possible without scrapping the site which won't scale.
Another issue is browser only client. If HN has set proper cors (I didn't check if it is the case), then you cannot build a web client through scrapping without proxy.
Do you have some updated version that I'm unaware of?
I'm reluctant to lose it as I'm about to release a browser extension that relies on it pretty heavily :)
[Edit]: this feature is undocumented in the link above, but it's a standard feature of FB and works as expected in all client libraries https://firebase.google.com/docs/database
To demonstrate, here's an example in 18 lines of code that v. efficiently listens to live updates to a profile (yours): https://codepen.io/theprojectsomething/pen/xxWBOWN?editors=1...
To that point there's definitely a few OSS solutions out there that provide this kind of functionality out of the box (recent YC alumni among them). But then who knows what problem you're trying to solve. Sounds like this is just rain on the tip of the iceberg.
Just being able to get an entire thread in a single request alone, and filter queries would make things much easier.
Related:
We're already helping to create a training set of slightly above average English speaker discourse for free.
This would include mod commutations and mod logs; unless it was security related, which would have a formal process with deadlines for release.
—-
To address you specific question as why the comment voting data would be of use, among other reasons, it would allow users to have custom filters. For example, posts and comments with high up/down votes tend to be flamewars. Mods also random boost/ding posts/comments too.
Lobste.rs actually exposes even less and doesn't provide an API because the users haven't consented to such intrusions.
Personally, I don't want to be able to review every vote for or against my HN account. It wouldn't be productive and would likely lead to me not wanting to be here at all. It'd be like having a super strong mirror and then looking at your skin and being confronted with too much information about the reality of how nasty looking and imperfect being human is. Basically not a healthy pastime.
FWIW, if you're truly passionate about this, it's not difficult to setup your own site with the lobste.rs open-source code. Maybe there is a whole subset of like-minded folks.
The lobste.rs origin story may also be of interest to you. The TL;DR is jcs was being singled out and abused by pg on HN. Thankfully there is no evidence dang has (or would ever) do this, it was certainly toxic behavior on Paul's part. Being a moderator isn't for everyone, it's a Certain Thing, and net-positive mods who avoid power tripping are truly rare gems on the Internet. Open information would guarantee such abuses couldn't happen in the dark, but at what cost to the best aspects enabled by the HN ecosystem?
References:
And yes, aware of Lobste.rs mod logs, but Lobste.rs is basically dead compared to HN — largely because it’s an HN clone.
My outsider opinion is that lobsters is dead because the sign up process is designed to keep people out. I appreciate there's also the network effect of people coming here because people come here, but I for sure would visit lobsters too if I could participate without having to know a friend to get past the velvet rope
In 2019 I set up a bot to store titles for HN stories as soon as they appear on the "new" page, so that I could then compare them to the current version.
I had also made a browser extension for FF and Chrome that compared the current version of any HN story to the stored one, and displayed a warning if they differed.
There was zero downloads of the extensions, so they fell into disrepair and don't work anymore, but the cron job for storing the titles is still active.
If there is any interest I could resuscitate the extensions, or some other system to make that information useful.
To me, if users have concerns about “unfair voting” to me the solution is not to hide them. Also, if comment points aren’t important, why even allow them? They are obviously important and if a users is not available to respect guidelines not to complain about them, what’s difference between that an guidelines related to comments about being downvoted?
If votes being public impacts voting patterns, listing top users, having usernames publicly associated with comments, being able to follow given user’s comments, etc — for sure does too. What’s the difference?
A moment later
Maybe it'll help if you get your head around the goal of HN. There are a bunch of things you personally seem to want to accomplish with HN. But HN itself has just one goal: to nurture curious conversation. Other sites have other, equally valid goals: propagating the news as quickly as possible, or relentless focus on a particular topic, or getting the best Q&A pairs to the top of Google's search results. HN, though: curious conversation. That's the whole ballgame.
Anything that drags on curious conversation is a non-starter on HN, even if it might make a bunch of things better for you.
We don't have to guess about whether making comment scores available is a drag on curious conversations; it manifestly was, for years. People have an innate, visceral reaction to comment scores they feel are unjust, and they talk about them, and those conversations choke the landscape like kudzu.
A lot of things about HN make more sense when you accept the premise of the site, and understand that HN will make most sacrifices it can come up with to optimize for that premise.
Personally I see this site as the Reddit of tech with 20% of its flavor coming from YC companies. That is how it behaves, as opposed what moderators say the site is about.
(https://james.darpinian.com/blog/scraping-my-own-hacker-news... gives one way of doing it, by crawling your /threads?id=‹username› page, which does contain the point counts for each comment by you.)
[1] https://github.com/superb-owl/hacker-news-comments
[2] https://superbowl.substack.com/p/commenting-on-hacker-news
[0]: https://hnhub.dev/
It was indeed barebones but I'd like to think it helped me find a job. I would link the code (it's on GitHub) but today the "Who is hiring?" posts include a few considerably more advanced and capable "searchers" right in the top of the post:
> Searchers: try https://kennytilton.github.io/whoishiring/, https://hnjobs.emilburzo.com, https://news.ycombinator.com/item?id=10313519.
It was fun; I enjoyed working on it.
I am definitely happy to share, but I didn’t expect any interest and I posted my comment under my HN pseudonym account. However when I link to the repo on GitHub that thin veneer of obscurity will disappear and I’d prefer that it did not. Being a programmer, I am quite familiar with workarounds and I am happy to offer you couple of different options:
1. You could add an email address to https://news.ycombinator.com/user?id=maxique and I can email you there - I’m happy to share my name and contact info in a one-off request, just not on an ongoing basis for anyone to see.
2. I could clone the repo onto my current machine and delete the .git dir, leaving the code and the README.md intact, produce an archive of that stuff, and link it here in this thread. While I’m sure it may be possible to use a code search tool to find the repo on GitHub after that, I am not terribly bothered by that.
I am open to other options as well if you have any suggestions. Finally, the code was written for whatever version of the Ruby interpreter (MRI) was current in 2016, in case that will dissuade your interest.
EDIT: I have a bias to action, appreciate interest in my work, and figured it was likely that you would see my reply at some later point and then I'd have to see your reply to my reply at some even later point, and we'd continue the loop, and eventually your interest would wane. So I went with option two. Here you go: https://we.tl/t-Bu257EdQn9
This is my first time using WeTransfer; let me know how it works out for you. I actually originally had a link that I hosted on file.io (which I've also never used before), but it looks like that gets automatically deleted after one download, which seems slightly excessive to me.
This repo started in 2015, and Firebase was bought in 2014. I'm really surprised; I thought it was way more recent than that.
Gives a few more variables, and the different sample times show a different set of transient effects.
https://gist.github.com/Q726kbXuN/15e61acc003bb6d46a458001fe...
Has been quite useful for the ops side of what I work on!
On GitHub: https://github.com/plibither8/hn-faves-api
You can try it out here: https://hn-faves.mihir.ch
https://firebase.google.com/docs/database/web/read-and-write
https://github.com/overshard/newtab/blob/master/newtab.js#L1...
Haven't had any issues with it!
I wrote a script that uses your browser cookies to scrape them from the website here: https://github.com/superb-owl/hacker-news-comments
Is there any way to query this API as if showdead is switched on?
I do wish it would support login though, but you can’t have it all.
E.g. I use Hackers [1][2] because it is much easier to read stories on mobile and supports dark mode. Unfortunately, it doesn't support commenting or downvoting (yet?)
[1] https://apps.apple.com/us/app/hackers-for-hacker-news/id6035... [2] https://github.com/weiran/Hackers
https://apps.apple.com/ca/app/hack-for-hacker-news-developer...
It would be a shame for this vast trove of expert knowledge and nerd sniping to vanish into the ether after the next hard drive apocalypse or similar.
Having a self-hosted private HN Algolia (perhaps via Elasticsearch) would be nice. Algolia's forced fuzziness and limited number of results per page are a drag. Why can't they provide any search controls? Or at least support the "-" operator to exclude a term.
TIL that HN doesn't use a "proper" / persistent DB. I wonder then how HN handles data persistence. Is it as simple as creating write-ahead logs + checkpointing every some duration?
Thanks!
I'm not sure it's linked anywhere visible, but it does have an entry in the html head....
I'm not sure if there are more tho, I know /ask doesn't have one.
See the link at the bottom of every HN item page.
Edit: I took a look, it seems like Firebase is owned by Google now, but it was original a YC-funded startup, judging by this page: https://www.crunchbase.com/funding_round/firebase-pre-seed--...