Show HN: Hndex.org – a full-text search engine of articles submitted to HN
hndex.org
hndex.org
The post in question: https://www.johndcook.com/blog/2020/07/25/worst-tool-for-the...
I think HackerNews and other link aggregators (e.g. reddit) have a kind of recency problem, where there is a lot of great content, but people only see the recent stuff. This seems like a great way to uncover some of the latent value of old HackerNews content.
Is there a way to suggest content too? e.g. If you liked that, you'd probably also like X, Y, and Z.
With Reddit and HackerNews, I want a relative ranking index. I can search by the top content, but something 5th today could have more votes than the top submission of 2015 because of forum/subreddit growth.
I want them ranked by something like the ratio of views to upvotes or upvotes compared to total upvotes for that day.
Yeah, this is definitely possible with public data. I did something similar on reveddit [1] for removed reddit content. Hovering over the graph shows the item with the highest vote ratio [3], and clicking skips to that point in time. Code is here [2] for anyone interested and I apologize in advance..
[1] https://i.imgur.com/p3Bi5IS.png
[2] https://github.com/reveddit/ragger/blob/master/revddit_aggre...
I believe HN's ranking system is extremely creative and works great for the main page and day-to-day use (my understanding is that it additionally makes use of time-decay terms for comments and stories). Like you said, it's really just historical search (i.e. algolia) that seems broken.
https://www.evanmiller.org/how-not-to-sort-by-average-rating...
https://news.ycombinator.com/item?id=15709405
https://www.evanmiller.org/bayesian-average-ratings.html
https://www.evanmiller.org/ranking-items-with-star-ratings.h...
https://redditblog.com/2009/10/15/reddits-new-comment-sortin...
Work in progress:
site:news.ycombinator.com some term
...but a many times "some term" is not in the title of the HN post, but in the body of the article. I didn't see any other tools that did this for HN.Most articles have no comments. Yet one can search on hndex for terms only found in the article, not in the title.
The correction I was making was about HN Algolia search (which is linked from the search box at the bottom of each HN page), which only indexes content on news.ycombinator.com itself – i.e., article titles, comments and text-only posts like Show/Ask HN – but not the external content in submitted articles.
For some reason I couldn’t find my own blog post[0], even when searching for the embarrassing typo I made in the title - acommodating.
- Add a search Button for convenient mobile use or if a user copy pasts things into to the search field using the mouse
- Add a comment counter on the result page, Since you index every article a lot of them have none or very few comments.
Oh and just a warning, depending on the jurisdiction providing the cache could be problematic under copyright laws since its basically a copy of the article.
site:news.ycombinator.com intitle:"some term"Go back and click "cached".
That was a cool experience.
Bug report: in Safari on dark mode, the text in the search box is almost the same color as the background (white on white).
From "GitHub was also talking to Google about a deal, but went with Microsoft instead" https://www.cnbc.com/2018/06/05/github-interest-from-google-... - Salesforce got MuleSoft for $7.5 billion, Microsoft got GitHub for $7.5 billion.
My only wish is that it should be possible to sort chronologically.
How does ranking work? Apologies if I’ve missed the explanation.
> How does ranking work?
Great question, it's something I'd love to know more about as well.
This is also why HN still looks like a site from the 90s, instead of New Reddit (thank God).
The OP site also has a “cached” link for each article, don’t know if you saw that.
Also, to the maker of HNdex, consider adding links to Archive.org and Archive.is next to the cached link you have, so that readers can click through and check if they have a version of it in case there’s images etc
HN feature request: Add submissions to archive.org (or equiv), include that cache link with story.
(Algolia was backed by YC btw)
Seems nice. Sort and filter functionality would probably add to this, but I will bookmark this for sure and try it out as a search engine for tech topics in general.
Thanks for creating this!
This seems to be a good complement to algolia, given that it searches through the linked pages instead of comments.
Minor nitpick: would it be possible to make it give a 'past' link, to search for all discussions on a result? Some of the 'comments' take you to duplicate posts with no comments instead of the more popular cases.
http://kakapo.susa.net:8080/cfs/
I abandoned it when Google deprecated the WebRequest API, but the code's still available on GitLab. https://gitlab.com/ksangeelee/cfs_build
It allows article score and uBlock Origin 'hits' as search criteria.
- What is your stack?
- How often is your database updated?
- And will it be open source in the future?
I often find myself searching HN for opinions on different technologies over time and this will be invaluable for that use case.
Is it possible to give a little inside peek of the techniques, algorithms and tech stack etc. you have used?
Great idea, I appreciate how fast it returns results! Just needs more more control over the search parameters and figuring out why articles like the example I posted above aren't working and you got yourself a nice HN search.
[1] https://chrome.google.com/webstore/detail/falcon/mmifbbohghe...
https://hndex.org/?q=Terence+Eden
Why? Because those old posts have a "what's new" set of links. One of which contains my name. I'd suggest only searching the `<main>` element, perhaps?
1. Show dates in results.
2. Are results better than searching on Google with "site:news.ycombinator.com"
Fantastic project and well worth creating!
As another commentor has noted - this totally disobeys recency bias and throws up interesting articles for a topic.
Edit - it is disappointing how many of these links 404. But even if that's the case the headline and intro is a set of time capsules of sorts nonetheless.
Ideally for this to be useful I'd want date on HN, date of article (although I see that'd perhaps be hard to extract from unstructured pages) and also the HN points, as those are a major proxy to quality usually.
As others have said it has a nice crisp UI and i like that it's so quick.
Ps: show the date / age of the posts on the results page
It's such an intense "dark mode" that I couldn't read the first full article of interest I found because the contrast was killing me, and then coming back to regular HN left me nearly blinded for a few secs
There are much fewer books in the last six or so weeks editions. Don't know why this is. Whether it is some scraper not working, or actually no books being mentioned.