Show HN: DeepHN – Full-text search of 30M Hacker News posts and linked webpages
deephn.org
deephn.org
[0] https://www.everycrsreport.com/reports/RL33810.html
[1] https://en.wikipedia.org/wiki/Fair_use#Text_and_data_mining
Alternatively, base your site out of a country which gives you the freedom to do things like this. Build it in China and you'll be applauded instead of sued. Copyright suits for petty things like this would laughed off when it's something actually useful.
Founders there have no choice but to sell or get copied by the platform that you're building on. It's baba or tenscent.
I agree, some anti-monopoly regulation is in order in China, and I'm not a fan of the Alibaba/Tencent monopolies on the ecosystem. However, I think there are other ways to go about fostering and encouraging small businesses and protecting creators and their intentions than American-style IP law, which allows you to sit on an invention or work and hinder society from having access to the fruits of science, which I'm vehemently against. IP isn't even real property, IMHO.
If you are afraid of being sued, you certainly should not do something like this.
On the upside, there is probably not much provable damage a publisher could ask to be repaired. So they probably are only liable to pay for the legal expenses of whoever sues them.
That said, meaningful fulltext search will need data. But both Bing and Google crawl pages.
It's not even immediately clear that displaying a copy of a page is materially different from caching (which http do in many, many layers).
Now, if présent the text as your own - that might be a problem.
Ed: see also: archive.org.
We mainly surface historical content, which receives additional traffic from DeepHN, instead of taking value away from the website.
Could the folks with JS disabled get a small semi-static page with the top interesting HN stats, and a message that say something like "Here are some simple stats. For the full interactive experience, please enable javascript"?
I doubt that it's worth it. If somebody can't be bothered to enable JavaScript to use the site, they're not a desirable customer anyway.
"Well this is an exceptionally cute idea, but there is absolutely no way that anyone is going to have any faith in this currency."
Written by co-founder/CEO of Pachyderm and first employee of RethinkDB none the less.
The first cryptocurrency thread on HN—about 9 months earlier—was even more interesting: https://news.ycombinator.com/item?id=253963. It contains another skeptical comment: "The bigger problem is that everyone has a incentive to run their computers day and night cranking out solutions, which burns up lots of natural resources and processor time for a zero-sum result." (https://news.ycombinator.com/item?id=253999), which seems astonishingly prescient now, even though the objection is still hotly (no pun intended) debated.
Cherry-picking the wrongest-seeming skeptical comment from that collection of data points is a case study in survivorship bias. (I don't mean to pick on you personally! This is common of course.) The infamous Dropbox comment from 2008 is similar in that it has gotten repeated out of context in a way that is unfair to the original commenter (https://news.ycombinator.com/item?id=23229275). When we repeat these things, it says more about us than it does about the thing we're repeating.
I like to see wrong “obvious” predictions to humble myself, even if not from me.
[1] https://www.wsj.com/articles/missed-teslas-12-551-rise-dont-...
Also from a creator's perspective it's important to see these cases of once in a generation invention (good or bad) and the comments on it.
> 21 million bitcoins seems like a low limit to set. According to [0] the exchange rate is currently 221 Bitcoins per USD, so the total value of all possible bitcoins is only 95000USD.
Today 95000USD is worth 1.51BTC
SpaceX aims to put man on Mars in 10-20 years
A long way from that mess of a site to Coinbase IPOing tomorrow for a zillion dollars.
What did you use to crawl the pages and how long did it take? Curious about your experience doing it and crawler integration with Seekstorm.
Is is possible to easilly expand the index with embeddings (vectors) and perform semantic search in parallel?
Your pricing indicates that hosting similar index as this demo would cost $500/month. Wondering what kind of infrastructure is supporting the demo? Thanks!
ps. Small quirk: https://deephn.org/?q=how+to+be+productive&filter=%7B%22hash...
First three tags seem not to be relevant to the post itself.
The pricing for an Index like DeepHN would be $99/month: While we are indexing 30 million Hacker news posts, for DeepHN we are combining a single HN story with all its comments and its linked webpage into a single SeekSorm document. So that we index just 4 million SeekStorm documents.
Yes, it would be possible to expand index with embeddings (vectors) and perform semantic search. This would we an auxiliary step between crawling and indexing.
The same article referenced here https://deephn.org/?q=how+to+be+productive&filter=%7B%22hash... contains the phrase 'well-defined'
Any idea why doesn't the article surface when searching for this ?
"Well-defined": Just a guess: We are doing key text extraction, i.e. we try not to index boilerplate stuff an and menu items. As "well-defined" is within a short list item, it might be accidentally skipped.
So, that is not yet perfect, and considering the diversity in web page structure it probably never will. But we will try to improve.
Stupid question, but you were crawling news.ycombinator.com, right?
Its robots.txt (https://news.ycombinator.com/robots.txt) contains
Crawl-delay: 30
Why did you not follow that?(I'm not being accusatory. I'm both curious about web crawling in general and have personally been archiving the front page and "new" every 60 seconds or so... (Obviously there's no reason for me to retrieve them more often, but my curiosity persists.))
No, for retrieving the Hacker News Posts we were using the public Hacker News API, which returns the posts in JSON format: https://github.com/HackerNews/API
The crawling speed of 100...1000 pages per second refers to crawling the external pages linked from Hacker news posts. As they are from different domains we can achieve a high crawling speed while being a polite crawler with a low crawling rate per domain.
Yes, it's an annoyance. Yes, sites should fix it. But not ever using a site because of it seems silly when it's literally a half-second longer click.
(Also, for this site I don't actually see the bug. So either they fixed it very rapidly, or GP was just referring to individual searchers being in the history, which is common for any search engine.)
This is not a good excuse for laziness.
> But not ever using a site because of it seems silly when it's literally a half-second longer click.
I highly disagree. With this site at least it is actually possible to leave. Many websites I've come across with this issue, it is entirely impossible to leave without physically holding down the back button in the browser to get a list of history items, and then clicking a site from earlier.
> (Also, for this site I don't actually see the bug. So either they fixed it very rapidly, or GP was just referring to individual searchers being in the history, which is common for any search engine.)
Looks like they've fixed it.
Full disclosure: I work on a similar fast, typo tolerant, fuzzy search engine search engine called Typesense (https://github.com/typesense/typesense).
So, you may be hitting SimpleFSDirectory instead, which does have issues with too many searches.
Could you share the reasons, MMapDirectory did not work for you?
The Lucene's own benchmarks are at: https://home.apache.org/~mikemccand/lucenebench/ , though I admit to not know enough about benchmarking to form a strong opinion.
Either way, good luck with the project/service. More competition is always great. The open-source components look interesting too.
I have one observation. Posts linking to twitter - https://deephn.org/?sort=score&filter=%7B%22score%22%3A%7B%2... - have most of the preview occupied with navigation and other not interested parts. So maybe it is possible to use some hacky solution to show the content part for the most popular websites, or if multiple pages are downloaded from the same domain the content part could be detected automatically.
More observations:
- When trying to go back after clicking the original HN link, there's no response.
- Clicking on what I understand are tags (hashtags) in a given post has no effect.
- Sorting by date is working, though it'd be nice to see the date.
Also a question: how do you generate the tags (hashtags) from the original contents?
The original HN link is opened in a new browser page/tab. The search results should remain unchanged in the previous tab. So you don't use the back button, but the previous browser tab.
>> Clicking on what I understand are tags (hashtags) in a given post has no effect.
Search for "google", go to the first result. there is a hashtag "oracle", click and the search results are filtered by that hashtags. Please post an example where its doesn't work, so that we can fix it.
The date is shown in the preview panel on the right hand side. But wa are thinking adding the time also to the result list.
We are deriving the tags from the terms and bigrams in title, text and parsed html of linked web pages. Top frequent terms per post, but only if terms are within top 65k tags per index. Stopwords are excluded.
Note: I'm trying this in a mobile browser (Firefox).
- Search for a word 'Decentralized'
- pick the second result: 'A decentralized web would give power back to the people online'
- On the result's (preview?) page, click the tag 'Decentralization'
I get no response, nothing updates, no new tab opens. Not sure what is the intended action (I'd assume it should be a new set of results corresp. to the hashtag). Same issue with 'google' example that you described.
But because you are still on the preview page you don't immediately see this as it's done in the background. But once you close the preview page (with the big arrow on the top/left of the preview page) you will see the updated result list and number.
Of course, this is a design flaw, we should automatically close the preview window once you select a tag to immediately show you the updated result list. In the desktop version, this is no issue, as there result list and preview window are always simultaneously visible.
That is related to our instant search, where you don't need to hit the return key to search. This causes the problem, that we need to identify when a query ends and the next query starts, to create the correct entries for the back button. We do this via timing - if the pause between key presses is too long, a new, separate entry is assumed. As the DeepHN is a single dynamic HTML page, we need to implement and manage the browser navigation logic ourselves.
If you scroll long enough within the search results you will finally reach results where the search term is not in the title, but only within the linked web page.
It is more obvious if you search for queries which are less popular than "friendship"
SeekStorm is intended that the user can index and search their own private data.
But if there is demand we could also allow to search in a central public index for web data or other public data.
A bug probably. Huge amount of text in title here.
I use email as kind of an activity stream - generally I can remember roughly the people involved and the time period. Would be lovely to have search that performs well in this regard.
Sometimes this is legitimate, sometimes spam.
I think it is great product.
The UI takes getting used to, but then again so do many actually useful tools.