HNHacker News
TopNewBestAskShowJobs

wolfgarbe

350 karma · joined May 7, 2017

Founder @ seekstorm.com Search-as-service Maintainer @ SeekStorm - sub-millisecond full-text search library & multi-tenancy server in Rust https://github.com/SeekStorm/SeekStorm Maintainer @ SymSpell spelling correction https://github.com/wolfgarbe https://www.linkedin.com/in/wolfgarbe/ https://www.quora.com/profile/Wolf-Garbe/answers https://wolfgarbe.medium.com/
submissionscomments
wolfgarbe··on Show HN: DeepHN – Full-text search of 30M Hacker News posts and linked webpages
>> you were crawling news.ycombinator.com, right?

No, for retrieving the Hacker News Posts we were using the public Hacker News API, which returns the posts in JSON format: https://github.com/HackerNews/API

The crawling speed of 100...1000 pages per second refers to crawling the external pages linked from Hacker news posts. As they are from different domains we can achieve a high crawling speed while being a polite crawler with a low crawling rate per domain.

wolfgarbe··on Show HN: DeepHN – Full-text search of 30M Hacker News posts and linked webpages
No, the preview is NOT the representation of what we have indexed. The preview ist limited to about 200 words to be compliant with fair use legislation. Indexing is limited to 1 MB per document.
wolfgarbe··on Show HN: DeepHN – Full-text search of 30M Hacker News posts and linked webpages
DeepHN does not copy the whole page, only the first 200 words or less. DeepHN does not copy the whole website but only a fraction of a single page of that web site. We always provide the original link. That should make it fair use. Of course we would prefer if there was a clear and concise legislation, valid in all countries, which could be followed.

We mainly surface historical content, which receives additional traffic from DeepHN, instead of taking value away from the website.

wolfgarbe··on Show HN: DeepHN – Full-text search of 30M Hacker News posts and linked webpages
Thank you all for your great feedback. We will fix all the bugs and add some of the suggested features and release a next iteration of DeepHN within the next few days!
wolfgarbe··on Show HN: DeepHN – Full-text search of 30M Hacker News posts and linked webpages
It should, but it currently doesn't in DeepHN. We will fix this asap.
wolfgarbe··on Show HN: DeepHN – Full-text search of 30M Hacker News posts and linked webpages
Thank you. We will look into this and fix it asap.
wolfgarbe··on Show HN: DeepHN – Full-text search of 30M Hacker News posts and linked webpages
SeekStorm uses dedicated root servers with NVMe SSDs, hosted by IONOS https://www.ionos.com/servers/intel-servers
wolfgarbe··on Show HN: DeepHN – Full-text search of 30M Hacker News posts and linked webpages
We will look into this.
wolfgarbe··on Show HN: DeepHN – Full-text search of 30M Hacker News posts and linked webpages
You mean performance in terms of latency or relevancy?
wolfgarbe··on Show HN: DeepHN – Full-text search of 30M Hacker News posts and linked webpages
Currently the maximum size of a single document is limited to 1 MByte. That is an artificial limitation resulting from the fact, that SeekStorm does not only indexes the content of the document, but also stores the original document. Somewhere there has to be a limit.
wolfgarbe··on Show HN: DeepHN – Full-text search of 30M Hacker News posts and linked webpages
https://seekstorm.com/about-deephn

https://seekstorm.com/

https://seekstorm.com/search-api-features

https://seekstorm.com/search-as-a-service-architecture

wolfgarbe··on Show HN: DeepHN – Full-text search of 30M Hacker News posts and linked webpages
"NIOFSDirectory and MMapDirectory implementations face file-channel issues in Windows and memory release problems respectively. To overcome such environment peculiarities Lucene provides the FSDirectory.open() method. When invoked, it tries to choose the best implementation depending on the environment." https://www.baeldung.com/lucene-file-search
wolfgarbe··on Show HN: DeepHN – Full-text search of 30M Hacker News posts and linked webpages
With the crawler which is part of the SeekStorm search-as-a-service.
wolfgarbe··on Show HN: DeepHN – Full-text search of 30M Hacker News posts and linked webpages
If you were entering deephn.org in your browser address bar, you couldn't use the browser back button to return to the previous page. That was a bug caused by the navigation code within our dynamic HTML page. It is fixed now.
wolfgarbe··on Show HN: DeepHN – Full-text search of 30M Hacker News posts and linked webpages
That would definitely work. We are working on local client that could index your local PDF and Word documents. So importing emails would be a nice addition.
wolfgarbe··on Show HN: DeepHN – Full-text search of 30M Hacker News posts and linked webpages
Good idea!
wolfgarbe··on Show HN: DeepHN – Full-text search of 30M Hacker News posts and linked webpages
The crawler is a part of SeekStorm.

"Well-defined": Just a guess: We are doing key text extraction, i.e. we try not to index boilerplate stuff an and menu items. As "well-defined" is within a short list item, it might be accidentally skipped.

So, that is not yet perfect, and considering the diversity in web page structure it probably never will. But we will try to improve.

wolfgarbe··on Show HN: DeepHN – Full-text search of 30M Hacker News posts and linked webpages
>> When trying to go back after clicking the original HN link, there's no response.

The original HN link is opened in a new browser page/tab. The search results should remain unchanged in the previous tab. So you don't use the back button, but the previous browser tab.

>> Clicking on what I understand are tags (hashtags) in a given post has no effect.

Search for "google", go to the first result. there is a hashtag "oracle", click and the search results are filtered by that hashtags. Please post an example where its doesn't work, so that we can fix it.

The date is shown in the preview panel on the right hand side. But wa are thinking adding the time also to the result list.

We are deriving the tags from the terms and bigrams in title, text and parsed html of linked web pages. Top frequent terms per post, but only if terms are within top 65k tags per index. Stopwords are excluded.

wolfgarbe··on Show HN: DeepHN – Full-text search of 30M Hacker News posts and linked webpages
Not necessarily a spam link. Linked pages from older post have sometimes updated their content.

Sometimes this is legitimate, sometimes spam.

wolfgarbe··on Show HN: DeepHN – Full-text search of 30M Hacker News posts and linked webpages
Crawling speed is between 100...1000 pages per second. We crawled about 4 million linked unique web pages

The pricing for an Index like DeepHN would be $99/month: While we are indexing 30 million Hacker news posts, for DeepHN we are combining a single HN story with all its comments and its linked webpage into a single SeekSorm document. So that we index just 4 million SeekStorm documents.

Yes, it would be possible to expand index with embeddings (vectors) and perform semantic search. This would we an auxiliary step between crawling and indexing.

wolfgarbe··on Show HN: DeepHN – Full-text search of 30M Hacker News posts and linked webpages
We can add that button, if you promise not to kill ;)
wolfgarbe··on Show HN: DeepHN – Full-text search of 30M Hacker News posts and linked webpages
There is no secret about the benchmark, its all open source: https://github.com/wolfgarbe/LuceneBench
wolfgarbe··on Show HN: DeepHN – Full-text search of 30M Hacker News posts and linked webpages
Yes, it does exactly that. If you check "Web" on the left sidebar, and uncheck "Stories" it will find the posts where "friendship" occurs in the linked web pages. But we still show the title of the original hacker news story in the search result and highlight the search term if it ALSO occurs in the title.

If you scroll long enough within the search results you will finally reach results where the search term is not in the title, but only within the linked web page.

It is more obvious if you search for queries which are less popular than "friendship"

wolfgarbe··on Show HN: DeepHN – Full-text search of 30M Hacker News posts and linked webpages
We are combining auto correction (SymSpell) and auto completion (PruningRadixTrie). That is sometimes tricky. We are using a static spelling correction dictionary which does not contains the term algolia, but angola. SymSpell is very fast, but requires a lot of memory. Therefore we are using a static dictionary, and not a dynamic dictionary derived from each individual index of each customer. Auto completion on the other hand we are dynamically generating for each individual index/customer.
wolfgarbe··on Show HN: DeepHN – Full-text search of 30M Hacker News posts and linked webpages
You are right. The preview is far from perfect. We just wanted to ship early. We will continue to fix and improve.
wolfgarbe··on Show HN: DeepHN – Full-text search of 30M Hacker News posts and linked webpages
DeeepHN is powered by our SeekStorm search-as-a-service, which has a full-featured API: https://seekstorm.com/docs

SeekStorm is intended that the user can index and search their own private data.

But if there is demand we could also allow to search in a central public index for web data or other public data.

wolfgarbe··on Show HN: DeepHN – Full-text search of 30M Hacker News posts and linked webpages
We are deriving the tags from the terms and bigrams in title, text and parsed html of linked web pages. Top frequent terms per post, but only if terms are within top 65k tags per index. Stopwords are excluded.
wolfgarbe··on Show HN: DeepHN – Full-text search of 30M Hacker News posts and linked webpages
We have just fixed the back button bug (make sure to clear the browser cache).
wolfgarbe··on Show HN: DeepHN – Full-text search of 30M Hacker News posts and linked webpages
Your are right. It IS a bug, and we will fix it asap.
wolfgarbe··on Show HN: DeepHN – Full-text search of 30M Hacker News posts and linked webpages
Thanks for pointing this out. That was not intentionally, probably a side effect from the navigation inside our search results within a dynamic page (using the back button to return to previous search queries). We will fix this.
← PreviousPage 2 of 3Next →