Ask HN: Web crawling theory
Hi, I'm looking for documents or books about web crawling, ideally not about a language but more general. Can you help me ? Thank you.
If you wanted to make a plagiarism detector for example, on the pages you index create a histogram of word triads for each indexed web page. Then compare a document for validation with the word triad signatures you created earlier to see if there is a potential match