How to deal with duplication is also a very difficult problem because there's loads of reasons why things could be duplicated. Take a textbook, I've seen duplicates which contain either one or several of the following: different editions, different printings (of any particular edition), added bookmarks/table of contents for the file, removed blank white pages, removed front/end cover pages, removed introduction/index/copyright/book information pages, LaTeX'd copies of pre-TeX textbooks, OCR'd, different resolution, other kinds of optimization by software that reduces to wildly different file sizes, different file types (eg .chm, PDFs that are straight conversions from epub/mobi), etc. Some of this can be detected by machines, eg usage of OCR but some of the other things aren't easy at all to detect.