Check out the Nutch crawler. It closely integrates with the Lucene search engine project.
If you want something which is very specific (like the hype machine) you can easily create a basic crawler in python. There is not much change in the crawling strategies. Just a basic programming knowledge and reading the documentation of the urllib module (in python) would be enough.