"How did you detect "decay"? Just based on HTTP codes, or by actually looking at the linked content? On my own blog, I found more often than I like that old links still "work" per HTTP, but now refer to something rather different from the content that I originally intended to refer to."
What would be some good way to detect such spam sites in an automated way? Looking for the link's title in the remote HTML? Check for common domain placeholder page contents and spam words? Maybe Google has some API one could use?