Thanks for the suggestions. robots.txt should now be blocking the /db/, which has all saved content, and a link has been added to the DMCA page on every generated page (putting it at the bottom would be obscure since the pages can get so long).
I'm not planning on copying any of the actual HN content, and don't present copy at all if it is on news.yc. At some point I'll hook into the API to grab comments/points every so often to update into the index pages and probably allow voting from the pages.