I've written scrapers over the years, mostly for fun, and I've followed a different approach. Re. "don't interrupt the scrape", whenever URLs are stable, I keep a local cache of downloaded pages, and have a bit of logic that checks the cache first when retrieving an URL. This way you can restart the scrape at any time and most accesses will not hit the network until the point where the previous run was interrupted.
This also helps with the "grab more than you think you need" part - just grab the whole page! If you later realize you needed to extract more than you thought, you have everything in the local cache, ready to be processed again.