Doesn't mention the hardest part I found when developing a crawler - dealing with pages whose content is mostly dynamic and generated client side (SPA's). Even using V8 it's hard to do reliably and performantly at scale.
Given this is from 2004 I'm not surprised.
Often those types of pages are difficult to link to as well, as they're often highly stateful, applications rather than documents, and what you're seeing may not be what you get when you click the link. Not really what you want in a search engine.