Looks like a good idea.
Is the idea to crawl locally? How do you crawl - sequentially or in parallel?
Is the idea to crawl locally? How do you crawl - sequentially or in parallel?
For the full text portion we send out a request to the URL for the full HTML, distill it (similar to "reader" mode in some browsers) and add that distilled content to the full text index.
The full-text generation happens with some concurrency but we intentionally didn't want to spam anyone's server so it's limited and takes a while to populate.