Katana: A crawling and spidering framework
github.com
github.com
IMO the hardest things in distributed crawling at scale are a good URL frontier, priorities, rate limiting and things like that, which are quite often overlooked.
Fast Ethernet is my favorite example.
[0] (translated title)
(Planck length is 10e-35. Even the strong nuclear force operates on a scale that's like 20 orders or magnitude larger (10e-15). And a hugantuan electron? Forget about it.)
Could contact me? We may have some interests in common. Check my profile.
Are any of you other HNers finding the web increasingly difficult to scrape from?
(EDIT: Yes, I figured it out. To get around the 301 loop. cURL needs to save cookies)
Large scale crawling is primarily a challenge in balancing the logistics in a way that is kind to both the crawler and the data consumers.
Distributed crawling, if you go that way, is also non-trivial as you're effectively juggling a shared rapidly mutating state in the dozens gigabytes.
It’s already _technically_ impossible to erase something from the internet, but if they removed the barrier to knowing where something was before in order to find it in the archive, it would be truly impossible in every sense of the word.
<https://en.wikipedia.org/wiki/Alexa_Internet>
<https://help.archive.org/help/wayback-machine-general-inform...>
A condition of that sale was that Alexa would continue to provide the results of its crawls, after a delay, to the Internet Archive. Those crawls form a substantial portion of IA's Wayback Machine archive.
I'm not certain that those archive are ongoing, as Alexa seems to have been largely shut down.
IA are a bit cagey on details, but I believe that there is a general IA-based archival service. There's certainly the "Save Page Now" feature:
https://web.archive.org/save/<URL>
And the independent but closely-cooperating ArchiveTeam (lead by Jason Scott) tailors crawlers specific to endangered / vulnerable online websites, its Warrior software: