Wayback Machine Downloader
github.com
github.com
Browsing through crawls has this neat side-effect of being able to serendipitously discover things that I missed back in the day just by having everything laid out on the file system.
PSA: There's a lot of holes in most crawls, even for popular stuff. A good way to ensure that you can revisit content later is submitting links to the Wayback Machine with the "Save Page Now" [1] functionality. Some local archivers like ArchiveBox [2] let you automate this. Highly recommended to make a habit of it.
1. The parent comment you're replying to links to the main page for the Wayback Machine, which includes a Save Page Now widget, but Save Page Now actually has a dedicated page <https://web.archive.org/save/>
2. If you have an archive.org account (lets you submit and comment on collections; the library is bigger than just the Wayback Machine) and you visit the Save Page Now page while logged in, you get more options, including the option "Save outlinks"
I had an old website of mine (my old video game portafolio) that I wanted to bring back to life. I have no sources and no backups. But it was still on the Wayback Machine! I first wrote a quick wrapper in Ruby and it worked fine. I then decide to open source it and publish it. It was a fun adventure to see this being used by so many! <3
PS: could you merge my two pull requests? ;)
That is not a tip, it's a dirty workaround.
Heh, I just realized SO is a potential source of energy if we could harness the copy-paste to SO question to copy-paste cycle. Blockchain! :-P
Yeah, but at the same time, if an attacker does have no-sudo access to a machine, everything interesting is most likely already compromised. Sudo does seem an hardly justifiable complication in most cases.
Filtering out unwanted domains (sale, spam etc) is a problem for a bunch of regexes, bayesian classification or machine learning.
Edit: I think I misread the original post quite badly and I don't understand the proposed feature.
On the surface a Yelp like system to rate domains as legit vs. click bait seems logical until you realize the scammers would just work at gaming that system too :p
This is where a universal ID would really help - but the other ways something like that could be used make me even more uncomfortable so here we are with no real good solution :(
To reduce the burden of writing specific scraping software, I investigated the software listed by the Archive Team[0].
The command line tools and libraries aside (as they would require much more specific tailoring to make them work), I was particularly interested in HTTrack and Warrick.
Warrick[1] is defined as a tool to recover lost websites using various online archives and caches. Warrick[2] requires some expert knowledge in order to get it up and running.
I found Warrick was a bit outdated, so decided to try something similar but more up to date, and came across this (hartator/wayback_machine_downloader).
I found this to be a bit easier to work with as it would allow me to download snapshots within a time period, which is what I needed for this project.
After running it for ~12 hours on my local machine, it still had not completed, downloading only 11768 of 94518 files.
Instead, I found myself writing a Python based tool that could fetch with much more accuracy using the CDX server[3] and filtering by date, targeting only certain fails and allow for multi threading.
In order to improve the process, and narrow down the data of what is needed, I scoped to just July and December timeframes per year, from 2010 to this year, only targeting html files.
For example: http://web.archive.org/cdx/search/cdx?url=%s&matchType=prefi...
Hopefully someone finds this useful.
[0] https://archiveteam.org/index.php?title=Software
[1] https://github.com/oduwsdl/warrick
[2] https://code.google.com/archive/p/warrick/wikis/About_Warric...
[3] https://github.com/internetarchive/wayback/tree/master/wayba...
[0] http://web.archive.org/web/20210119093354/http://riotfish.co...
wget -r -np -k "https://web.archive.org/web/DATENUMBER/URL"
The CDX basically lists all the pages of the site that are archived in the wayback, as well as when each page was archived (e.g: "page x was archived 3 times, on these dates"). Using the CDX allows the tool to download a specific copy of the site (e.g: the latest) rather than trying to download every copy of the site that the wayback machine has.
This is important because for most sites, the wayback has multiple copies, and they're all interlinking. For example, the copy from May 2020 might not be complete so one of the links in that copy will take you to the January 2018 copy. Not a problem for a human viewer, but a bot / crawler will see pages in the January 2018 copy as separate from those in the May 2020 copy, so will begin downloading the January 2018 copy (because wayback URLs are of the form web.archive.org/<timestamp>/<archive-url> rather than web.archive.org/<archive-url>/<timestamp>). This copy will (inevitably) lead to other copies made at different dates, and before you know it you're downloading hundreds or even thousands of copies of the same site.
[source: tried to download a site from the wayback machine several years ago using wget - it didn't end well!]
I don’t understand this. Did they mean “e.g.” instead of “i.e.”?
This project started in 2015 btw. Another similar project called waybackpack started in 2016. There are probably more projects. IMO wayback-machine-downloader is the better project though.
https://github.com/jsvine/waybackpack
The Wayback CDX Server API these projects are based on is quite simple to use btw, just some JSON responses to decode.
https://archive.org/help/wayback_api.php https://github.com/internetarchive/wayback/blob/master/wayba...