I have a `wget-mirror` shell function invoking wget with all the trimmings that takes care of 99% of sites. I’ll edit the full command into this comment when I get home if anybody else wants to start doing the same :)
I have a `wget-mirror` shell function invoking wget with all the trimmings that takes care of 99% of sites. I’ll edit the full command into this comment when I get home if anybody else wants to start doing the same :)
https://en.wikipedia.org/wiki/WWWOFFLE
https://ftp.netbsd.org/pub/pkgsrc/distfiles/wwwoffle-2.9j.tg...
The way the www is going, it seems like downloading a copy of libgen, i.e., nonfiction books, and scimag, i.e., academic journals, via torrent, would be more valuable than archiving websites, in general. These primary sources are part of the material used to train so-called "AI" anyway. The problem is that this so-called "AI" also includes all the garbage from the www.
Worst case is eventually these books and journals will again become publicly inaccessible but "AI" will be offered as a bogus substitute; a future where few people will do research using primary materials anymore, they will just submit questions to a remote "AI" server. Truth will be decimated.
https://blusharkmedia.medium.com/the-ongoing-battle-against-...
https://techhq.com/2023/09/can-libgen-shadow-library-survive...
https://www.twitter.com/theshawwn/status/1320282152689336320
https://qz.com/openai-books-piracy-microsoft-meta-google-cha...
https://qz.com/shadow-libraries-are-at-the-heart-of-the-moun...
https://goodereader.com/blog/e-book-news/authors-file-lawsui...
When asked about whether this was true, they refused to answer based on confidentiality concerns, then said they had deleted all copies of the dataset, stopped using it, and no longer employed the individuals that compiled it:
https://www.businessinsider.com/openai-destroyed-ai-training...
We do know for a fact that the (non-OpenAI-controlled) "Books3" dataset is just "all of bibliotik":
https://www.twitter.com/theshawwn/status/1320282149329784833
https://github.com/soskek/bookcorpus/issues/27
And we also apparently know for a fact that this was included in the datasets used to train LLAMA:
https://en.wikipedia.org/wiki/The_Pile_(dataset)
https://aicopyright.substack.com/p/the-books-used-to-train-l...
https://aicopyright.substack.com/p/has-your-book-been-used-t...
https://news.ycombinator.com/item?id=40258584
https://arxiv.org/pdf/2005.14165.pdf
https://www.wired.com/story/battle-over-books3
https://www.washingtonpost.com/technology/interactive/2023/a...
https://www.theguardian.com/technology/2023/apr/20/fresh-con...
https://storage.courtlistener.com/recap/gov.uscourts.cand.41...
See 40-45.
https://storage.courtlistener.com/recap/gov.uscourts.nysd.60...
See 87-116.
(Oh, never mind YouTube videos that I once added to playlists ... that later disappear leaving only holes in my playlists.)
It is probably easiest to save the render as a picture and then store text separately for searchability?
Chromium's MHTML "Save as…" and the SingleFile WebExtension should both save copies of the rendered DOM.
Apparently Safari has WebArchive and Mozilla had MAFF for similar use cases.
I think WARC is supposed to save enough data about network streams for dynamic pages to work. At least on the Wayback Machine, infinite scrolling and "Load More" buttons do kinda work sometimes. You may have to load the archived pages in a browser and try to use each dynamic feature at least once, to trigger requests for needed resources.
SingleFile: https://github.com/gildas-lormeau/SingleFile
LWN on WARC, tools: https://anarc.at/blog/2018-10-04-archiving-web-sites/
Self-hostable web archives: https://awesome-selfhosted.net/tags/archiving-and-digital-pr...
Wayback Machine addons, bookmarklets: https://help.archive.org/help/save-pages-in-the-wayback-mach...
I donate to The Archive. More people should too.
Plus for as great of a service as Wayback Machine is, it can be very unpleasant to actually browse. I dislike how it injects its own toolbar into every page (yes I know how to massage the URLs to get the raw page data, but it isn't browsable that way). Have you never encountered sites in Wayback Machine where certain pages were just randomly not archived? Or when you click a link and get a page from years earlier or later than the one you came from? Never encountered a page or an entire domain that was blocked from Wayback Machine? Why do you think I would get started doing something like this in the first place if I didn't find it more fun to browse my own archives than Somebody Else's?
wget \ --recursive \ --mirror \ --timestamping \ --page-requisites \ --html-extension \ --convert-links \ --restrict-file-names=windows \ --no-parent \ $url
I would love that. I have a little for parameter version, but I feel yours is more tried and true.
wget-mirror() {
wget --mirror --convert-links --adjust-extension --page-requisites \
--no-parent --content-disposition --content-on-error \
--header="Accept: text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8" \
--user-agent="Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:129.0) Gecko/20100101 Firefox/129.0" \
--restrict-file-names="windows,nocontrol" -e robots=off --no-check-certificate \
--no-hsts --retry-connrefused --retry-on-host-error --reject-regex=".*\/\/\/.*" $1
}
Some notes:— This command hits servers as fast as possible. Not sorry. I have encountered a very small number of sites-I-care-to-mirror that have any sort of mitigation for this. The only site I'm IP banned from right now is http://elm-chan.org/ and that's just because I haven't cared to power-cycle my ISP box or bother with VPN. If you want to be a better neighbor than me, look into wget's `--wait`/`--waitretry`/`--random-wait`.
— The only part of this I'm actively unhappy with is the fixed version number in my fake User-Agent string. I go in and increment it to whatever version's current every once in a while. I am tempted to try automating it with an additional call to `date` assuming a six-week major-version cadence.
— The `--reject-regex` is a hack to work around lots of CMS I've encountered where it's possible to build up links with an infinite number of path separators, e.g. an `www.example.com///whatever` containing a link to `www.example.com////whatever` containing a link to…
— I am using wget1 aka wget. There is a wget2 project, but last time I looked into it wget2 did not support something I needed. I don't remember what that something was lol
— I have avoided WARC because I usually prefer the ergonomics of having separate files and because WARC seems more focused on use cases where one does multiple archives over time (as is the case for Wayback Machine or a search engine) where my archiving style is more one-and-done. I don't tend to back up sites that are actively changing/maintained.
— However I do like to wrap my mirrored files in a store-only Zip archive when there are a great number of mostly-identical pages, like for web forums. I back up to a ZFS dataset with ZSTD compression, and the space savings can be quite substantial for certain sites. A TAR compresses just as well, but a `zip -0` will have a central directory that makes it much easier to browse later.
Here is an example of the file usage for http://preserve.mactech.com with separate files vs plain TAR vs DEFLATE Zip archive vs store-only Zip archive. These are all on the same ZSTD-compressed dataset and the DEFLATE example is here to show why one would want store-only when fs-level compression is enabled.
982M preserve.mactech.com.deflate.zip
408M preserve.mactech.com.store.zip
410M preserve.mactech.com.tar
3.8G preserve.mactech.com
Also I lied and don't have a full TiB yet ;) [lammy@popola#WWW] zfs list spinthedisc/Backups/WWW
NAME USED AVAIL REFER MOUNTPOINT
spinthedisc/Backups/WWW 772G 299G 772G /spinthedisc/Backups/WWW
[lammy@popola#WWW] zfs get compression spinthedisc/Backups/WWW
NAME PROPERTY VALUE SOURCE
spinthedisc/Backups/WWW compression zstd local
[lammy@popola#WWW] ls
Academic DIY Medicine SA
Animals Doujin Military Science
Anime Electronics most_wanted.txt Space
Appliance Fantasy Movies Sports
Architecture Food Music Survivalism
Art Games Personal Theology
Books History Philosophy too_big_for_old_hdds.txt
Business Hobby Photography Toys
Cars Humor Politics Transportation
Cartoons Kids Publications Travel
Celebrity LGBT Radio Webcomics
Communities Literature Railroad
Computers Media README.txt
Some of this could stand to be re-organized. Since I've gotten more into it I've gotten better at anticipating an ideal directory depth/specificity at archive time instead of trying to come back to them later. Like `DIY` (i.e. home improvement) there should go into `Hobby` which did not exist at the time, `SA` (SomethingAwful) should go into `Communities` which did not exist at the time, `Cars` into `Transportation`, etc.`Personal` is the directory that's been hardest to sort because personal sites are one of my fav things to back up but also one of the hardest things to try and organize when they reflect diverse interests. For now I've settled on a hybrid approach. If a site is geared toward one particular interest or subsulture, it gets sorted into `Personal/<Interest>`, like `Academics`, `Authors`, `Artists`, `Goth` (loads of '90s goths had web pages for some reason). Sites reflecting The Style At The Time might get sorted into `1990s` for a blinking-construction-GIF Tripod/Angelfire site or `2000s` for an early blog. Some times I sort personal sites by generation like `GenX` or `Boomer` (said in a loving way — Boomers did nothing wrong) when they reflect interests more typical of one particular generation.
I have encountered "GnuTLS: The TLS connection was non-properly terminated. Unable to establish SSL connection." multiple times, and retry options seem to be useless when that happens. Some searches suggest it could be related to tls handshake fragmentation, but nonetheless wget could retry if related options are used. Manual retry seems to download the missing URLs, otherwise mirroring jobs are randomly incomplete.