HTTrack Website Copier – Offline browser
httrack.com
httrack.com
I haven’t followed httrack since, but it seems like scrapy and similar are much better replacements.
You'll probably find better approaches, and while I never tried scrapy, it seems to be using a javascript engine for hard cases, which was something I thought about (but this was way above my skills at that time).
The hard parts remains however, if you want a functional site: you need to rewrite links, or use an external proxy-like mechanism. Having a fully functional offline, file-based site, is the real tricky part. Cases will remain unsolvable, as the inside code logic can produce whatever external link resource based on randomness, time, etc.
The approach in httrack was both ugly and pragmatic: attempting to recognize link/files patterns within javascript and fetch/replace what can be replaced with local links. Javascript producing html will typically be analyzed with really dumb - yet sometimes effective - js parsers. (parental advisory: don't look at the parsers code, your eyes would melt)
And obviously this is not going to solve all cases and will even break pages with tricky js
But, generally speaking, being able to preserve "the internet" by saving whole websites offline should be something we give more attention to.
Just read this recently: https://www.theatlantic.com/technology/archive/2021/06/the-i...
httrack was extremely helpful and there really was no equal. The “modern” web requires a live JS engine, but as you point out, even the “old” web had server-side logic that couldn’t be captured.
In that light, I think httrack has stood up pretty well and nobody expects you to go rewrite it or clean it up. If someone today has a mostly static site they want to archive without writing custom code, I would still recommend httrack (it’s more controllable than wget or similar). I just assume that those sites are mostly gone :(.
[0] https://github.com/chowderman/hyperfiler
* disclaimer: I created HyperFiler
wget -E -r -k -p --span-host http://mycoolhomepage.comA tool to crawl all of the links within a website and submit each one of them to Wayback...
is an interesting solution to GP's issue, works better than httrack and can also pull down multiple timestamps of archived site
Disclaimer: I tried to get involved with ArchiveTeam to help get some websites that I care about properly archived and saved to The Way Back Machine, but they weren’t allowing new members to sign up. They were happy to talk to me about what needed to be done and one of their members set things up to archive everything that was available, but I wasn’t able to help in that process.
I really like the tool. I doubt if that is helpful today, bc. of the raise of the Javascript stuff...
[0] https://addons.mozilla.org/en-US/firefox/addon/single-file/