Waybackpack: download the entire Wayback Machine archive for a given URL
github.com
github.com
This definitely makes it a whole lot easier, wish I had access to it from day 1. Great work guys/gals!!
Shameless plug, if you found it interesting, please consider donating:
Also, a few new users even compliment me on the great design. I guess it isn't that bad ;)
"Sorry, you've already viewed 3 Startup Timelines"
Yes, it is. Why else would you put a 3-startup limit for non-registered users, and present them with a registration page? If the script is malfunctioning, turn it off. Maybe if I could actually see the content I would sign up, but since I have to sign up first, I guess I'll do what most others who are visiting your site for the first time are doing: Go away and never come back.
Do yourself a favor and get rid of intentional annoyances. You're already funding this thing with donations.
Sorry about that - this is a known bug we're trying to fix.
This is no way intentional to try to get you to sign up
Be civil. Don't say things you wouldn't say in a face-to-face conversation. Avoid gratuitous negativity. (-:I'm sure this bug will be fixed shortly, right bakztfuture?
Startup Timelines was always made to be free and accessible, it didn't even ask you to create an account until last month (been running the site for a year now). I don't want anyone to be upset so I've quickly made an account you can use to browse if you've gotten this rogue error:
username: hn_user
password: startuptimelines1
(all accounts are full btw)
I'm sorry and hope this doesn't ruin your take on the site forever. There's a tour that walks you through the site when you register, so, here are screenshots of the pages:
Tour page 1: http://i.imgur.com/5DCwdbg.png
Tour page 2: http://i.imgur.com/o7ghamJ.png
Tour page 3: http://i.imgur.com/iHN775V.png
let me know if there is anything else I can do guys bakz[at]bakzdesign.com ... sorry and thank you again
I can guess what it means but maybe someone here has some insight?
It certainly looks like their Tengine (nginx) servers are configured to expect pipelined requests. It has no problem with greater than 100 requests at a time. See HTTP header above.
Downloading each snapshot one at a time, i.e., many connections, one after the other, perhaps each triggering a TIME_WAIT and consuming resources, may not be the most sensible or considerate approach. If just requesting the history of a single URL, maybe pipelined requests over a single connection is more efficient? I'm biased and I could be wrong.
However their robots.txt says "Please crawl our files." I would guess that crawlers use pipelining and minimize the number of open connections.
I have had my own "wayback downloader" for a number of years, written in shell script, openssl and sed. It's fast.
IA is one of the best sites on the www. Have fun.
Hopefully it gets patched to have a built-in rate limit (X requests per minute/hour).
(It only is traditionally, because so many sites do nothing to protect themselves from "being too nice", so arbitrary-backend mirroring-client devs allow their users the option to ask for less than they want. This isn't a sensible protocol design, on either side; it doesn't optimize for, well, anything.)
It'd be nice if it had identification in the UserAgent, so that we could complain to the right people if it was a problem.
And, yep, the library is intentionally designed only to request one snapshot at a time.
Thanks again for the feedback. Really appreciate it — and the existence of the Internet Archive and Wayback Machine.
Granted, that is not the Wayback Machine but I am sure they love people using them.
There are big privacy issues to getting data from browsers. A lot of websites depend on "secret" URLs, even though that's unsafe, and we don't want to discover or archive those. That means we need opt-in, and smarts.
We do have a project underway with major browsers to send 404s to us to see if we have the page... and offering to take the user to the Wayback if we do.
I call APIs like this "accidental APIs"! From looking at our traffic, we have quite a few programmatic users of it.
More and more people are starting to rely on "archive.is" as it handles Web 2.0 content without issue. But I'm concerned about the survivability of that survice, and whether it can handle big growth.
Using --continue with wget doesn't work (I'm guessing they turned it off).
Also stuff that's been censored for other reasons.
I'm not sure what technologies would be required to implement something like this, but I feel like the Internet Archive would be important and BitCoin might be away of encouraging verification from a globally distributed network of 3rd-party verifiers.