I see WP is not proposing to run its own.
I see WP is not proposing to run its own.
Like Wikipedia?
1) provides a snapshot of another site for archival purposes. 2) provides original content.
You're arguing that since encyclopedias change their content, the Library of Congress should be allowed to change the content of the materials in its stacks.
By modifying its archives, archive.today just flushed its credibility as an archival site. So what is it now?
As an end user of Wikipedia there are occasions where content has been scrubbed and/or edits hidden. Admins can see some of those, but end users cannot (with various justifications, some excellent/reasonable and some.. nebulous). That's all I'm saying, nothing about Congress or such other nonsense. It seems like an occasion of the pot calling the kettle names from this side of the fence.
An archival site (by default definition) promises you that it will not modify its content. And when it does, it's no longer an archival site.
Wikipedia has never been an archival site and it never will be. archive.today was an archival site, but now it never will be again.
Meanwhile their IMA on Reddit: no promises, no commitment. Just like Microsoft EULA :)
https://old.reddit.com/r/DataHoarder/comments/1i277vt/psa_ar...
I'm quoting all of that because is lacks an explicit promise of non-modification /i
Meanwhile seriously, if you were disappointed not to see e.g. "We explicitly don't promise not to modify", then perhaps you should consider why, regardless, this site was trusted enough to get a gazillion links in Wikipedia... and HN.
And I'm quoting all of that because it lacks an explicit (or implicit) promise of modification. :)
It was (emphasis on past-tense) so-trusted because it advertises itself as an archival site. (The linked disclaimer is all about it not being a "long-term" archival site. It says it archives pages for latecomers. There is an implication here that it archives them accurately. What use is a site for latecomers if they change the content to be something else?) If they'd said or indicated they would be changing the content to no longer reflect the original site, Wikipedia would not have linked to them because they wouldn't be a credible source.
In any case, now I can't use them to share or use links since we can no longer trust those archives to be untampered. When I share a link to nyt content on archive.today or copy and paste content into email, I'm putting my name on that declaring "nyt printed this". If that's not true, it's my reputation.
Just like it was archive.today's.
What if the nyt article itself is the problem? How does that square?
What's your better idea?
Archive.org snapshots may load javascript from external sites, where the original page had loaded them. That script can change anything on the page. Most often, the domain is expired and hijacked by a parking company, so it just replaces the whole page with ads.
Example: https://web.archive.org/web/20140701040026/http://echo.msk.r...
----
And another example: https://web.archive.org/web/20260219005158/https://time.is/
The page "got changed" every second. It is easy to make an archived page which would show different content depending on current time or whether you have Mac or Windows, or your locale, or browser fingerpring, or been tailored for you personally
Isn't there a substantial overlap with the copyright holders?
> Internet archives wayback machine works as alternative to it.
It is appalling insecure. It lets archives be altered by page JS and deleted by the page domain owner.
Nonstarter for anything that you actually want to be preserved, especially anything controversial.
Yes, they are essentional, and that was the main reason for not blacklisting Archive.today. But Archive.today has shown they do not actually provide such a service:
> “If this is true it essentially forces our hand, archive.today would have to go,” another editor replied. “The argument for allowing it has been verifiability, but that of course rests upon the fact the archives are accurate, and the counter to people saying the website cannot be trusted for that has been that there is no record of archived websites themselves being tampered with. If that is no longer the case then the stated reason for the website being reliable for accurate snapshots of sources would no longer be valid.”
How can you trust that the page that Archive.today serves you is an actual archive at this point?
Oh dear.
> How can you trust that the page that Archive.today serves you is an actual archive at this point?
Because no-one shown evidence that it isn't.
Wikipedia does not have a project page with this exact name.
I assume that is weasel words for 404 Not Found.
To https://en.wikipedia.org/wiki/Wikipedia:Requests_for_comment...
I read that up to the first "proof", https://web.archive.org/web/20260218135501/https://www.googl...
It lands "503 Service Unavailable No server is available to handle this request."
But this one is not credible either so...
ArsTechica just did the same - removed Nora from older articles. How can you trust ArsTechica after that?
I don't know what you're talking about re: Ars removing her name from old articles.
2. We learned about Nora's involvement from Patokallio. We learned about Nora's non-involvement... also from Patokallio. They could have reached a settlement with AT that includes hiding Nora's name.
3. Regardless of who Nora is, it is interesting to see the extent of this censorship: so far only gyrovague.com and arstechnica.com, but not tomshardware.com and not tech.yahoo.com. This shows which sites are working closely with the AT defamation campaign, and which are simply copywriting the news feed.
If AT is appropriating some random person's name as an alias, it seems helpful to report on that publicly in order to expose the practice and help clear up the misinformation.
One with title 'Archive.today CAPTCHA page executes DDoS; Wikipedia considers banning site'
I'll try to add the link with comment edit:
This has Nora's name https://web.archive.org/web/20260210195502/https://arstechni...
The current version has not
I am lost here. It is definitively an organized defamation campaign.
“You are guilty simply because I am hungry”
And again, the accusation against Archive.today isn't just that they removed their "Nora" alias from a snapshot, but that they replaced it with the name of the blogger they were quarreling with. There's no defensible reason to do that outside of petty revenge (which tracks with the emails and public statements from the Archive.today maintainer).
Oh, yes, by removing the name in the context of "Streisand Effect".
> petty revenge
How does it "revenge"? Was it a porn page? Or something bad?
It is likely to be just a funny placeholder name of the same length to come in mind.
--
We could find good and bad motives for both AT and Ars.
The bias against AT was here apriori. Paywall-story for CondeNast, russophobia for the rest.
The porn smear threats came later, via email.