Tell NYT, Atlantic, USA Today to keep Wayback Machine
savethearchive.com
savethearchive.com
I'm not sure how to articulate my thoughts on this exactly, other than to say it's disappointing that doing the right thing (i.e. respecting robots.txt) is rewarded with the burden of soliciting responses to a petition while at the same time others are rewarded with profit for ignoring those same directives.
The only reason "others are rewarded with profit" in cases like these are because pinkie-promise-style obligations don't affect players too small or shadowy to bother litigating.
I think you're looking at the wrong end of the spectrum there. It's some of the biggest players who flaunt the rules.
"Several AI companies said to be ignoring robots dot txt exclusion, scraping content without permission: report" (2024) https://www.tomshardware.com/tech-industry/artificial-intell...
This is going to go in a boring direction with an argument thread that's been made since Internet time immemorial, and before. The argument goes: Pirating articles off nyt.com leads to lost sales of subscriptions, so it's not harmless. The response is, inevitably, no it doesn't, it leads to more sales. Or, people who weren't going to pay weren't going to pay anyway, so might as well give it to them for free, and be happy (as the NYT) for free advertising. And then the follow up, "No, it's a lost sale and journalism needs the money." HN is for thoughtful and substantive discussion, not for rehashing the same boring argument we've all read a thousand times. So my question isn't which camp is right. Both camps are firm in their beliefs. Copyright infringement is fine, copyright infringement is not. My question is in today's AI-fueled digital hellscape, how do we support journalists and the arts? If journalism only exists because eg Jeff Bezos pays for the Washington Post, we're going to get biased reporting (which has existed since long before the Internet); If art only exists because the artists come from rich families or have patrons like the Renaissance era, is society better off?
Even if you believe what the AI companies are doing is or should be a copyright violation, the Internet Archive is redistributing in a more direct manner.
User-agent: archive.org_bot
Disallow: /I wonder how archive.org_bot behaves when <meta name="robots" content="noindex, noarchive, nocache" /> is present.
Just out of curiosity, why don't you want your public blog archived? not questioning, just trying to understand the logic/motivations?
Also, I think you're being unfairly downvoted.
> A few months ago we stopped referring to robots.txt files on U.S. government and military web sites for both crawling and displaying web pages (though we respond to removal requests sent to info@archive.org).
Of course not, did you ignore the lines right after? “As we have moved towards broader access it has not caused problems, which we take as a good sign. We are now looking to do this more broadly.”
The announcement is from 9 years ago. I already mentioned they ignored the robots.txt for my own blog.
All of the LLMs would be massively less useful if it wasn't for scraping the latest news.
Every LLM company can afford to spin up a new subscriber account every day, proxying to appear different IPs from all sorts of ASNs, do some crawling until the account gets banned, and then do it again, and again, and again.
What's the conclusion from this train if thought? Just because some burglars can pick locks doesn't mean you should leave your front door unlocked.
Locking a door (or robots.txt) is how one can establish mens rea for those who bypass the barrier.
The actual root cause is that we're allowing LLM companies to completely disregard copyright laws for their profit. Whether the LLM companies scrape the Web Archive or the original source doesn't change the copyright infringement implications in any way, and cutting off the web archive doesn't practically change anything (because as I understand, LLM scraping is already prolific all over the web).
That isn't relevant to ordinary media outlets because a) they don't have enough content volume for rate limiting to be effective since it's possible to get everything they publish even at a slow rate limit, and b) getting AI scrapers to subscribe to their bulk download API instead is not the objective in their case.
There really isn’t even a defensible argument as to how this even should be illegal. The idea that someone can read words about a concept, and then rewording an explanation of that concept somehow violating the rights of the original author, is absurd.
The issue here and elsewhere isn’t LLMs. It’s that IP as a concept has always been a dystopic farce. Despite this we have not only kicked the can down the road on addressing this, we’ve doubled and tripled down and built our society around the concept. The advent of AI has simply blown the scale of the problem up to the point where it cannot be ignored any longer.
How many people do you think use LLMs in some fashion at all in their daily lives? Genuine question, I'm sure my personal experience is a biased sample, but so is everyone else's. Stats from AI companies isn't going to be (seen as) objective either. OpenAI and Anthropic are pushing a feature where I get a situation report at 9am like I'm an important official. With both labs pushing that, I think some people are getting their daily news from LLMs, the question is how many would it take for it to be meaningful, and how would we know if/when that bar gets crossed? What are the implications of that?
Which means LLMs have a zillion sources to get the story. Removing any given subset isn't going to prevent it from having the information in the training data, all it does is prevent that subset from being archived for future humans.
Be a pirate, because a pirate is free...
In the end, we settled on agreeing that making such stuff available after 30 days, and possibly with access restrictions (can’t be pulled more than N times a day, in case it becomes relevant in the future) struck the right balance.
To my knowledge, the Internet Archive hasn’t done any outreach on this issue. In addition to pressuring the publications, I’d put some pressure on them to negotiate.
Solution would be to restrict LLMs from training on the archives. Libraries and universities paying for back access is more complicated.
It's flipped right now. There's no single source of ground truth, but data and information are abundant. Yes, that abundance that includes false data and lies, but it is still abundance.
The work The New York Times and The Atlantic do at their best days, i.e. their investigative journalism team adds to this world, but they try to hide / cloister that work away even though the journalists themselves want to make it accessible.
In an ideal world, every child would learn how to read english via the NYT and The Atlantic, they'd grow up with these sources of record, learn from them, and watch the world through them. But the current model doesn't allow for that.
I think a patronage mixed with wikimedia-style foundation might be a better fit. Readers who love the institution and its mission are invited to pay as much as they want with scaling benefits (let's say you love the NYT so much that you want to give $10k/mo for their work, you should get commensurate access / get to ask questions). And these contributions flow into the endowment, which is invested and the outputs of that are distributed as a part of their operating budget.
I don't think classical journalism can survive an information abundant world without a patronage-based approach.
Maybe. The alternative is most people simply aren’t going to engage with long-form journalism. Keeping the analysis behind subscriptions while video summaries make ad revenue on YouTube and Twitter might be the best fit.
I think the point is all of society isn’t that. Some people still pay for proper journalism. Those who don’t want to don’t have to.
In case it "becomes relevant." Wouldn't that benefit you either way? It makes you wonder if they have a dashboard of unfortunate digital statistics on display somewhere and worship of these numbers have replaced the underlying spirit of journalism.
I’m not in journalism. Becoming relevant down the road would benefit the journal, but they’d also want to get the views in case it converts into subscriptions. Hence the pull limit.
> the underlying spirit of journalism
The places that ignore the business of journalism cable fund the fight true journalism requires. To the extent I think there is a betrayal of fundamentals, it’s in the papers who went with free content driven by ads.
Is the Internet Archive regularly used as a paywall workaround? Generally it's archive.is, which has no connection to the IA.
Too often they’ve been caught selectively reporting details and quotes, or reporting facts from an unreliable source that turned out to be outright false. In the latter case they quietly retract the article, so most readers continue believing the lie (maybe that’s why they don’t want to be archived).
Even posting a small blog is better, while it can also be biased and untrustworthy, if it has original thought, supports an individual, and doesn’t have ads. Although the amount of obvious LLM blogs submitted here is another issue.
The primary source of investigative journalism is the newspaper.
If a NY Times article is corroborated or even paraphrased itself by a more trustworthy organization, or has direct links to multiple primary sources, I wouldn’t mind. Except the NY Times article is still paywalled, and there may be a source that’s not, in which case I still think that source should be submitted instead.
It becomes a research resource. It also creates a high-friction interface for potential subscribers.
I wound up subscribing to Le Monde Diplo because of a HN comment referencing a paywalled article. I didn't want to sign up just for one article. So I bypassed using one of the circumvention sites (I think outline was popular then). The article was compelling enough that I signed up for the paper, and remain subscribed to this day.
You can cryptographically verify a timestamp though by piggybacking on bitcoin like opentimestamps do.
A pie chart showing the times I used the wayback machine to read an old NYT article vs the times I visited it due to a highly upvoted top HN comment linking to a relatively new article so we all can bypass the paywall is a solid circle.
That’s how I signed up to The Atlantic. I wanted to read the Signalgate reporting. There are other publications which get upvoted here frequently that have the paywall workarounds. I generally click around their paywall.
I recommend you actually go and read those fiches. The press was not historically high quality. Mass media has had the same problems for decades.
What it used to have was genuine independent competition.
The work of independent journalists is more important than ever before.
Correct, and they're stealing viewers from these corporate media orgs like NYT. A lot of young people are getting their news from independent journalists on Youtube like Nick Shirley.
What else you got? Alex Jones was an independent journalist? Rush Limbaugh? Tim Poole? Crazy.
They have a robust paying subscriber base that supports them and don't have an owner whose last name rhymes with Pesos who can axe a story just because he doesn't like what it says.
That they publish articles that put Republicans in an unfavorable light is I think because Republicans are doing things that put themselves in an unfavorable light.
To your point, there have been at least a few articles I've seen that put Democrats in an unfavorable light as well.
And for what it's worth I consume news outlets that lean both ways. What's more important to me is factual accuracy.
NYT had $2.82B in revenue in 2025.
The NYT is of course guilty itself. It did not investigate the possible murder of its star witness Suchir Balaji and is too reserved in examining the consequences of AI in general.
If they don't fulfill their journalistic and societal obligations, soon its own journalists will be replaced by AI bullet point slop like Axios.
I'm grown up now, I understand how things work, and I'd rather see Tide and Coke ads than pay $20/mo to 8 different orgs, while maintaining that ad free option for those who want it.
The children of the internet probably won't sign a truce, so let's just cut them out and let intellectually honest people have a decent internet.
I dunno. That seems like a pretty big fuck you to a paying customer already when all they have to do is provide a sub for a few more bucks a month. But I guess I'm a child of the Internet.
If you have an argument against adding a ad-free sku for news outlets I'd love to hear it.
How much faster would consumer software be if adware was made illegal? How much faster would our devices be if we didn't have half the code base supporting malware?
Acting like an ad enabled internet was the only option is extremely foolish, especially when the ad enabled internet was fully chosen and pushed onto the public by very specific people (thanks Newt Gingrich!).
That era vastly predates the Internet, let alone the (relatively) ad-free pre-1980s Internet, neither of which we can return to in any meaningful fashion.
Nope, two problems
1- Ads is privacy issue not only convenience issue. Targeted ads should not normalized.
2- Companies figures out that even paying doesn't means you don't get ads. You probably are bigger target with more disposable income than average in such case.
Ah, so, take the money out of it completely? No subscriptions, and no ads? Sounds like a good idea to me.
I guess I don't really care. As soon as it becomes unworkable to view these publications through archivers I'll just stop viewing them altogether. I don't see this helping their bottom line though.
They also preserved old books. But now I guess they're becoming middlemen for access to limited ebook platforms that ensure books disappear when publishers lose interest.
The "Information Age" is proving to be the setup for a dark age, when nonprofitable things are just thrown out and efforts to preserve them are actively fought.
…why would they go under if the people who don’t pay for news stop reading them?
The paywalls were one thing, but disallowing archival is practically suicide.
The Times alone pulls a multiple of the Internet Archive’s visitors [1][2].
The whole point of archiving is so that people can review it later. People living in the future are the vast majority of readership (and no they didn't pay for it).
The article's place in historical context is far more important than the paper itself. Writing that stands the test of time and that gets cited frequently is where all the authority and credibility comes from. It's absurd that the NYT of all places can be this boneheaded, but I guess it's a sign of the times.
If posting the link instead implies that the 97% of people not currently willing to subscribe can't read it, then people instead post a link to a publication their audience can read, in which case the first publication gets actually 0%.