Reddit will block the Internet Archive
theverge.com
theverge.com
The datahoarders of the world probably have somewhat decent archives of quite a few youtube channels and such, but since it's not publically availible, and reuploading to youtube or IA isn't really viable, it's lost as far as any of the rest of us are concerned.
Doesn't help that video is really big and cumbersome to archive. Audio's a lot smaller and thus easier to keep around. Text is easiest, but there tends to be a lot of it, and archiving one page at a time is usually not worth it in the same way that a video or song might be, so blocking automation is usually a pretty good bet for anyone who really doesn't want their stuff to be archived.
We are talking archives, older archival processes were not much better. At best it would arrive in a public information nexus, but be filed away in a dark filing cabinet, not much different than data hoarding. At least data hoarding can interface with the wider internet incredibly easily should there be a desire.
Although I would agree that you are right and that there needs to be a better long-term infrastructure for this.
I think back to the original peer-to-peer applications and almost think that that loose framework would be an improvement - where different people with files with the same digital signature can serve as mirrors for a reference file through a decentralized network.
Good luck finding someone's archive of a Youtube channel. Sure it probably exists, probably even several copies of it, but unless XxX_PussyDestroyer69_XxX on reddit wrote in a comment that they personally archived that particular channel, you're never going to be able to get in touch with the relevant people to actually get a hold of what you're looking for.
I wish peer-to-peer video ended up being more popular than it is, if only to have an obvious place to put those archives people otherwise juat hoard foe themselves, but I have little faith that will happen for as long as Youtube remains the only viable video platform out there for your average Joe.
Is it? Its very easy to produce (hence there's too many of them) and they are extremely fragile (bit rot, complicated format that no one knows how to parse etc). Seems to be this is inevitable. I personally think Youtube is going to start pruning their database in the next decade.
Whether you classify that as "AI-related" or not, I don't know.
Basically laundering outright wrong information into something the next generation is now going to believe as scientific truths.
I often wonder how many people/organizations are seeding places like Reddit with malinformation/beliefs for it to become canical truth in the AI age once it's too late to tell the difference for most people?
Lord knows I've made trolling-level posts that are only marginally accurate back in the day that are now part of the AI corpus of knowledge. Mix those in with some of my well-researched stuff and you couldn't even really filter it based on "this account is a shitposter" to weight it lower. Nevermind plenty of earnest posts made that were outright wrong simply due to... being wrong in the moment and later learning better.
Metadata is what makes gold out of poo, I assume model developers can "train negatively" too if metadata suggests they should.
Internet Archive has been terrible as capturing full pages on Reddit for a while. So it's not a real loss. Unfortunately right now these AI companies have full freedom to do whatever they want. Taking paid content, artistic works, and your own posts on social media. So Reddit trying to charge them is a good idea as it's some form of quid pro quo put on AI scraping companies.
Except that what Reddit is really doing is selling content they didn't produce and don't own. I don't think they're walking some kind of high road here like they would be if they were actually fighting against the scraping.
That being said I can still take some satisfaction in seeing AI scapers get jammed up considering how they face zero consequences right now.
See: https://www.reddit.com/r/internetarchive/comments/1gpn54q/is...
> They are not specifically targeting Wayback Machine. Anything other than residential IP's are blocked, to my information. Such as IP's of cloud services like Hetzner, GCP, AWS... The list goes on. (from my comment there)
A way to balance archiving (in case something happens to Reddit) and exploitation (of copyright/material).
Yeah, the front end for de-enshittification looks a lot like that other archive site,
In the summer of 2020 I was driving to Buffalo a lot with my son and getting cheap hotel deals thanks to the pandemic and thinking about missile defense systems and I was sick and tired of the awful shape of the web and dreaming up a system that would "archive" 100% of web pages before I read them. I spent two weeks on a spike prototype and concluded that an "archiver" can never really know if a modern web page is done loading so it at best uses heuristics to make the page load completely and waits a long time -- which makes following a link even slower than waiting for all the ads and trackers to load. I finally got Fiber-to-the-Node at home so downloading all the trash of the annoyances economy became more tolerable, a lot of the ideas I had that the time made it into my RSS reader a few years later.
See, I don't think this is right either. Back during the original API protests, several people (including me!) pointed out that if the concern was really that third-party apps weren't contributing back to Reddit (which was a fair point: Apollo never showed ads of any kind, neither Reddit's or their own) then a good solution would be to make using third-party apps require paying for Reddit Premium. Then they wouldn't have to audit all of the apps to ensure they were displaying ads correctly and would be able to collect revenue outside of the inherent limitations of advertising.
Theoretically, this should have been a straight win for Reddit, especially given the incredibly low income that they've apparently been getting from ads anyway (I can't find the report now so the numbers might not be exact, but I remember it being reported that Reddit was pulling in something like ~$0.60 per user per month versus Twitter's slightly better ~$8 per user per month and Meta's frankly mindblowing ~$50 per user per month) but it was immediately dismissed out of hand in favor of their way more complicated proposal that app developers audit their own usage and then pay Reddit back.
My initial thoughts were either that the Reddit API was so broken that they couldn't figure out how to properly implement the rate limits or payment gating needed for the other strategy (even now the API still doesn't have proper rate limits, they just commence legal action anyone they find abusing it rather than figure out how to lock them out; the best they can really do is the sort of basic IP bans they're using here), or the Reddit higher-ups were so frustrated that Apollo had worked out a profitable business model before them that they just wanted to deploy a strategy targeted specifically at punishing them.
But it quickly became clear later that Reddit genuinely wasn't even thinking about third-party apps. They saw dollar signs from the AI boom, and realized that Reddit was one of the largest and most accessibly corpuses of generally-high-quality text on a wide variety of topics, and AI companies were going to need that. Google showing an intense dependency on Reddit during the blackout didn't hurt either (yes, at this point I genuinely believe the blackout actually hurt more than it helped by giving Reddit further leverage to use on Google, hence why they were one of the first to sign a crawler deal afterwards).
So they decided to use any method they could think of to lock down access to the platform while keeping enough people around that the Reddit platform was still mostly decent enough to be usable for AI training and pivoted much of their business to selling data. All of this while claiming, as they're still doing today with the Internet Archive move, that this is somehow a "privacy measure" meant to ensure deleted comments aren't being archived anywhere.
The same thing basically happened with Stack Exchange, except they had much less leverage over their community because the entire site was previously CC licensed and they didn't have any real authority to override that beyond making data access really annoying.
The good news is that it really does seem like "injest everything" big model AI is the least likely to survive at this point. Between ChatGPT scaling things down massively to save on costs with the GPT-5 update and the Chinese models somehow making do with less data and slower chips by just using better engineering techniques, I highly doubt these economics around AI are going to last. The bad news is that, between stuff like this and the GitHub restructuring today, I don't thing Big Tech has any plans on how they're going to continue functioning in an economy that isn't entirely based on AI hype. And that's really concerning.