I'm almost certain they won't. IIRC, Facebook and Twitter are very resistant to scraping. Eventually they'll shut down or pivot, and all their existing data will go poof.
Plus people think about social media differently than newspapers. Newspapers were an open public record, and some effort was always made to archive them (e.g. the local library binding them into books or microfilming them). Social media is this weird amalgam of public and private, that people are more jealously guarding from the public.
1. https://blogs.loc.gov/loc/2017/12/update-on-the-twitter-arch...
They're relying very heavily on their existing network of users if that is not some weird thing that only affects me. I'd assume most readers don't create accounts.
even today we have issues determining if what sources from antiquity say is true.
as a general example, there wasn't much propaganda value in lying about food, so for the most part we can trust those descriptions to be true, particularly in guides intended to be written as cookbooks. in 2024, we have legion accounts making fake recipes for the sole purpose of getting clicks for ad dollars.
Myspace, the once mighty social network, has lost every single piece of content uploaded to its site before 2016, including millions of songs, photos and videos with no other home on the internet.
The company is blaming a faulty server migration for the mass deletion, which appears to have happened more than a year ago, when the first reports appeared of users unable to access older content. The company has confirmed to online archivists that music has been lost permanently, dashing hopes that a backup could be used to permanently protect the collection for future generations.
...
Internet Archive generally allow websites to control if they get indexed/mirrored or not, via the robots.txt, so websites can decide for themselves.
Luckily, we have other grassroots movements like ArchiveTeam that doesn't care and archives anything deemed valuable to be archived, website owners be damned.
<https://blog.archive.org/2017/04/17/robots-txt-meant-for-sea...>
<https://blog.archive.org/2016/12/17/robots-txt-gov-mil-websi...>
No? This is wrong? The site owner is the person who morally decides whether their site should get mirrored or not.
It's stuff like this that makes me actively cheer forward otherwise-harmful things like Google's web attestation efforts.
Alternatively - you make an excellent case for the paywalling of vast swaths of the internet.
No, as good as it is, the Internet Archive in a single point of failure. Which was put on stark display when they decided a few years ago to pick a legal fight over copyright that they could never win and that put their organization at risk.
Also, I've tried to use the Internet Archive to grab Facebook posts. It doesn't work, even for public ones (all I got was pages and pages of the Facebook login screen).
Since when were newspapers and magazines considered to be accidental historical sources? They were part of the war effort. It was literally state war propaganda in an ongoing war.