Internet Archive generally allow websites to control if they get indexed/mirrored or not, via the robots.txt, so websites can decide for themselves.
Luckily, we have other grassroots movements like ArchiveTeam that doesn't care and archives anything deemed valuable to be archived, website owners be damned.
<https://blog.archive.org/2017/04/17/robots-txt-meant-for-sea...>
<https://blog.archive.org/2016/12/17/robots-txt-gov-mil-websi...>
No? This is wrong? The site owner is the person who morally decides whether their site should get mirrored or not.
It's stuff like this that makes me actively cheer forward otherwise-harmful things like Google's web attestation efforts.
Alternatively - you make an excellent case for the paywalling of vast swaths of the internet.
No, as good as it is, the Internet Archive in a single point of failure. Which was put on stark display when they decided a few years ago to pick a legal fight over copyright that they could never win and that put their organization at risk.
Also, I've tried to use the Internet Archive to grab Facebook posts. It doesn't work, even for public ones (all I got was pages and pages of the Facebook login screen).