One problem with archive.org is that it retroactively obeys the robots.txt file, even if the files have been spidered and archived [0].
For example, consider the case when a new domain owner attempts to block all bots from spidering their site, by adding something like this to their robots.txt file
User-agent: *
Disallow: /
This is actually a fairly common case when domain resellers purchase expired domains.Now when you try to visit the archived link, because the live robots.txt file disallows bots, you won't be able to access the archived site (which may have been owned by someone completely different).
[0] https://archive.org/post/406632/why-does-the-wayback-machine...