Robots.txt is usually explicitly permission for a robot
not to crawl. But a robot crawling your site and an archive, cache, or duplicate page are all different propositions.
Google and others have enhanced robots.txt to enable permission for crawling (allow, sitemap), meta tags can deny archiving and various means allow permission to be explicitly denied for caching.
To use your analogy of raising a sign: if you don't put up a 'no trespassing' sign then it doesn't make trespassing legal.
FWIW I disprove of this state of affairs and consider copyright to be hugely defective in these respects.
>but they clearly can keep a copy //
It's nuanced but permission to access a page =/= permission to keep a copy. Just as you have explicit permission to access a video on YouTube but in most jurisdictions will not have permission to download it for later (commercial) use.