I think it comes down to use. Web crawlers like Google are fine because they index the web and then the search engine directs users to the original source. If instead Google recycled all the content they crawled and hosted everything on google.com while scrubbing all attributions from the pages then they’d fall afoul of copyright law (specifically the moral rights [1]).