OpenAI and Anthropic are ignoring robots.txt
businessinsider.com
businessinsider.com
https://newmedialaw.proskauer.com/2022/05/24/doj-revises-pol...
As a practical matter, if web site owners don't like particular HTTP requests then they can just ignore them or return errors or junk responses.
The DoJ is explicit in saying that something like a Cease and Desist is enough, so if for example the NYTs found OpenAI's bot then that would likely be prosecutable.