Fun fact: Germany's IP law has a provision that allows AI training by default for everything that's reachable on the Internet, if the website operator hasn't published a "nope" in machine-readable form (i.e. robots.txt).
If you read the law it says it’s only allowed if the rights holder doesn’t disallow it (in machine readable form) - I would argue robots.txt falls under machine readable
So to legally copy a website, all you need is to just pass it through an AI filter and then you can legally publish the rip off?
No - data mining copyrighted material and republishing copyrighted material under your own name are two very different things
Passing data through an AI filter for "training" is a different thing (legally and ethically) from publishing the output.