What has it got to do with deduplication? I'm talking about crafting some kind of alternative (not necessarily duplicate) data. I agree some kind of post data collection cleaning/filtering of the data before training could potentially catch it. But maybe not!
Ah fair enough. The OP here mentioned having highly similar content on each of the many domains.
The funny way to do this would be to use an LLM to generate the content you respond with. Have 2 smallish LLMs talk to each other about topics chosen at random and generate infinite nonsense pages that have a few hundred words.