Detection of watermarking requires access to the watermarking key, a secret in the current suggested scheme (leaking it would amount to being able to strip the watermark).
So, there will need to be a watermark checking service. The checking service will of course be rate-limited for common folk (and model distillers). OpenAI/Anthropic/Google/other privileged model builders need to filter out AI slop at scale, so need access to others' service without rate-limits (or the watermarking keys need to be shared).
This creates an in-group with pristine datasets, and an outgroup whose models will collapse on the slop outputs with no good ability to filter.
But all the chinese labs who are hot on the heels of american labs thanks to "distillation" seems to be able to work without "pristine datasets"?
No it's EU law.
This is simply false. You are underestimating the utility of synthetic data and the ability to learn from the mistakes the current weights make.