>This creates an in-group with pristine datasets, and an outgroup whose models will collapse on the slop outputs with no good ability to filter.
But all the chinese labs who are hot on the heels of american labs thanks to "distillation" seems to be able to work without "pristine datasets"?