> it makes preparing training material for Machine Learning significantly easier
How?
How?
Right now, the Internet training set is becoming more and more contaminated with better and better generative AI images and video.
It makes the models more screwed up, and makes it very difficult for humans to figure out what is original and not too.
If there was some signal that could at least make it easier to identify ‘original’/real images…
The original model collapse paper assumes you train networks on 100% synthetic data produced by the previous generation. But if you maintain some portion of real data then the problem is mitigated.
I remember the original paper showing issues with even a couple percent of certain kinds of synthetic data too, not 100%.