If the training data contained sexual images of consenting adults and legal images of children, many intelligent models could interpolate between the two.
To be clear, I find child sexual abuse appalling, but maybe synthetic images would keep some people "satisfied" and leave the real children alone?
Besides, if it worked like that, training on anything under copyright would have a similar effect. These models have a bigger problem if they get “tainted” by a tainted training source. (Fingers crossed they do! But I doubt it)
But "CSAM" also includes "character looks under 18" even if they look fully developed, which an AI model could do without training on actual CSAM.
Whether that makes the model illegal is a separate question. Is a LLM illegal if a clever prompt can get it to output a copyrighted poem? Is a human illegal if they can draw "CSAM"?
The inability to extract a training set from a model adds a lot of ethical ambiguity to generative images, as does the ambiguity in who’s responsible for what the models produce. I think it would be utterly ridiculous to say that someone training models with CSAM has no culpability in what it produces— only for possessing the CSAM to begin with— but I also think it would be utterly ridiculous to hold people accountable for everything their models produce, given their flexibility. What about writing a prompt that generates CSAM inadvertently? What if nobody involved intended to make CSAM but through some algorithmic shenanigans the prompt produced it? Should we legally require some amount of model testing before it’s used? Would the tester be violating the law if it failed the test, even if they reviewed every single image in the training set? Who’s responsible for deliberately poisoned models with secret key terms, or malicious data that is not CSAM but can trick the model into creating it? Would some entity like Midjourney that not only provides a model, but a complete appliance for this process be responsible for the images it produced? Does it matter if they authored the models they use? What if users can upload or train their own models? How does automation ethically affect these considerations?
Someone with legal expertise obviously would have a better grasp of these situations than I do, but I do know we’ve got a lot of growing pains en route with this technology.
As far as I know it's still OK to make images of murder and torture.