I guess CSAM detection will remain hard as the training data raises so many ethical concerns.
I wonder if there's a way in which we can forward-hash sensitive material like CSAM and train a classifier that only consumes the 'hashed' version of the material for training and detection.