For the former, I guess when training the original model, a bunch of the Reddit images weren't available at crawl time. Wouldn't it make sense to somehow weed those out from the data set before the training?
For the former, I guess when training the original model, a bunch of the Reddit images weren't available at crawl time. Wouldn't it make sense to somehow weed those out from the data set before the training?
I also didn't notice this issue until after I had trained the model and tried inverting hashes (didn't want to look through the NSFW dataset myself). And I thought it was amusing to leave it as-is (it demonstrates a point that issues from the dataset can be carried over to the results).
For the "image not available" ones, removing those in an automated way probably would be straightforward (could have used perceptual hashing for this task!).
For removing watermarks, there's some neat work from Dekel et al. in CVPR 2017: https://watermark-cvpr17.github.io/