Show HN: An AI program to check videos for NSFW content
github.com
github.com
It will over-report on video of baby's first bath, beach volleyball, and high school wrestling; it will under-report on many of the kinkier or more subtle forms of porn.
Humans get this stuff wrong all the time simply by having different cultural contexts; there's no way the 2022 state of the art is up to the challenge.
That's part of the issue.
That doesn't always stop corps productizing it, and making it their official NSFW filter. Good enough is good enough, just so long as they can pay less salaries.
In the blog, I suggest this could be the result of an uncultured data set when training the CNN. Or perhaps the dataset was fine, and this is pushing the hard limit of what ResNet50 CNN architecture can do (the off the shelf model I use for this is an extension of ResNet50).
Some of the anomalous results are amusing. One day, I uploaded a video of a female violinist in concert, and the model flagged every close up of her as NSFW! Just those closeups. Wide shots, and closeups of other musicians were absolutely fine.
Again some of that might be down to me (clunky code / very low NSFW threshold). And I suspect the model I used was itself a PoC (https://github.com/emiliantolo/pytorch_nsfw_model). But it does make you wonder how the bigger labs with critical products, like Palantir, handle doubts like this.
Watching the olympics at lunchtime is commonplace and accepted in most workplaces. Being a creep about it is not. It's almost always a pretty clear line between the two.
It only becomes sexualized by the behavior of the people watching - professional athletes competing for their country isn't something intrinsically inappropriate.
If someone is sexualizing something in the workplace, the problem is the person's behavior, not the thing they are sexualizing.
The situation we find ourselves in came from setting boundaries around content instead of behavior. Our society treats nudity itself as if it were the act of sexualization, assuming that people have no control over their own behavior.
In this case, I think it's best just to hold companies accountable for discrimination that isn't based on some protected attribute. Companies lean to heavily on "right to refuse service (for any reason)" and "we can choose who to do business with".
The fact we distinguish between "things you are allowed to discriminate on" and "thins you aren't" makes for a very convoluted process. Same with at-will employment tbh. We should just flip it from a "you need a reason to thing you were unjustly discriminated against" to "you need a good reason to fire someone".
But when it comes to online video checking, it's another tool in the arsenal; it's the first lines of defense. First check: does the signature match a previously known and confirmed video marked as porn or otherwise unacceptable. Second check: does the more fuzzy AI thing consider it porn with a high probability. Third check, which will be things the AI mark as 'not sure' or incorrectly marks as 'porn', and things humans flag up manually, will be humans checking and judging things.
And of course it'll be flexible, because the definition of porn - or what is 'not acceptable' - will vary by country and culture, and these companies want to be active everywhere.
except Europe, unless they are given tax breaks and data, amirite
- Babies (https://github.com/wingman-jr-addon/wingman_jr/issues/22)
- Beach volleyball (but this definitely has SFW and NSFW variants, based on a somewhat subjective line)
- Athletes in general. The model particularly thought some American football players were NSFW for a long time.
- Swimming
- Yoga - again, most SFW and some NSFW here but it still struggles
- Wrestling was a tough one for sure
- Pokemon
While indeed tough, I've seen definite progress. So it's not just a matter of tech, but also of considering the human element - the state of the art may not be up to the challenge of perfection, but it is definitely up to a point of true utility for some use cases. I'm happy about that.As a note, it uses an EfficientNet Lite L0 backbone - I'm a bit limited in what type of scanning I can perform in a sufficiently speedy manner.
I also agree on the context for sure - one reason I haven't tried switching to an object detection method (and that I don't rely heavily on truly random crops) is that the focus of the image is highly important for the NSFW-ness in some cases. True, two images may contain the same content ... but one is far worse than the other. The nature of CNN's still has some of this location-invariance baked in, but I don't want to exacerbate it.
One challenge I think the OP may run into here that may also not be immediately obvious is that accuracy on image stills does not translate that well to video. I have basic video support in my addon, and while I knew there would be some differences, I was surprised at how many discrepancies there really are. As two examples:
- Images in video are often blurrier. In true still images, there is a somewhat higher prior involved with amateur NSFW content and blurriness. This can be a source of false positives.
- The opposite of the note above about focus. Taking stills of moving images will have many transitory frames that seem inappropriate on their own because it seems as if they are focusing on something when in reality the camera is just panning - obvious to the human, less so to the model trained on stills.
At any rate, given how well your list of edge cases coincided with failures I've grappled with, I'd be interested to see how well you think my addon stacks up for still images when set to stay in "normal" mode. I'd love to hear any feedback you have via GitHub so I can make it better.One of the technical issues that you pointed out, is that a model trained on still images, shouldn't be expected to work on video. While I did not train a custom model for this this project, I'm current working on another DNN model for a completely different purpose, where I think feeding frame deltas into the model, will improve the outcome.
As a hobbyist, I would reckon for porn and the like, analysing frozen frames is probably just enough. For violence however, I would agree with you and say that some effort to encode motion would be essential.
Focussing on NSFW content generally, I would guess, depending on the scale of your project, that you will forever run into 'edge cases' for NSFW images, even before you run into the soft wall of subjectivity.
I agree that the tech is improving all the time, and I think something like this can be made to be truly useful one day. Possibly soon. But it would need a large, active development team, a great deal more compute, and a LOT of data. In much the same way that no home/garage coder can hope to put together a model like GPT3 right now, I would think that a foolproof NSFW classifier would need more resources than you or I have access to at this moment.
But things change all the time.
Thinking about what you're doing, one thing I might suggest, if you have time to develop it, is to add some kind of 'recording' mechanism to your plugin, so that the users themselves can add to your dataset... But you have to wonder how many users will allow that! XD
I'm also wondering if a Firefox extension is the best place for your model? To that end, I would suggest putting the app on a server (which is what I originally wanted to do with my hack) which will give you the opportunity to crowdsource data collection. People might be more willing to volunteer data in that way (in a similar way to how people use https://builtwith.com/).
You're also very welcome to take the UX work I've done on this opensource project (because this hack was ultimately just a UX experiment), and plug your model in. If your model and trained weights are available, I'd like to try and create a branch myself, if I have time.
Also, as hobbyist building knowledge, I hadn't heard of `EfficientNet Lite` before. I'd been considering Darknet - https://pjreddie.com/darknet/ for embedded stuff until reading your post.
Regarding the current state of tech: I agree the tech still has quite a ways to go. I think one of the most interesting aspects here is how e.g. NSFW.js can get extremely high accuracy - but not necessarily perform better in the real world. I think it speaks in part to the nature of how CNN's work, the nature of the data, and the difficulty of the problem. Still, having seen how incredibly good "AI" has gotten in the last decade, I have quite a bit of hope here.
Regarding putting it on a server: that is indeed a fair question, but my desire is to keep the scanning on the client side for the user. In fact, it was actually the confluence of Firefox's webRequest response filtering (which is why I didn't make a Chrome version) and Tensorflow.js that allowed me to move from dream to reality as I had been waiting prior to that time. I can't afford server infrastructure if the user base grows, and people don't want to route all their pictures to me. So I guess I see the current way it works as a bonus, not a flaw - but it DOES impact performance, certainly.
Regarding data collection with respect to server - yes, this is something I've contemplated (there's a GitHub issue if you're curious). There are, however, two things that I've long mulled over: privacy and dark psychological patterns. Let me explain a bit. On the privacy front - it is not likely legal for a user to share the image data directly due to copyright, so they need to share by URL. This can have many issues when considering e.g. authenticated services, but one big one also is that the URL may have relatively sensitive user-identifying information buried in its path. I can try to be careful here but this absolutely precludes sharing this type of URL data as an open dataset. On the psychological dark patterns front - while I'm fine with folks wanting to submit false positives, I think there's a very real chance some will want to go flag all the images they can find that are false negatives (e.g. porn). I don't think that type of submission is particularly good for their mental health or mine. So, in general, I think user image feedback is something that would be quite powerful but needs a lot of care in how it would be approached.
Regarding the UX - thanks! And you're welcome to try the model as well - I've tried to include enough detail and data to allow others to integrate as they wish: https://github.com/wingman-jr-addon/model/tree/master/sqrxr_... Also, let us know how things go if you try out Darknet.
Good luck!
Seems like we already acknowledge good enough practices and are okay with them. Lets not demand perfection in the same way we shouldn't assume perfection of ~~AI~~ ML.
If this was for a product, the first thing I would have done, was drop half the tools in the stack for their inefficiency! For example, the GUI shell is using Node Webkit (which I prefer to Electron, but is essentially the same thing). That in itself is quite bad, but it's not the worst approach to building a desktop app, as Microsoft and Slack have already proven.
But the .exe's you mentioned are also quite large.
The code is very transparent though. I doubt you'd have been the only one to note how clunky this all is!
> Please link to <https://raskie.com>, if you intend to redistribute this .exe.
Now that's probably okay so long as it's a nice request ("please"), and not a requirement; AIUI, GPLv3 doesn't (generally) allow you to add conditions (ex. a license file that says "GPLv3 but no commercial users" isn't allowed by the text of GPLv3). But IANAL and I'm sure this is more complex than I think it is. Which brings us to a reasonable useful point: I would suggest never putting anything in a license file other than the exact standard form of the license; adding anything else, at best, makes people have to read it and try to understand what you've done rather than being able to say "oh, it's GPLv3 and I've cached that that's 100% okay for me" and stop thinking about it.
The license file needs to be in the root of the repo.
The first thing many developers who have been bitten by this in the past do is run all the boilerplate files like licenses through diff. Some probably have it hooked to every git clone request.
> it determines the Baywatch intro to be somewhat NSFW at a glance
As nothing in the Baywatch intro is NSFW (https://www.youtube.com/watch?v=O0nqwgu_Us4), it seems the model is failing at large here.
Interesting subject and lots of praise to be had if the model can be made accurate, but seems it's not really there yet. I wish you luck in getting there!
a debatable claim. There are several scenes in the intro I would not want to have paused fullscreen in a work environment.
Just doing so would suggest you aren't actually working (on its own a problem in many cultures).
In an office setting the background noise can be highly disruptive, for example.
Which is also a debatable because I don't think there's much resarch indicating that images of naked people are in any way harmful for children.
Maybe NSFPP? Not safe for parents and puritans?
I would suppose that depends on the workplace and area(and time) where you live. It is definitely "somewhat" sexualised content in my opinion (close up of a womens body stripping down), even though there is no nipple to see. But sure, lots of normal advertisement sells more sex. So is the red line, when a nipple is shown? That would classify biological and medical content as nsfw.
My point is, human morals are very different. Humans fight about, what content is NSFW. A computer model can never be "acurate" in that sense.
This reminds me of a question I've long had about this sort of thing you reminded me of - in the US at least, we seem to draw the line for showing off breasts at the nipple, but only female nipples. Setting aside any merit or logic to that, can AI actually determine female vs male nipples? I assume not through the nipple itself, but the breast it's attached to or the face of the person?
Not with 100% accuracy, but on average yes, as a female breast is indeed quite different from a male breast. Shape, hair, size, ...
But there are of course exeptions. Muscular women, with small breast who did not breast feed yet, are probably quite hard to detect.
Yes, through context. A bearded woman holding a chainsaw would probably fail the test.
Clearly incorrect by any reasonable definition unless you work at a swimwear company or a few bits of the media.
I mean, you or I might think Alexandra Paul in a one-piece is an image approaching the beauty of classical sculpture, but quite a lot of classical sculpture is NSFW in most of the world, and the brow is a bit lower where Baywatch is concerned.
Very many workplaces have a definition of NSFW that is a bit more restrictive than Facebook's content rules.
And rightly so. Go to work, do your work, go home. It's not adult daycare.
Watching video footage of people running around in swimwear is not really appropriate content unless you work in media or swimwear.
And the title sequence to Baywatch is hardly an advert for sporting endeavour; Baywatch is titillation. It would be entirely reasonable for almost anyone, male or female, to complain about that content.
It is far more sensible to define NSFW as "things I shouldn't be looking at just in case someone has an unjustifiable complaint to make about me that they are looking to bolster with additional offences".
Per the Baywatch example - there's clearly a grey area here, and in my opinion the model is working as intended. False positives for NSFW are better than false negatives.
Or do you mean if you're looking for NSFW scenes of a particular actor?
The solution can also be deployed on-promises for real-time, local video analysis without leaving the deployment environment: https://pixlab.io/on-premises.
Surprised there is no free tier here but I’m not experienced in this space.
The fact that our modern society thinks nudity and sex is bad but violence, torture and death are OK I just don't understand.
I would much rather have my kid see the Pamela Anderson sex tape than someone being murdered or even water boarded.