Content Based Intelligent Cropping
engineering.curalate.com
engineering.curalate.com
There are a lot of things that if you get wrong, you'll ruin the image. This is true from an artistic perspective, but even with a lower bar, it's very easy to crop an image to make the content distracting or disturbing.
There are rules of thumb, for example, about how to crop images with people. It's not just that you want to preserve the faces. You need to be careful about where on their bodies you crop: if you do it right at a joint, it will tend to look like the person is an amputee. So you've got to crop a leg at mid-thigh rather than at the knee, for example.
There's also an idea of "look space". If you crop so that the direction of the subject's gaze is toward the outside of the image (rather than toward the middle), the subject will tend to look crowded and the overall composition unbalanced.
And there are overall compositional rules, like rule-of-thirds, leading lines, etc. Cropping an image is likely to damage these, thus harming the aesthetics of the picture.
I won't make the claim that this stuff is impossible to automate (the very fact that I can list these rules of thumb suggests that much of it can handled without creativity as such), but doing so in a way that preserves the desirable qualities of a good image will require a LOT more sophistication than what's being brought to bear here.
Some similar previous work that suggests to me this may be possible.
https://research.googleblog.com/2015/10/improving-youtube-vi...
There's face/features detection (and other CV) in here, too, but you can get by with some really cheap tricks in a lot of cases. In fact, the meat is it finds things that are chromatically different or more, uh, feature-full.
https://github.com/nkozyra/smartcrop
(again, don't be fooled by the name. at some point I'd hoped it would be smart. it was a class project.)
Plastic fumes :(
> But it doesn’t have to be this way! In this post, we present a technique that we use for intelligent cropping: a fully automatic method that preserves the image’s content
This seems like overengineering. Why not just build a UI to quickly crop a set of images exactly as desired, with a scalable box given the required aspect ratio?
Since it's likely you'll want to verify these kinds of images before posting them anyways, and might go back and fix some later, it seems like you're not really cutting down on anything other than maybe a few seconds per image to move that initial bounding box.
Though it's cool nonetheless.
My own use-case for this, though, involves ~500k images (think company logos appearing in profiles like LinkedIn, fit to various layouts.) I certainly couldn't do that (or even verify that) by hand.
http://www.fmwconcepts.com/imagemagick/multicrop/index.php (scroll down to see examples)
Pretty neat field and projects. I think we haven't really seen machine vision reach ubiquity / full deployment yet.
You should use them as centers of gravity, measure normalized size, compare against other features, etc.
In the past, cropping to remove faces was done for costs reasons. The retailers paid the model for printed use only and had to pay extra to use them on ecommerce websites.
We ran A/B tests, paid the models rights for ecommerce usage and demonstrated a steep increase in conversion.
The rights issue is the reason some companies provide "paste a face on pictures" automated services. Basically you buy the rights on a few headshots and the service stitches the faces & clothing together (see OpenCV stich functions...).
The problem is that while decent, the services often end up in the uncanny valley, actually decreasing conversion.