Face detection in pure PHP (without OpenCV)
svay.com
svay.com
1, the algorithm use something called haar filters, its simply box filters with a "White" and "black" area - for each region of the image, you take the sum of the pixels in the "White" area and minus the sum of the pixels in the "black" area. Thus the filters output a simple single number;
- First you generate a gazillion of these filters for different regions of the image with different shapes. such as a 2 horizontal rectangle or 2 vertical ones ( one black and one white ).
- You take the result of the output of the filter to find the threshold which differentiate the faces and non-faces in your training images the best.
- Using all the filters and outputs you have, feed it into an machine learning algorithm called "Adaboost", which attempts to minimize an exponential loss function of the error of classification. These different filters are then assembled with different weights.
- The final structure of the detector is a "detection chain", which is a degenerate decision tree (like a chain) with nodes consisting of the aforementioned filters assembled together. This is how the algorithm achieves its speed, by rejecting non faces early in the detection process. Only when a image region passes all the nodes in the detection chain that its labeld as a "hit";
- After that you scan all regions of the image at all sizes ( brute force ) and then assemble the detection results.
To be honest, the first time that I saw the description of the algorithm, it seemed a little "magical". The underlying reason why these "box filters" work so well is because the human face is well-defined by "boxy" feature such as our eyes, nose, eyebrows, lips..etc. Its a wonderful wonderful application of Machine learning to a specific domain.
This is also why this algorithm has MUCH lower detection rates for cascades trained to detect side view faces, because these "boxy features" that are so well detected by these filters are simply not as prominent in the side view
For more, refer to the Viola and Jones paper , of all the versions out there, I find this the best:
http://lear.inrialpes.fr/people/triggs/student/vj/viola-ijcv...
to be perfectly honest, this detector sucks.... Although its fairly illumination invariant, its not rotation invariant and sucks for side-view faces detection. I've been trying for the longest time to implement the histogram based detector outlined in this paper: http://www.cs.cmu.edu/afs/cs.cmu.edu/user/hws/www/CVPR00.ps
which is also what pittpatt uses for their detector and IMO its a much better detector. However the lack of training images and time has been impeding my progress. It'll be opensourced when I'm done.
I think I still have access to some data if you need it, but I agree there isn't much to be had. If only Facebook would open up their API for us to snag faces from user's photos using the tagging/notes system.
edit: it is more 10/10 since one that kind of recognizes a face is on a multi face image and it kind of frames a torso with head. I've tried your JS version and it works on single face and multi face images to recognize one image, however multi face recognition seems to be broken - blue bar comes to the end and that's it, nothing happens ever.
1) The detector doesn't like faces that are slightly rotated, either in-plane or out of plane.
2) The detector is "racist" in the way that most of the training images used were Caucasian people and therefore it works best on Caucasian.
3) It fails badly on kids because they aren't present in training data.
Try normalizing your data a bit for that and see how it works out. If you want a detector that works better , simply use the one in openCV since it does a more thorough detection. The performance optimizations I did also cuts down the accuracy of the algorithm ( assuming his port has those too ).
1. Run the adaptive boosting code which creates a complex decision function using known positive and negative data.
The code creates what is called an Integral Image from each piece of data, and runs it through the algorithm.
Each iteration adds or subtracts a 'feature' from the decision function. So in the end we build up a general detector for faces. The algorithm only keeps the features which work best together not individually, so in the end you have the best 'team' of features.
Essentially we are training a function so it can decide what is and is not a face. This is likely what the .dat file output was, a descriptor of the resulting decision function.
The more data used to train, the better the decision function will be.
2. Run the decision function against unknown data.
Ta da, an answer. It could be right, it could be wrong. Sometimes these algorithms find faces in the darnedest places like walls and other textured surfaces.
For more info on Viola-Jones adaptive boosting check out the original paper:
Robust Real-time Object Detection http://research.microsoft.com/en-us/um/people/viola/pubs/det...
Looks like I might be porting parts of this code into processing!