1, the algorithm use something called haar filters, its simply box filters with a "White" and "black" area - for each region of the image, you take the sum of the pixels in the "White" area and minus the sum of the pixels in the "black" area. Thus the filters output a simple single number;
- First you generate a gazillion of these filters for different regions of the image with different shapes. such as a 2 horizontal rectangle or 2 vertical ones ( one black and one white ).
- You take the result of the output of the filter to find the threshold which differentiate the faces and non-faces in your training images the best.
- Using all the filters and outputs you have, feed it into an machine learning algorithm called "Adaboost", which attempts to minimize an exponential loss function of the error of classification. These different filters are then assembled with different weights.
- The final structure of the detector is a "detection chain", which is a degenerate decision tree (like a chain) with nodes consisting of the aforementioned filters assembled together. This is how the algorithm achieves its speed, by rejecting non faces early in the detection process. Only when a image region passes all the nodes in the detection chain that its labeld as a "hit";
- After that you scan all regions of the image at all sizes ( brute force ) and then assemble the detection results.
To be honest, the first time that I saw the description of the algorithm, it seemed a little "magical". The underlying reason why these "box filters" work so well is because the human face is well-defined by "boxy" feature such as our eyes, nose, eyebrows, lips..etc. Its a wonderful wonderful application of Machine learning to a specific domain.
This is also why this algorithm has MUCH lower detection rates for cascades trained to detect side view faces, because these "boxy features" that are so well detected by these filters are simply not as prominent in the side view
For more, refer to the Viola and Jones paper , of all the versions out there, I find this the best:
http://lear.inrialpes.fr/people/triggs/student/vj/viola-ijcv...
to be perfectly honest, this detector sucks.... Although its fairly illumination invariant, its not rotation invariant and sucks for side-view faces detection. I've been trying for the longest time to implement the histogram based detector outlined in this paper: http://www.cs.cmu.edu/afs/cs.cmu.edu/user/hws/www/CVPR00.ps
which is also what pittpatt uses for their detector and IMO its a much better detector. However the lack of training images and time has been impeding my progress. It'll be opensourced when I'm done.