Chernoff face
en.wikipedia.org
en.wikipedia.org
The idea is that you take features of your dataset, and use those to represent a face. Say for example, you want to classify 100 people based on different features. And let's say you've collected 15 features for each person (e.g., height, weight, shoulder width, length of first name, length of last name, type of car driven, etc.). Now try mapping each of these features to Chernoff faces. You'd map it in the following manner: height->area of face, weight->shape of face, shoulder width->length of nose, length of first name->location of mouth, length of last name->curve of smile, type of car driven->width of mouth, etc.
Once you've mapped in that fashion and visualize the faces, you can observe how discriminative your features are. How do you interpret this? If your Chernoff faces tend to show a lot of variation in expression (e.g., smiling vs. sad), you say the length of last name is more discriminative. On the other hand, if the faces all appear to have same area, your first feature (i.e., height) is not very discriminative.
Other features used for Chernoff faces could be: location, separation, angle, shape, and width of eyes; location, and width of pupil; location, angle, and width of eyebrow, etc.
One drawback (as listed in the Wikipedia page) is that we humans perceive the importance of these faces by the way in which variables are mapped to the Chernoff facial features. If the feature mapping is not carefully chosen, your largest varying feature may be ignored because we appreciated the change in expression more than the change in eyebrow length.
I don't get why that is easier or more revealing than doing a principal component analysis (?)
If the primary utility is to be able to quickly visually discriminate values once you know how they're encoded into facial features, then I can see the value, but again you'd have to know the encoding. Or have I missed the point completely?
Checkout the Chernoff Fish demo posted below by the user meagher here: https://news.ycombinator.com/item?id=16664051. Play with different features, say for example, 'performance'. When you change the value of 'performance', the eye size changes. However, the eye size doesn't mean anything except for you to visually understand variations in data. If 'performance' was mapped to, say, fin size, it doesn't change its meaning.
Would people more readily recognize that say, "Large Spiky Orange Fish" strategies lead to greater returns, compared to if the strategies were presented as "Short | Value Investment | Large Market Capitalization"?
This could also be an interesting way of eliminating inherent bias while leveraging human pattern recognition abilities. Represent values pictorially, and hide data labels.
Here are the Lawyers' Ratings of State Judges in the US Superior Court:
https://www.rdocumentation.org/packages/datasets/versions/3....
And here is the R code:
library(TeachingDemos)
svg("Chernoff faces for evaluations of US judges.svg", height=6, width=8)
faces2(USJudgeRatings[1:12,c(2:11,2:9)]/5-1, scale="none")
dev.off()
Could someone who knows R and where to find the faces2() function give us a legend (i.e., a table of judge's property to facial feature)? Maybe we can add that into Wikipedia afterward too.> The features are: 1 Width of center 2 Top vs. Bottom width (height of split) 3 Height of Face 4 Width of top half of face 5 Width of bottom half of face 6 Length of Nose 7 Height of Mouth 8 Curvature of Mouth (abs < 9) 9 Width of Mouth 10 Height of Eyes 11 Distance between Eyes (.5-.9) 12 Angle of Eyes/Eyebrows 13 Circle/Ellipse of Eyes 14 Size of Eyes 15 Position Left/Right of Eyeballs/Eyebrows 16 Height of Eyebrows 17 Angle of Eyebrows 18 Width of Eyebrows
So there are 18 visually distinguishable features, and the code maps the columns 2–9 redundantly to two features each, and columns 10–11 to one feature each. I don’t understand how the scaling was chosen — the function requires the values to be in the interval [0, 1] but since the original scores appear to be in [0, 10], a more natural transformation would be to just divide by 10 instead of `x / 5 - 1`. Obviously these two transformations result in markedly different faces.
If I find time later I’ll upload an updated plot that includes a legend. Unfortunately that’s not easy since the `faces2` function overrides R’s plot layout so there’s no space to fit the legend.
[1] https://commons.wikimedia.org/wiki/File:Chernoff_faces_for_e...
All the code is on GitHub: https://github.com/tmm-archive/chernoff-fish
> Using a neural net... would be a useful way of generating custom Ross-Chernoff plots, and also allows me to include “deep learning” in the author keywords.
A vampire had hologram faces of tortured humans as graphs
A sea of tortured faces, rotating in slow orbits around my vampire commander.
"My God, what is this?"
"Statistics." Sarasti seemed focused on a flayed Asian child. "Rorschach's growth allometry over a two-week period."
"They're faces…"
He nodded, turning his attention to a woman with no eyes. "Skull diameter scales to total mass. Mandible length scales to EM transparency at one Angstrom. One hundred thirteen facial dimensions, each presenting a different variable. Principle-component combinations present as multifeature aspect ratios." He turned to face me, his naked gleaming eyes just slightly sidecast. "You'd be surprised how much gray matter is dedicated to the analysis of facial imagery. Shame to waste it on anything as—counterintuitive as residual plots or contingency tables."
I felt my jaw clenching. "And the expressions? What do they represent?"
"Software customizes output for user."
[1]: https://gnarmis.github.io/chernoff-faces/ -- a simple toy
For example, here are Chernoff Faces of some Turkish Universities: https://ibb.co/nFAoC7
The size of the eyes represents the number of projects submitted to TUBITAK(Govt science body that coordinates and provides funding), the size of the nose represents the number of projects that are accepted. The size of the face is the number of professors in the institution and the size of the mouth is the number of publications.
So, if you are looking for Turkish Universities that have lot's of accepted projects look for a big nose. Big eyes and small nose will be an indicator for a large number of failed applications and a face that has a huge mouth but small nose and eyes is an indicator for an institution with lot's of publishing going on without seeking funding to projects.
FWIW, as much as I liked the idea of Chernoff faces after running into them in _Blindsight_, I was never able to find a productive use for them where they were really all that much a better visualization than, say, a bunch of scatterplots or fancier techniques like t-SNE.
It's still available, and there are pirate versions if you want to try it before buying it. It's a great book, it's chock full of odd ways to visualise data.
Chernoff Faces https://imgur.com/gallery/ES3in