Yale researchers reconstruct facial images locked in a viewer’s mind
news.yale.edu
news.yale.edu
The most sure and certain finding of any preliminary study will be that more research is needed. Disappointingly often, preliminary findings don't lead to further useful discoveries in science, because the preliminary findings are flawed. If the technique reported here can generalize at sufficiently low expense, it could lead to a lot of insight into the workings of the known-to-be complicated neural networks of the human brain used for recognizing faces.
A useful follow-up link for any discussion of a report on a research result like the one kindly submitted here is the article "Warning Signs in Experimental Design and Interpretation"[3] by Peter Norvig, director of research at Google, on how to interpret scientific research. Check each news story you read for how many of the important issues in interpreting research are NOT discussed in the story.
[1] http://www.phdcomics.com/comics.php?f=1174
[2] http://www.sciencebasedmedicine.org/index.php/related-by-coi...
I would like to see drastic changes to the journal publishing model, but I don't consider article inaccessibility to be correlated with flaws in the findings. To bring up the latter out of frustration over the former is poisoning the well of debate.
Incidentally, I found the paper in under 30 seconds by searching for the journal name, the senior author, and a few technical terms, all of which were in the press release and which gave me a correct first search result. It would be nice if the press release also contained a DOI link and other identifying information, but I can't really blame press departments for supplying journalists with the information they want and omitting that which most of them don't.
The left column is the original image. Next to it is a non-neural PCA/Eigenface reconstruction. Next to that are the reconstructions from a variety of brain regions.
He seems to have a pretty good track record of posting his papers there. This one isn't up yet (for one thing it isn't really fully published yet). Maybe you can check back later...
I know, it's super shitty. I have no idea what I'm going to do when I leave school.
http://www.youtube.com/watch?v=nsjDnYxJ0bo
Their particular process is described in the YouTube caption:
The left clip is a segment of a Hollywood movie trailer that the subject viewed while in the magnet. The right clip shows the reconstruction of this segment from brain activity measured using fMRI. The procedure is as follows:
[1] Record brain activity while the subject watches several hours of movie trailers.
[2] Build dictionaries (i.e., regression models) that translate between the shapes, edges and motion in the movies and measured brain activity. A separate dictionary is constructed for each of several thousand points at which brain activity was measured. (For experts: The real advance of this study was the construction of a movie-to-brain activity encoding model that accurately predicts brain activity evoked by arbitrary novel movies.)
[3] Record brain activity to a new set of movie trailers that will be used to test the quality of the dictionaries and reconstructions.
[4] Build a random library of ~18,000,000 seconds (5000 hours) of video downloaded at random from YouTube. (Note these videos have no overlap with the movies that subjects saw in the magnet). Put each of these clips through the dictionaries to generate predictions of brain activity. Select the 100 clips whose predicted activity is most similar to the observed brain activity. Average these clips together. This is the reconstruction.
With the actual paper here:
http://www.cell.com/current-biology/retrieve/pii/S0960982211...
On the other hand, this is presented as in the press release as mind reading, but the reality is more like trying design something similar to a cochlear implant.
Mind to explain to me why the right clip is inconsistent with the left clip?
https://www.youtube.com/watch?v=nsjDnYxJ0bo
Starting at 20-second I feel like the right clip is out of touch with the left clip. between 20th and 22nd second I see at least three individuals rendered from the reconstruction.
From 26th to the end of the clip I also see multiple individuals. The names also look different from one another... When you say find the closest is that an expected result?
Based on my perspective, there were three sets of videos:
1) The several hours of "training" video, that they used to learn how the test subject's brain acted based on different stimuli. (The paper (which I've only skimmed) says 7,200 seconds, which is two hours)
2) 18,000,000 individual seconds of YouTube video that the test subject has never seen.
3) The test video, aka the video on the left.
So, the first step was to have the subject watch several hours of video (1), and watch how their brain responded.
Then, using this data, they predicted a model of how they thought the brain would respond for eighteen million separate one second clips sampled randomly from YouTube (2). They didn't see these, but they were only predictions.
As an interesting test of this model, they decided to show the test subject a new set of videos that was not contained in (1) or (2), the video you see in the link above, (3). They read the brain information from this viewing, then compared each one second clip of brain data to the predicted data in their database from (2).
So, they took the first one second of the brain data, derived from looking at Steve Martin from (3), then sorted the entire database from (2) by how similar the (predicted) brain patterns were to that generated by looking at Steve Martin.
They then took the top 100 of these 18M one second clips and mixed them together right on top of each other to make the general shape of what the person was seeing. Because this exact image of Steve Martin was nowhere in their database, this is their way to make an approximation of the image (as another example, maybe (2) didn't have any elephant footage, but mix 100 videos of vaguely elephant shaped things together and you can get close). They then did this for every second long clip. This is why the figure jumps around a bit and transforms into different people from seconds 20 to 22. For each of these individual seconds, it is exploring eighteen million second-long video clips, mixing together the top 100 most similar, then showing you that second long clip.
Since each of these seconds has its "predicted video" predicted independently just from the test subject's brain data, the video is not exact, and the figures created don't necessarily 100% resemble each other. However, the figures are in the correct area of the screen, and definitely seem to have a human quality to them, which means that their technique for classifying the videos in (2) is much better than random, since they are able to generate approximations of novel video by only analyzing brain signal.
Sorry, that was longer than I expected. :)
Edit: Also, if you see the paper, Figure 4 has a picture of how they reconstructed some of the frames (including the one from 20-22 seconds), by showing you screenshots whence the composite was generated.
I suspect the algorithm always outputs some face generated by a paramaterized face model (neutral net based?). Therefore even random output would generate a face. Then with some "tuning" and a little wishful thinking you might convince yourself this works.
Am I being too skeptical?
"Working with funding from the Yale Provost’s office, Cowen and post doctoral researcher Brice Kuhl, now an assistant professor at New York University, showed six subjects 300 different “training” faces while undergoing fMRI scans. They used the data to create a sort of statistical library of how those brains responded to individual faces. They then showed the six subjects new sets of faces while they were undergoing scans. Taking that fMRI data alone, researchers used their statistical library to reconstruct the faces their subjects were viewing."
So yes, it will always output something like a face. It's more like they are using the FMRI to select among preset options. It's still potentially a great result, but we need more detail than this article provides.
Generating a result which looks like some combination of the inputs, has a little bit of 'Wow' factor if you just look at the pictures and don't think too hard about it. But ultimately its not an exercise in which they can be wrong, and since the image returned is always going to be a face, they'll always kind of be right.
An impressive result would be if they trained their system on the the 300 pictures, and then we're reliably able identify which picture the subject was looking at. That would be quantifiable and testable, and I presume the result would expose that this is all nonsense.
So, non-falsifiable...?
Note the cherry picking of the output for the bottom picture in the Yale press release (forth row in this image).
The paper includes a quantitative evaluation, in which they take a set of 30 distractor face images not used elsewhere in the study, and for each of them determine whether the reconstructed face is more similar to it, or to the original face the subject was looking at. On average, the reconstructed face is closer to the correct face than the distractor 62.5% of the time.
So it's better than random, and I think it's pretty cool work, but the quality of the reconstructions is pretty terrible, considering that a randomly chosen distractor will usually be very different from the test face (~half the time opposite sex, frequently different race, different age, etc). For comparison, it would be interesting to evaluate some simple, obviously terrible reconstructions by the same metric. For example, we could "reconstruct" the face as a image that is a single, solid color, the average RGB value of the pixels in the original face. Another "reconstruction" that it seems would very likely do better under this evaluation metric is something like the "race-gender-age" of perp descriptions in the news ("white male in his thirties").
So this raises many questions: How diverse were the faces in the training corpus? How close were the new images to those in the corpus? When you're looking at hundreds of images to train the machine, are you also unknowingly being trained to think about images in a certain way? What happens when you try to recreate faces based on the fMRI responses of subjects who didn't contribute to the initial training set?
The implications of the last question are pretty interesting . If different people have different brain responses to looking at the same image, does that help us begin to understand why you and I can be attracted to different types of people? Does it help begin to explain why two people can experience the same event but walk away with two completely different interpretations?
Hypothesis confirmed!
I thought the woman looked close enough to be able to identify, but the man was not. Still, very impressive work.
They're reconstructing faces the subject is viewing, not remembering. To replace police sketch artists, the criminal would have to be present...
Actually, there's an obvious scifi plot: investigator is a serial killer but excluded from memory extraction because the witness just saw them (again).
In all seriousness this is a great advance in neuroscience that would help understand many things about brain. On the other hand, potential for misuse is enormous. Can you even prove you had been interrogated if such a device is used on you?