UC Berkeley Scientists Translate Brainwaves into Imagery
abcnews.go.com
abcnews.go.com
The description of the first video: "The left clip is a segment of the movie that the subject viewed while in the magnet. The right clip shows the reconstruction of this movie from brain activity measured using fMRI. The reconstruction was obtained using only each subject's brain activity and a library of 18 million seconds of random YouTube video that did not include the movies used as stimuli. Brain activity was sampled every one second, and each one-second section of the viewed movie was reconstructed separately."
So they gathered a lot of fMRI data from people watching several hours of YouTube videos (the training set). They then use this to train some sort of machine learning algorithm to make a model. The pictures you see in the article are from a running the model on a test set which does not contain any of the videos from the training set.
To borrow and extend an example from Ender's Game, if you know what the train schedules are, you can figure out troop movements; even if you don't necessarily know which particular unit is going to which place, you still have a pretty good idea of what the military is gearing up for.
He was referring to AI algorithms, but seriously, who would have thought that having YouTube would lead to this?
It seems valid to me. There's no reason to ask them to somehow extract this visual data "unbiased", without bootstrapping off of video clips like that.
Actually, I'd commend that link to anybody posting complaints here, it covers everything everybody is saying as of my writing here.
I agree that their approach seems valid. There is a reason to ask them to extract the visual data in an even more unbiased fashion, though: if we understand how the brain is wired, then it should be "trivial" to back out the image from the patterns of activation.
Of course, the previous sentence is making a couple assumptions that I don't think are anywhere close to valid. 1) "the brain" implies that there is a single, nearly completely conserved architecture that is remotely similar from one person to another; 2) I think you'd need to get the activity to much higher resolution than fMRI can give you; 3) the stimulus <--> response mapping is moderately close to bijective, so for a given input, there's only one set of activity, and vice versa. Still, this study is an interesting first step on what will, no doubt, be a very long journey to improve the technology.
The image on the left is what the subject is being shown.
The image on the right is a composite from a set of other clips that match the brain activity observed when the subject is shown the clip on the left.
It's important to note that they're generalizing from a few hours of training data to millions of videos. So the classifier has to be picking up on something deep for it to be re-applied in such a flexible way.
I sort of imagine this approach as being akin to the way Bumblebee (the yellow VW Beetle in the first Transformers) lost his voice, but was able to communicate by switching between radio stations. As that recomposition process being richer and richer, it starts to approximate the real signal...
On the other hand, this whole thing is starting to look like the Deus Ex Trailer...
Source: "In practice fitting the encoding model to each voxel is a straightforward regression problem. " (http://gallantlab.org/)
Among the caveats: really high-T magnet, not feasible for general use; the visual stimulus was actually a warped version that would be reproduced as the nicely shaped activation upon retinotopic projection (it's a log-polar mapping)... Still really cool.
(ps: no-paywall version here: surfer.nmr.mgh.harvard.edu)
I suppose if it were properly anchored / built into the machine, you might be able to mount an LCD panel, and then somehow calibrate around it. Alternatively, something complicated involving projection and mirrors, maybe. The scanning tunnel is pretty damn narrow though.
It's strange, but when reading about all sorts of interesting science, I end up wondering about the methodology sometimes more than the actual results.
[1] something a bit like these: http://www.scansound.com/xcart/product.php?productid=16172... Gave me flashbacks to X-Men: http://comicattack.net/wp-content/uploads/2010/02/42.jpeg :)
There are two methods we've used, one was a goggle system where an array of optical fibers, one per pixel, are brought from the scanner tube to the control room and coupled to a LED display. The resolution is atrocious, and the thing is heavy, but it gets attached to the head coil so the person inside does not have to bear the weight.
The method we are using now is to attach a mirror to the head coil, and have a huge flatscreen outside the tube. You show mirrored images on the screen, and since the screen is big enough it covers the entire visual field. Works better.
The article states; "The reconstructed videos are blurry because they layer all the YouTube clips that matched the subject's brain activity pattern."
They could have just shown the best match, this would not have been a cool blurry image that is very easy to misinterpret as being generated from the brainwaves directly. More than a little slippery.
Not totally scifi awesome, but still pretty cool and I think it is a valid approach. Sounds like the bigger the library of clips mapped to brain activity, the more the technique converges to a desirable result.
It's easy for laypeople to misinterpret things they don't fully understand, but I think they can be disabused of misunderstanding by careful explanation. And we can get from here to there without invoking slipperiness.
Reconstruction of what the subjects are currently looking at is interesting but a direct window into the imagination would be something else.
I am not sure how well the algorithm would work without the V1 activation, since V1 is retinotopically organized, making it quite easy to decode.
So V1 is 'raw' and later cortices have performed more processing on the data causing it to be higher level, which in turn makes it harder to translate it back to a visual?
http://www.flickr.com/photos/brevity/sets/164195/
It's kind of a similar hack. There's a many to one relationship of images to a tag. Then that relationship is reversed and averaged out to get a consensus image. Of course this only works at the linguistic/labelling level, not at a brain level.
Potentially, it could be much easier than recognizing full images, as there would only be two possible outputs to discriminate.
Anyway, my point is that if there are some neural activity pattern that is correlated to lying, it could be easier to extract that information than it is extracting full images from the visual cortex.