Extracting audio from visual information
newsoffice.mit.edu
newsoffice.mit.edu
The device worked, but had a bug. a hummmm on the receiver.
For the longest time I could not figure out what was the "hummm", till I noticed it was there even with the transmitter off.. and then, I heard strange voices even with the transmitter being off!
Then I realized the humm was the reflection of incandescent light on the window, and the voices where the people in the room making the window vibrate too..
Ahh.. those where great fun days early on "hacker" life as a hacker. :)
...and that is when little Johnny knew he had to lay off the shrooms.
:P
I love hearing stories like yours though. I, too, miss the child-like wonder of making discoveries. We take the physical world we live in for granted, but there's just so much out there to see and learn from.
BTW: not everyone uses exactly your own personal guidelines.
BTW: HN Guidelines include "Resist complaining about being downmodded. It never does any good, and it makes boring reading."
For this particular film the rather characteristic look of digital video is appropriate, whereas most productions would rather shoot with much fancier gear. But you'll find plenty DSLRs working on film sets these days, and you will see more and more convergence as high-quality sensors become a commodity. Black Magic have a pocket camera that delivers 13 stops of dynamic range.
With two time samples you shouldn't be able to learn anything about the state of waves in a pool, but if each sample is a photograph with lots of pixels you can actually tell a lot.
> Because of a quirk in the design of most cameras’ sensors, the researchers were able to infer information about high-frequency vibrations even from video recorded at a standard 60 frames per second. While this audio reconstruction wasn’t as faithful as it was with the high-speed camera, it may still be good enough to identify the gender of a speaker in a room; the number of speakers; and even, given accurate enough information about the acoustic properties of speakers’ voices, their identities.
But generally what you hear is very close to how the person actually sounds - although their accent or inflections may be adopted for the purposes of their role. This can be a bit jarring; I've worked with method actors who maintain their screen accent at all times during production until the film is done, so when they switch back to their regular accent after a month or so it's extremely disorienting, since I've been listening to them in my headphones day in day out for weeks, and am paid to pay as much attention to their voices as the cinematographer pays to their faces.
Some actors go even farther in support of their public image. Rock Hudson had a somewhat high voice that producers deemed incompatible with his looks, so during production he would warm up every day by shouting for 20 minutes and gargling with orange juice to inflame his vocal cords, and of course he smoked a lot too. What actors will do to themselves in pursuit of screen presence far exceeds anything I've ever been asked to do in post. Editing dialog is more than enough work without trying to sculpt people's voices.
In the Indian film industry (maybe not so much Bollywood, but more of the local ones like Tamil, Malayalam, etc) there's a LOT more post being done. Some entire Malayalam films have dubbed voices (for the original movie). Not to mention nearly every song track does not have the singing recorded during the shooting of the dance sequences. So those sets and actor voices I would bet would be completely different.
Edit: They mention capturing frequencies up to five times higher than the 60Hz frame rate, which would mean a maximum frequency of 300Hz, which would suggest the equivalent of 0.6kHz audio, which is a 73.5th of the audio rate of a CD. I doubt you'd get intelligible speech from current consumer hardware using this technique.
Suppose the camera scans 720 lines in HD every 1/60 second. Each row is offset in time by 1/43200 second. A rigid object could be slightly offset in space on each line of pixels, indicating that sound waves perturbed it in the time gap between when the camera captured each line. So that subframe video data can be turned back into audio at a much higher frequency than that apparent 60 Hz video sampling rate.
In other words, we're not just talking about 60 frames-per-second from a camera. It's really perhaps 43,200 rows per second, an enormously higher sampling frequency.
Let's say that it would read the entire image in 1/120 second, then it is waiting and does nothing another 1/120 second before it starts reading next frame.
The real number would be significantly smaller. Therefore they can not bump the sample rate more then five or six times. And I imagine they are using some intelligent algorithm to evenly space out the captured samples already.
Yes, yes, that was completely obvious from the article. We are getting thousands of "measurements" per second.
However, each of those measurements is incredibly inaccurate. Each one is trying to detect the change of colour of 1/200 of the colour range in a single pixel. You may be getting less than a single bit of entropy per measurement.
An advanced signal processing technique will look at the longer-term picture. Sound vibrations are not a random walk - they tend to be a combination of sine wave vibrations, where the rate of change of magnitude of each wavelength is significantly lower than the vibrations themselves. Therefore they are to a certain extent predictable, and this predictability is used by audio compression algorithms. The signal processing algorithm will have to make use of the extremely limited information coming from the measurements, and match up possible sets of varying sine waves that could be causing those measurements. This may be sufficient to reject some of the noise that we could hear on that video, and clean up the sound a bit, but it is quite a hard (and CPU-intensive) processing task.
http://en.wikipedia.org/wiki/Cocktail_party_effect
http://en.wikipedia.org/wiki/Independent_component_analysis
Edit: Spacing.
http://www.youtube.com/watch?v=ZbpwBTDvXrI
It's in French, but this makes it even funnier.
Because of a quirk in the design of most cameras’ sensors,
the researchers were able to infer information about
high-frequency vibrations even from video recorded at a
standard 60 frames per second.
Can anyone explain what this quirk is?So having a rolling shutter is good for this specific application because it trades off resolution (most of which is redundant or insignificant information) for sampling rate.
- Actually, doing a quick calculation shows that at 1khz a 1/2 wavelength is just 17cm. I wonder how precise spatial scene/source information has to be to allow this diversity to contribute significantly to the sampling. If you had a planar source and precisely spaced two objects it shouldn't be too hard to increase spectral resolution. The complementary possibilities are also be interesting -- with precisely laid out N objects and a good spectral resolution for each afforded by the shutter you could perhaps resolve the sound into N distinct sources, allowing to determine the origin of the sound; with precisely known source locations you may be able to extract some object location information.
If you want to record a particular person through a window, for instance, you can get a laser microphone that catches the vibrations of the glass.
I often wonder what makes people to pass off things like that as "typical tropes", where they are obviously realistic and doable.
Just yesterday I chuckled when recalling a James Bond movie involving Bond driving a car in reverse by viewing a back-up camera. At the time, it formed an instant "trope" because it was so cool and novel. Yesterday, I was doing exactly that with the backup camera on my car - obviously realistic and doable technology.
Is the stability required from the camera a dealbreaker when it comes to outdoor mounted cameras in moving air, or would it be pretty easy to algorithmically filter that out? i.e. are winds and drafts predictable enough that they could be removed accurately enough for smaller vibrations to remain?
What would those ferns look like? What would they be made of? I'm imagining christmas trees made of cellophane fibers or something.
Kind of off topic but just last night I was watching some (silent) lightning in the clouds, thinking how much more localized sound energy is when compared to light. In other words, I can see for light-years but can't hear much anything 100ft away. Or perhaps it's just our sensors are more sensitive to light waves than air pressure waves. Now I realize sound isn't so localized! It leaks. /rant
1) A lot of the tests were done with a camera that costs thousands of dollars. You can see images of the camera used in one of the experiments.
2) They were able to use consumer-grade cameras to also capture sound. Even frequencies up to 5x higher than the actual 60 FPS of the captured video.
But the GoPro has a rolling shutter as well, so their second approach would be applicable. However, that effectively relies on rows per second and while you have a higher frame rate you have a lower resolution. In the end they could cancel each other out.
There are also some consumer cameras that can capture almost 1,000 fps from a small sensor window. I wonder if those would work.
Here is the original paper:
Now, a 2000 fps camera can see things that a naked eye can not.
Almost missed your comment btw.
http://en.wikipedia.org/wiki/Archaeoacoustics#Past_interpret...
$ youtube-dl -f 18 http://newsoffice.mit.edu/2014/algorithm-recovers-speech-fro...
You can use the -F option to list the available formats.
Or maybe it's time to think how to adapt to a world without privacy?
Not long ago there was a spate of HN articles about apps that could measure your heart rate via the camera (it watches for & measures subtle changes in your skin color which occur during the pulse cycle). This is exactly the same idea, just with a much faster "pulse".
I expect the researchers will next discover the "rolling shutter" (a "that's not a bug, it's a feature!" of cell phone cameras) and discover how to extract the audio info without the need for high-framerate cameras. atomatica found a perfect example: http://youtu.be/TKF6nFzpHBU?t=10s