DIY Acoustic Camera
navat.substack.com
navat.substack.com
The resolution is fine enough that with COTS parts, I can record my signature simply by sketching it out with my fingernail on a table.
Every few years I dust this off and play with it, wondering if there's some application or other way to "turn this into money" (an increasing concern in the coming months...<tiny>PLZ HIRE ME</>), but I'm not a "product guy".
I'll answer some questions about the technology, but would really love to know if anyone here has advice on somehow using this achievement to pay rent. :)
The problem I have trying to find the niche application is that there's not much (at least, that I could think of) where you can have high quality audio data, but where a simple camera wouldn't work. Also, full imaging (as opposed to just tracking the largest/loudest source via TDOA) is quite different math, stupendously more computationally intensive.
Monitoring structural vibrations is also useful, and I think is an ongoing research area. I mention this because it's possible to sell it and research it at the same time.
What about synchronized cameras in different locations?
The current tech is spooky trash, parlor trick-quality from what I've used. Every time we use some of the automatic gizmos in conference rooms, we get tickets to make it stop.
Pick your favorite top video conferencing platform or camera maker and they'll want to improve what they have. The Creepy - "Just works" jump is a big one.
PS., the industry trade show is happening RIGHT NOW at https://www.infocommshow.org/
Could you put some sound sensors inside of some mechanical structure and use the acoustics to figure out where some physical contact is happening on the outside of the structure?
Specific Application: prosthetics devices that can - with only a few acoustic sensors - determine where the 'touch' was on the outside of the device.
If viable, it may be similarly useful for robots - or any machine in general - that needs a low-hardware (and thus low cost) method of getting course tactile information around it's boundary.
In my ideal (and in the best jobs I've had in the past), someone finds me (or I find someone) who I can share a list of "cool things I've figured out how to do, but don't know the usefulness of", and that person then tells me what to build.
For your "where is that annoying sound coming from?" service, what sort of scale do you imagine, and what form factor?
A handheld consumer device with a range of ~10m which points in the direction of the loudest thing?
How much would you pay for such a device? (would love your thoughts in email)
After setting up the array, once could answer questions like "where is that darn squeak coming from!", and even characterize the undesirable noise, both spatially and spectrally. You could also measure how effective noise isolation materials are, but I don't see what the spatial information gets you.
However, if weaknesses due to "perimeter gasketing, frame, door leaf, or wall construction" result in some sort of localized noise, then a system like this would certainly pinpoint it.
Acoustic cameras like from Noiseless Acoustics are on the market though they seem to be marketed to industrial customers. There are similar mapping systems using a scanning mic like from Soft dB.
It's sensitive enough to noise that I can pick up (and locate) the air vents in a room, even when the sound is at the threshold of hearing. Noise (pink and white), and even more so MLS (maximum length sequence) really "jumps out" (it's very obvious), well below my threshold of hearing.
There's so many interesting areas of research I've never had the time/money to fully investigate. I'd love to play with an "active" system, not just "passive", with a goal of experimentally finding modes of resonance of objects in a room.
I bet once can tell the relative contribution to noise of one physical object over another, but I don't know enough about construction to know if one would be able to separate the door frame from the perimeter gasketing. You do need line of sight for it to work. At the least, you'd have a way to quantify the sound leak, with numbers and reproducibility.
FWIW the Noiseless Acoustics camera costs ~$18k(!)
I had the "opportunity" as a patient in a very new hospital wing some 7 years ago for about 13 days. While I had a private room, the door was always open to the walkway, I guess as is normal to allow quick response by medical staff. But at one point I really felt overwhelmed by all the external noise that I could hear from the other rooms in the ward, nurse stations, and so on. I really felt that the hard surfaces and even the angular nature of the floor layout was conspiring against me - almost focussing noise into my room. I imagine given enough data you could show the poorer rest of patients prolonged their recovery period and hence increased bed occupancy and cost to the health system (important in a state run public hospital service that is predominant in Australia).
As such it would be nice to think hospitals, schools, offices would include thorough acoustic assessment to at least allow appropriate mitigation of noise during design (before having invoke more active measures like soft furnishings, etc).
Schools very often have acoustics reviews, although more often in cities than rural areas. Classrooms in addition to auditoria, gyms and common areas. Standards exist for those too.
Office buildings are hit or miss. The developer may hire us for a base building review. Tenants' architects hire us as they design their workplaces. There's a lot of push and pull to find a balance of the modern open ceiling industrial aesthetic and glass conference rooms with reasonable acoustical goals.
Working in conference rooms, people far too often are concerned about the look ( big windows, natural light, great table ) and less concerned about acoustics than would be reasonable. In my experience, Architects think people just sit around the table, and chat.
I once had a brand new office upfit project with a "flagship room" that was a large flex space, (could be a board meeting, could be a hackathon). There were three sides of glass, a concrete floor and metal ceiling structure. If you clapped in the room, it could be heard for 1.5 seconds as it bounced across all those surfaces before decaying.
My bosses were furious at the microphone quality, the installer was unhelpful and bailed on us. We hired a consultant to perform an evaluation and he told us that the room was awful, with lots of numbers ( figures for acoustic reverberation at different frequencies ) and told us of certain products that could help in those ranges.
It is a lot easier to design rooms with acoustic features than it is to retrofit them in terrible sounding rooms.
Kids must tiptoe across in front of the camera.
If they make any noise, the camera uses a powered gimbal to aim a hose at them and squirt them with water.
And a terrifying siren goes off.
I've been considering putting multiple microphones on one side of the house to track birds. I mainly want to isolate the audio for recognition. But setting up a PTZ camera to get good shots would be even cooler.
This was one of my first at home demos! Tracking birds and cars. If you want to get started DIY and figuring it out yourself, try about 5 microphones, each 1m apart. The trick is having all the mics in line of sight to the object in question. Otherwise you're left with doing an "intersection of angles" method, which is much simpler, but horrible in terms of precision.
In a later demo I used a PTZ laser pointer to track the moving sound sources, so, it's at least possible!
It's a big field, but if you could, for instance, set up an alert for when a noise source crosses a property border, or when something that sounds like a human comes within x meters of a particular building, I think that would be valuable. I'm no longer there, but happy to talk if you're interested.
I know very little about drones, but, in a demo I made years ago (trying to convince someone to give me money for using this in a drone defense product), I was able to "fingerprint" different drone models, even sometimes distinguish between two different drones of the same model. As long as there's line of sight, you can sometimes "see", in the data, slight changes in the speed of the propeller, as well as the rough "shape" of the drone itself.
But for all those applications, even though I love sound, I always think to myself.... "wouldn't a pair of cheap cameras do this better?" :)
> no longer there, but happy to talk if you're interested.
Thanks!
That's part of where they've gone (though the cameras are far from cheap) as well as RF, with some AI magic sprinkled on top.
I wouldn't suggest doing drone defense, but smaller-scale asset protection might be more approachable. There's what, 3000 local governments in the US? I'm sure a lot of them have had a tractor stolen from a road construction site. Or maybe if they lease their equipment you could sell a solution to the leasing company.
[1] https://www.youtube.com/watch?v=rEoc0YoALt0 Explainer Youtube video about Motion Amplification
I bet if you really worked on tuning and filtering you might even be able to pick up the vibration of a persons throat to hear what they say.
I'm curious how you would combine acoustic localization in 3 space with motion amplification. I unreservedly agree that they are both "super cool", but don't see how they tie together to make something greater than the sum of their parts.
The only thing I thought of is, if two data channels (video, audio) are registered accurately enough, one could maybe combine the spatially limited frequency information from both channels for higher accuracy?
For example: voxel 10,10,10 is determined (by the audio system) to have a high amount of coherent sound with a fundamental frequency of 2khz. Can that 2khz + 10,10,10 be passed to the video system to do something.... cool? useful? If we know that sound of a certain spectral profile is coming from a specific region, is it useful to amplify (or deaden) video motion with a same frequency?
And since Intel, Google, Facebook etc keep buying startups that produce cool things and preventing them from producing more cool things (North Focals being the most recent I'm aware of) it's gonna be a while
Does anyone know if the array used here supports timestamped samples and/or clock sync to support multiple arrays? Or is it a single 16-channel stream?
Having done some very primitive dabbling with this stuff, the DSP programming is always the most interesting part to me. These folks are killing it with some really cool 3D scanning integration to the acoustic analysis
You can do that, but the gain isn't as pronounced as you'd like. A 12-16dB gain doesn't sound that dramatic.
Now, combined with some other newfangled math, like neural source separation, you might be able to do something spectacular...
Another approach to this is the Ambisonics method of capturing the directional soundfield at a point. But you'd need to use a high degree multipole expansion to get resolution anywhere closer to video.
Link to a previous post about this mosquito turret concept: https://news.ycombinator.com/item?id=27552516
My cat would serve a similar role for larger insects. Her eyes would track them for me so I could locate and destroy them. Unfortunately she either cannot, or more likely chooses not to, track mosquitos.
In demo, the two angles drove a pair of servos steering a laser pointer. Followed the loudest object around the room :)
IME, finding a way to communicate the information to the user is often non-intuitive. That is to say, once a device has located birds in trees, how would you like it to inform you?
Keep in mind "the black box" can output the position in 3 space (x,y,z, measured in mm) of coherent sound sources, but to know where those are relative to the camera, so that once can draw a little arrow, can be hard.
I'd like to try hooking it up to a VR/AR headset, since I imagine those already handle the task of knowing precisely where my head is and where it's looking.
1. mount the array on a tripod somewhere in the frame of the camera 2. the array is covered with an assortment of fiducials, 3. software uses the known intrinsics and extrinsics of the camera to figure out the array position relative to the camera 4. do the obvious thing with chaining transforms until you get the sound source position relative to the camera
If so, I think that would work, but would be a lot of coding to do all that CV...
> phone that has AR support
I take it cell phones now do much of this work?
https://acousticstoday.org/wp-content/uploads/2020/06/Battle...
The fan blades should cause doppler shift and changing amplitude that varies based on the location of the sound.
I suspect that after just a few seconds, this would give better information than an array of 16 microphones.
Not to say wouldn’t work, you would get results, but they will be based in a different strategy.
Sure, the maths is complex... But there is only one source soundwave and location which causes a given smearing. The challenge is to find it...
When I first saw a popular science article about them, I got excited about incorporating them into an array, but couldn't find any technical details, just a lot of what looked like vaporware. Is it anything more than three orthogonal pressure sensors? (aka.... 3 microphones?)
This PDF may be helpful http://past.isma-isaac.be/downloads/isma2010/papers/isma2010...
I was interested in making my own alexa-like device, but it seems mic arrays are sooooo expensive - more than the cost of an alexa device for the least expensive one i can find :/
I made a 16 mic array out of a bunch of trash-picked and cheap 4ch ADCs.
The actual mic capsules are likly far cheaper than $1 a piece (probably closer to $0.10 than $1) but the mics in an array need to be phased-matched. The two approaches to getting phased microphones are 1) building them using precision techniques so they are phased-match from the start (which is expensive and why pro phase-matched mics are around $1,000 each), or 2) get a whole pile of cheap mic, test them one-by-one (or really, pair-by-pair) and select the mics that are best phased-match to use them in the array. The #2 approach is cheaper, but does add cost.
I've never used phased matched mics in my arrays (can't afford it!) and also have never needed to "bin" them. ("pair-by-pair" testing).
0. https://www.minidsp.com/products/usb-audio-interface/uma-16-...
1. https://www.digikey.com/en/products/detail/knowles/SPH1668LM...
2. https://www.digikey.com/en/products/detail/analog-devices-in...
Not that I can find! Building the array is way more expensive than it needs to be.
I have limited EE knowledge, so have been stumbling through it on my own, building my first array out of reference microphones, another with $10 omnis from guitar center, and one with 8x, cheap, repurposed webcams.
Right now, my limiting factor on driving the cost of a future array down is that I haven't figured out how to get a lot (at least 8) I2S inputs to a micro-controller. If that were solved, it would be easier.
Main limitation is USB 1.1 IO, so ~1MB/s, unless you are fine with recording to SDcard then probably around 10MB/s. Pico itself can interface 29 microphones with no sweat (30 GPIOs, 2.0 GB/s internal bandwidth).
> Pico itself can interface 29 microphones with no sweat.
I... had no idea. I thought that since it didn't have an i2s peripheral I was going to have to either find a micro that did, or do something bitbangy using SPI and perhaps an external buffer. I see that it might be possible to get a few I2S inputs using the PIOs. Thanks for this, will certainly give it a shot.
(though I don't see how you're getting 29 microphones :P "prove it" ;)
With high enough frequencies I can see reflections, but not at any distance, and the sound source has to be loud. Of course, I'm relying on line of sight, and perfect reflection. Any bumps in the wall would add some phase error I think.
If I ever get a chance to work on the problem again I'd love to see if anything interesting can be done with multipath.
$275 doesn't really seem exorbitant for niche hardware given than you need 16 decently high quality microphones. I eagerly await a ShowHN using $2 mics and cardboard instead!
But... you don't. :) The challenge I find is getting the data into the computer. That's what always costs the most. I've done it with 8x $1 mics and a used $100 sound card.
I may just order a bunch of i2s microphones...
Can you? :)
I know enough EE to do simple things, about the amount you'd expect someone who's worked as a firmware engineer to know. The fact that you're saying "Can't you ... cheap..."? make me think this must be a viable path though.
My hands shake too much for anything but the simplest soldering; is there a cheap FPGA board you'd recommend? And getting all the data into the computer... it could easily be ~70Mbps. (16 mics, 192/24) Making a custom USB class could be a mess, I wonder how hard it would be to just dump ethernet.
I don't have any specific recommendations for a dev board, though a selection of some cheap dev boards is here: https://www.joelw.id.au/FPGA/CheapFPGADevelopmentBoards
If you search around on Alibaba you can find some cheap dev boards that appear to have ethernet already (instead of you having to provide an ethernet PHY, which I think takes up something like 18 IOs?). Beware that even though most dev boards have a USB interface for programming it probably doesn't function as a USB for comms (i.e. you'd probably need to provide another USB PHY). Each I2S device probably requires 3 GPIOs. Someone who knows a lot more about this can probably make a much better recommendation.
If you do the beamforming on the FPGA then you probably only need to output a lot less data that might be easily done over a simple UART.
In the end, designing your own board with the right-size FPGA is probably the right solution, but that's only cheap per-unit in quantity and requires someone who more than sort of knows what they're doing. Though for someone who knows what they're doing it'd be a relatively quick project...
$5 dev boards on ebay with free shipping.
https://www.cypress.com/file/138911/download
https://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.73...
Are you kidding me???! It costs so... so... so... much less. I thought automotive might be a good application, considering all the doors opened by using more DSP tricks layered in addition to source localization. (I can localize coherent sound patterns s well as coherent sound)
I would love to chat with you, happy to buy a coffee or beer for your time. My email's on my profile.
Here's an article about a large installation at Porsche. https://www.azom.com/article.aspx?ArticleID=18378
HRTF stuff is fun, if that's what you're referring to! :) I've worked with some of that stuff before, including the stupidly overpriced mannequin heads.
> Some of these devices for automotive are large enough to surround a car on 3-4 sides, with several hundreds of microphones and the associated cables and positioning arms. Depending on where the devices he mentioned are being used, there are things like mannequins with heads and models of how humans hear for identifying sources inside a car.
Do you work in a field that would benefit from the same results, for a fraction of the cost? Or, if not, do you have any advice on how to find and talk to these mythical industries that could pay me? It looks like Porsche wanted to build their own, in house, but I'm hoping if it costs less than a tenth as much, maybe more people would want one.
But c'mon, they are not $500k. More like $20K.
https://www.fluke.com/en-us/product/industrial-imaging/sonic...
Obviously, Fluke, and the positive reputation that brand is known for, and reliable product support are worth a LOT, but I'm sure there's some $$ divisor beyond which you/someone like you would take a risk on something substantially cheaper.
In my implementation, there are multiple stages using a dataflow approach with lots of compile time optimization. In 2011 I could image a roughly 2m^3 space using 8 microphones at ~10fps in real time on 3 desktop computers, 2015 I was able to do 12mics, 3m^3 space, on 2 laptops, but that involved a LOT of custom numeric programming to shave cycles.
If I had access, I'd love to see what could be done given a well tuned implementation and modern GPUs. An efficient scatter gather OP (like what AVX3 has) would increase performance by an order of magnitude.
*OK, I've skimmed Acoular.
Edit: To clarify, the "opposite" of beam forming means using processing you can choose which direction you want to listen at any one time, like a beam. Then you can scan the beam across x,y and make an image.
what we do we make sure that all receivers are synchronized, i.e take samples at the exact same time
then you can correlate the signal received between dishes (which will arrive at different times due to delays in propagation), and find out the time difference of the signal which then points out to signal origin (beam forming) - this is how phased radar works
once you align the signals you can use the minute differences in the signals to compute a synthetic aperture, i.e improving the angular resolution
The major difference between a microphone array and an imaging sensor is the availability of phase information for the received wave. A microphone oscillates with the sound pressure wave, and that oscillation is translated directly to a voltage. Your software can see the full time series of that wave, so the information about it is 'complete'.
An optical image sensor, essentially, turns photons into electrons. The optical wave is too fast to turn into a voltage time series, so you only see the wave's amplitude at a given sample in time. Therefore, in order to turn it into an image, you need to recover some fraction of the phase information in some way.
A pinhole is one way to do that. One way to think of a pinhole is that it maps every source point to a distinct imaging plane point, so the phase of the wave doesn't matter as much to the final image. It acts as a filter that cuts out ambiguous information that phase would have disambiguated.
A lens performs a similar operation by interacting with the light wave's phase to bend wavefronts in a way that maps points on the object to an imaging plane.
Those approaches don't recover 100% of the phase information, but they recover or filter enough to form the image you care about. Light field cameras attempt to recover more complete phase information through various ways better explored in the wikipedia link.
Could you create a sound blocking plane with a pinhole that makes an acoustic camera that follows similar principles to an optical camera obscura? I bet at some level you could, but I also bet it would not be very advantageous. You still need a microphone array to act as the imagine plane. The size of the pinhole is probably very constrained by sound wave diffraction (it's a pretty long wave after all, compared to light). Using the directly available acoustic phase information is more compact and efficient.