Noise-canceling headphones without the headphone
spectrum.ieee.org
spectrum.ieee.org
This sounds pretty fundamentally different from normal ANC. A common misconception is that ANC headphones "predict" the incoming noise and cancel it (and thus "they're good with constant noise but not sudden changes"). In reality they're just hybrid open-loop/closed-loop feedback systems. Specifically, the mic inside the cavity works with the speaker as a closed loop system to equalize the pressure inside in real time. This works for wavelengths smaller than the cavity size, give or take. So there is no prediction, and they don't care whether the noise is constant or not. They just have a given frequency response, and above a certain frequency all you have is passive muffling (no active cancellation).
But this can't do that, because the speakers are much further away from the user. So it had to predict to get any kind of decent frequency response. I wonder how they're doing that. A lot of noise sources are uncorrelated and not predictable...
Edit: They cheated. They're taking direct feeds from the source loudspeakers, and then all they have to do is track the transfer function to cancel that out from another speaker. This will never work in the real world, where sound sources are not coincident with your microphones (or lack thereof as in this case). Cool research, zero real world application as a true ANC system.
I have no idea why they're selling this as ANC. This looks more like a system for computing real time room/response calibration as the user moves around their head. That's valuable, in a completely different use case.
I encourage you to read at least those before you dismiss the research and its real-world applications.
TL;DR: The purpose of the research is to try out a new way of figuring out which sound is going into the user's ears, so that it can be cancelled out. Current methods work for lower frequencies but not for higher frequencies.
They "cheated" because they're testing one component of the system and the "cheating" component is immaterial to the research they're doing.
What they're handwaving away is the real problem. The reason why current methods work for lower frequencies and not for higher frequencies is that your microphone, emitter, and ear all need to be enclosed within a volume << the minimum wavelength you can handle, to eliminate sound field and propagation time effects and let you treat the noise cancellation problem as a point problem in space.
Once you put the ears in open air and the emitters away from them, the problem isn't figuring out the sound that's going into the users' ears, it's figuring it out in advance. The ANC speakers have to send the complement signal before the original signal arrives at the ears, so they both arrive coincidentally.
In other words, the "missing part" to make this research work is time travel (if you use their sensing technology as the only input). Or maybe full-volume sound field sampling and characterization technology (if you don't). Neither of which exist, and which are orders of magnitude more difficult than this research.
The handwave part is "The constraint of the possible locations [of microphones] would affect the control performance and this remains a topic to be further studied in the ANC community."
Yes, it would completely destroy high frequency performance, which is what they were trying to achieve, in anything but lab scenarios (such as the sound sources still being loudspeakers and the microphones being directly in line between them and the user). Real world noise doesn't come from loudspeakers, it comes from all around you. Good luck using distant microphones to compute the expected sound at a user's ear in advance, with any kind of accuracy in the high frequencies.
At best I imagine they could achieve adaptive cancellation for a set of slowly moving pointlike sources in an otherwise simple room, with a number of microphones greater than the number of sources. It's like RF MIMO systems. But again, this is orders of magnitude more work to implement, and real world constraints are going to kill your high frequency response. And for larger sources - forget it. You just can't characterize the transform for that. Not enough dimensions in your input data to solve for it. So anything mechanical, stuff where the noise isn't coming from a literal speaker with a 1-dimensional input signal - nope. As the uncertainty and source size grows, your high frequency response goes down the drain.
Speakers, located in the seats, if they knew where ears were, could play sound to cancel that.
Decent microphones are $2, and what was expensive were things like microphone preamps, ADC, and the whole data pipeline. From there, we need a shit-ton of computation which really wasn't practical until... big-ass computation systems came out for the rise of ML.
The problem is hard today, but definitely not impossible. I don't think I would have given that same answer a decade ago. An NVidia Titan V brings over 100 teraflops. That's a lot of cycles one can throw at trying to predict sound by my ear from sound at the skin of the airplane in real-time.
https://www.youtube.com/watch?v=te3pyLk_wBs
Spectrogram sample: https://mrcn.st/t/flight_noise.png
I get that this is HN, but no, ML and AI do not magically solve all problems.
Thankfully, most of the energy is in the low frequencies, which existing noise cancellation systems can already do a god job of dealing with in the near field.
Instead of focusing on the fact that these researchers didn't singlehandly revolutionise ANC in a single study, how about we focus on what they did do?
Like, the idea of remotely sensing where a user's ears are and what they're picking up is useful. Just not for far field ANC.
(A) "This" works for a wavelength < L.
> and above a certain frequency all you have is passive muffling
(B) "This" does not work for a frequency > F. <=> (B') "This" works for a frequency < F.
---
Don't A and B' (or B) contradict each other? Because an upper limit for the wavelength equates to a lower limit for the frequency - not an upper limit.
What is your background that justifies your very confident demeanour?
A long wavelength implies that the phase of the reference input and feedback input are similar, which is what allows you to just subtract them.
Your processing delay, the distance between the speaker and the microphone, and the speed of sound are all fixed, so as the wavelength gets shorter, the relative phase difference between the two signals you're trying to subtract increases. At a certain point, you run out of phase margin and can no longer just subtract them anymore.
If you were designing an ANR system, you'd probably want to put a low pass filter somewhere around this point to prevent your feedback system from oscillating.
My background is audio processing is one of my hobbies, I have torn down and repaired noise canceling headphones and understand how they work (and have previously talked about this on Reddit and been corroborated by an actual engineer at a company designing such headphones), I have implemented laser galvanometer closed loop feedback systems and worked with them (which is both a similar control system problem, and a different part of OP's research paper on the light path), I have attempted room and speaker response cancelation (and wrote my own DSP firmware for implementing biquad filters, which I use for my DIY living room rig), making an open source speaker calibration system is in my future project list and I've recently done a lot of thinking about this... and, most importantly, since this is neither my professional nor academic field I have no reason to overstate my claims or mislead about the importance of my work, unlike way too many academic researchers (see also: that guy who I recently called out for being a paper mill about side channel leakage methods after he successfully spun a false media story about generating Wi-Fi signals from RAM using a misleading paper title; in that case infosec is in fact one of my professional fields).
I don't like to brag or anything, I'm just a curious guy who likes tinkering with (lots of) things... but if you insist on asking "what are your credentials", well, there you go.
I'm a silence junkie living in world getting noisier with every day it seems - so this subject is of high personal importance to me. That's why I'm interested in getting an idea about where to place your assessment.
But room-scale or free-space noise cancellation only works when either 1) your noise source is extremely predictable, or 2) your noise source is low-dimensional. It's just never going to work for the general case unless you can somehow sense the entire soundfield. The entire sum of what waves are traveling in which directions. Not just a few microphone inputs. It's going from 1D data to 3D data, the difference between a light sensor and a CAT scanner. And it needs to work at 20kHz response.
I guess hypothetically you could do something like wrap a room in microphones and compute the sound coming in in all directions, then with a very good adaptive model be able to cancel noise from outside within it. But if you're going to draw a boundary and wrap it in a massive array of microphones... Aren't you better off just adding insulation material? :-)
Edit: just as a baseline, I very crudely tested my Bose QC20 earbuds (which aren't exactly bleeding edge tech) and I think they get about 10dB active cancellation at 1kHz, which goes down to nothing at 2kHz. Human voice fundamental frequencies are typically 85-180 Hz for adult cis men and 165 - 255 Hz for adult cis women, so this level of ANC does get rid of the fundamental and a few harmonics. I might test a friend's AirPods Pro later and see if they're better. 1kHz has a wavelength of ~34cm, so I think there is room for improvement here in the earbuds case. I'm not experienced enough in this field to have a good feel for the numbers beyond order of magnitude estimates, but I think going up to 4kHz might be doable, given that for earbuds we're talking about cavities in the ~5cm range. Might need more mics/drivers and more miniaturization to pull it off, not sure.
For those wondering, the test methodology was to play a tone, turn on cancellation (which has a delay), then turn off cancellation (which is instant) simultaneously with decreasing the amplitude of the tone, and trying to match the perceived loudness between both cases.
At my current place I've pretty much concluded that anything but deep bass just doesn't transfer. Even at 5 in the morning, I can listen to stuff at a reasonable listening volume without disturbing my neighbor, as long as I'm not pumping out bass below 80Hz or so. And the neighbor recently got a dog and I haven't heard her bark even once when she's home :)
Well, let's do the math here. I can't find the data sheet for the weight of an electret microphone cartridge, but let's assume 4 grams. A thousand microphones -- assuming we want 100 the length of the airplane and 10 around -- is around 4 kilos. Add another 2 kilos for computation, and a few more for wiring, and you've added the weight of a piece of luggage.
Now, let's say you want to add 2 inches of mineral wool sound insulation. I'm assuming 100 meters x 10 meter circumference (I have no idea, but the numbers scale the same). You've added (quite literally) half a metric ton to your airplane. And made it 2" thicker, either reducing cabin space or increasing drag.
Plus, sound insulation does very little for low frequencies, which is the majority of airplane noise.
This is all assuming this whole idea works, which it probably won't, because it's not just about the microphones on the fuselage but also how sound is transmitted inside the plane and other noise sources.
1) Assuming system-level model
10,000 microphones * 8 kilosamples per second = 800 megasamples per second.
NVidia Tesla does 100 teraflops. I get about 125,000 FLOPS per sample. That feels adequate to me!
2) One filter per microphone per passenger. We need to divide by 250.
500 FLOPS per sample per passenger. That's more than enough for a very fancy IIR.
Of the two, I think #1 is more likely to work than #2, precisely because you want a coherent model. If you want to adjust your model for someone walking down the aisle (or any kind of system ID), that's a lot easier with 125,000 FLOPS per sample than 500 FLOPS per sample.
Again assuming this all works, which is a massive if.
Look, this isn't practical.
The answer to that might be right now (we couldn't do it before, and we probably can today) or it might be in a decade. But electronics will keep falling in price. 10,000 microphones * $5 per microphone+electronics = $50k.
The advantages of cancelling an entire wavefront go well beyond passenger comfort too. Industrial noise cancelling systems are more about equipment life than about employee comfort. I had laptop screws unscrew on airplanes before, due to vibration. If planes need less maintenance as a result of active vibration reduction throughout the airplane, you'll make up that $50k virtually overnight.
As a footnote, the crazy part here isn't the 10,000 microphones, but the speaker-at-every-seat part. You'd almost certainly want both the microphones and speakers in the skin of the airplane. But that's a story for another day.
Noise-cancelling headphones work the way OP described, and are limited to working in a cavity which is smaller than the wavelength.
There are predictive noise cancelling systems too. These are commonly built into industrial equipment. You might, for example, have a fan or motor which spins at 5000 RPMs, and what do you know, it has a consistent noise profile on each of those rounds. If you cancel that noise profile, you do pretty well. These systems both reduce noise and extend equipment life by reducing vibration. They also work to higher frequencies (in part because speed-of-sound in metals is > 10x that in air, and what matters is wavelength more than frequency).
There are ones which are open-loop used in HVAC systems. If I have a pipe, I can record sound on one end, and play a cancelling waveform on the other end, with an appropriate delay. I haven't followed the field, and when I looked, these were mostly expensive prototype systems (I'm not sure if they ever made it into mainstream use), but they did work. They needed to be calibrated for each HVAC, which made them impractical for most real-world applications.
There are all sorts of ways to do active noise cancellation. Many of these could be implemented on airplanes. Most of these techniques were implemented decades ago, when computation was a lot more expensive. A simple feedback loop is a few opamps, capacitors, and inductors, so pretty cheap to toss into headphones with 1990's-era technology. One could do a lot more with teraflops....
We're talking headrest ANC here (for use cases like consumer ANC headphones), not built-in noise reduction for industrial systems.
And in either case, the design of algorithms depends on much more on context than on whether it's "consumer" or "industrial." In this case, the context looks a lot more like industrial machinery.
I dont think its due to intentionality just that low pitch is easier to cancel than high pitch things like voices.
But what if the software/ai is advanced enough to reproduce a sound, but erase a certain aspect of it? Like how photo editing can edit out an object or background? Then the headphones can use a seal to completely block out all noise, and play only the sounds the user selects!
Earplugs below my bose headset work wonder to block out anything short of cataclysmic, but it's hardly practical. And I can't exactly play music with those on...
The problem is you can’t possibly work out exactly how and when the sound will hit each set of ears around you to be able to direct a beam of sound to their ears (assuming it’s even possible to have such directional sound)
The problem is significantly easier with headphones. You have a headphone and speaker in between your ears and the sound. The distance between the microphone and your ears is constant so you can perfectly time the negative sounds.
Directional sound is possible to an extent, Woody Norris gave a TED talk about it a long while back, it is called hypersonic sound. I think it uses ultrasound to generate compressions and rarefactions far away. I realize it sounds like science fiction, but you could perhaps transduce sound to electricity and again to sound, sending the signal of your voice away from you faster than the speed of sound, in time for a cancellation wave to be generated and have an effect. I suppose the analogy would be quiescing a ripple in a pond.
The utility is limited by the limitations of passive noise attenuation devices. These things are great for the shooting range, because bringing gunfire down to a level that won't damage your hearing is pretty easy. There's no kind of earphones that can just completely block out all sound, though.
That is simply not true - at least in my experience with an old model of Sony headphones. It's possible that prediction plays some role, but I bought mine with specific purpose of silencing neighbor's kids.
They work really well with this kind of sporadic, random noise and the difference between turning on active cancellation and just passive attenuation due to ear muffs is very noticeable. I'm talking about low-frequency noises that are mostly transmitted through the vibrations in the walls and floors - feet hitting the floor while running, ball bouncing of the walls and so on.
Only few care for the latter and millions/billions care for cancellation of external sources.
Besides, "cancelling the human voice" would just provide a false sense of security, as said voice would still need to be received by others -- and like any message it could be intercepted.
Worth noting that the feeling of privacy is a lot about perception. People feel better if they know those they see around them cannot hear their private convo. That is why they speak softly and put their hands to their mouth talking in public places (the other part was to boost the signal, before the advent of differential microphones).
Don't get me wrong, I most definitely agree external cancellation is terrific, be it active or passive. In the passive realm people use traditional dampening materials, as well as MPPs (micro-perforated plates) in certain industrial applications.
If you want to cancel another human, just build a wall.
Does anyone know what type(s) of injury is this referring to, how long is too long and how widespread/likely it is?
The Sightlines Gel Ear Pads from NoiseFighters is a popular option for earcup replacement.
The IceVents Classic Ventilated Headband from Qore Performance is a popular option for headband replacement.
The skin around ears have been super sensitive due to what I assume is just the natural heat build up from wearing headphones and pressure on skin not used to having pressure on it all the time. The skin is raw and irritated.
It's not terrible, but not comfortable for sure and has been limiting how much I can wear any over the ear or even on the ear headphones.
My bose quiet comfort are comfortable for maybe an hour or two before i even notice I am wearing them.
From there small in ear with noise cancelling can be good
If your ears stick out or are bigger you should really go to a best buy or music shop or whatever and try on cans until you find ones that fit
The only kind of injury I can think of is related to being less aware of your surroundings.
I'm aware of how noise cancelling technology works it's by creating an opposing sound wave of equivalent energy to nullify the unwanted sound wave.
My theory is when I am off a call and I have the headset on my head with ANC enabled it's working to cancel the background noise. It's only needed on a call but it's obviously on all the time. These are not cheap headphones I've seen them for $300 so not cheap but I'm sure not the highest end.
I've tried taking them off, put around my neck, on my head between calls. But a few times I've fumbled them and hung up on someone or picked up before I was ready. My solution was to turn off ANC which made the sound clearer and my ears don't ring as bad as with ANC on.
My own anecdotal evidence but there is an obvious effect on my ears. I was about to go to the doctor the ringing was getting so bad.
That said, like with most IEMs, sleeves should be replaced sometimes (probably once a month or if you get any dirt on them—be careful with handling and don’t let that stuff get into your ear canal), so that should be factored in the price.
Which sleeves you choose makes a big difference—they are more or less the noise isolation part of IEMs, and they make or break your comfort when wearing. I swear by the combination of SE215 + Shure’s original black foam sleeves (or, if those cannot be obtained, Westone’s foam eartips). Passive noise isolation in this setup is better than ANC in AirPods Pro for me subjectively (although according to rtings.com’s measurements they go head to head), and is easy to wear for long periods of time.
(I guess I’m replying to thepostoffice but for the most part this is addressed to dghughes’s upstream post.)
IEM seal the ear canal causing a build up of moisture and ear wax that can lead to infection in some people.
I am assuming this is a fairly common feature for headphones? If so you might be able to switch models to one that has an off mode, if yours don’t.
I wear Sennheiser HD1 M2's (with ANC) for ~8h/day, 5d/week, and have been for over 2 years -- with no discomfort or unwanted side- or after-effects. Ofc I'm careful about volume level, and I take frequent breaks from both screen and headphones (albeit for the sake of my eyes). I'm a healthy 47yo software architect by trade, also a lifelong musician and music afficionado. I care about -- and for! -- my ears and hearing, and would be grateful for any good links to credible sources (ie peer-reviewed scientific studies) which highlight hidden dangers or damage.
Was this ever verified quantitatively? It sounds like the kind of thing that could easily be generated by group confirmation bias.
Are there any other cars that do this?
And yes. Mostly luxury cars ATM but it’s growing.
Noise is only a little bit better. It’s still embarrassing.
There's a huge runway in front of us for both products and technology.
'Acoustic Holography' has worked for some time, commercial implementations have been few and far between, mostly because what's in the market is already selling boatloads without it.
It'll come eventually, there are a few groups dedicated to it.
A successful anchor in that space will allow companies to go deeper into applications.
Acoustic room splitting, Silence on demand, etc.
https://en.m.wikipedia.org/wiki/Wave_field_synthesis
The math says, if implemented correctly, WFS is better than sitting in the sweetspot of any speakers. Also, there's no sweet spot with this technology.
I've been using ANC headphones since I started woodworking in the garage and it works for me but not for my neighbours. Soundproofing the garage is difficult and not always possible (not when the door is open anyway), so I would pay top dollar for something like this.
So, assuming you had a magical beamforming speaker array around your circular saw, you'd still only be able to cancel frequencies around under 1kHz, because there's no way to get the thing in phase above that.
This is why noise-canceling earphones have much better active cancellation at high frequencies than headphones, which rely more on passive isolation until lower frequencies. They're smaller.
[1] Not sure exactly what the name of this device is.
https://www.flightglobal.com/quiet-revolution/31793.article
https://apps.dtic.mil/sti/pdfs/ADA446777.pdf (see plots at end)
https://en.wikipedia.org/wiki/Laser_Doppler_vibrometer "Security – Laser Doppler vibrometers (LDVs) as non-contact vibration sensors have an ability of remote voice acquisition"
Play enough noise, and the system will explode into fire...
Those massive stacks of speakers at an arena rock show are (in total) perhaps 20kW. The lights use far more.
Earbuds can we well south of one watt.
You can get the equivalent of state of the art ANC performance with a pair of passive IEM's that fit properly.
Try using memory foam tips if you have trouble with tips fitting your ear shape.