AI headphones let wearer listen to a single person in a crowd by looking at them
washington.edu
washington.edu
In groups and with friends, it's inevitable that you end up in a busy restaurant or a bar, and it always frustrates me that I don't hear something, I ask the person to repeat only to not hear it again, usually because they repeat it at the same low level (considering the circumstances). Missing jokes and throwaway comments is even worse ("hey what are you all laughing about, I didn't hear it, could you repeat it for me like three times until I hear it").
After a whole day of tests the audiologist comes in and says I have good news and bad news and good bad news. The good news is that my hearing was beyond great, it was at the level of a 5 year old. The bad news: I could hear so well I was unable to differentiate sound; my hearing hadn’t gotten worse, my brain’s ability to separate sound had. The good bad news is that my hearing would inevitably deteriorate, as all ours does, and for several years I’d be able hear in public places!
I think part of what has made this worse is that restaurant and public space designers have stopped thinking about sound. Most bars opened in the last 15 years have cement floors, very little sound insulation, and they’re based on the idea that you’re not having a good time unless your ears are ringing.
I’ve stopped patronizing these places if only because I literally cannot maintain conversations.
Recently listened to a really good podcast about this phenomenon https://podcasts.apple.com/us/podcast/gastropod/id918896288?... (or pick your favorite podcast app)
Couple takeaways I remember:
- "Silence is the new luxury" -- restaurants can have good sound design, but it doesn't come cheap. Upscale restaurants are starting to differentiate themselves with sound design
- The modern clean aesthetic (glass, concrete, stainless steel, minimalism) promotes loud, echo'ey spaces
- "You're not having a good time unless your ears are ringing" was an intentional design choice popularized by some restaurant guru in the 90's. Growing awareness of the problems is starting to create a backlash
- Loud restaurants are damaging for the waitstaff's health. You can work for hours in an environment so loud that OSHA would demand hearing protection.
- The luxury sound design studios can be so good at isolating ambient noise that they also sell an "anti-noise-cancelling" sound system that actually selectively re-amplifies crowd noise for when you do want to tune up some sense of busy-ness (with too much sound dampening in an unfilled room, it starts to feel too isolated... being "out and about" is some of the reason people go out dining)
I should preface by saying that I have some existing tinnitus that developed from playing drums in rock bands in my 20's without proper ear protection. It's manageable and quiet enough that it can be largely ignored.
Recently I visited Las Vegas, and I ate at the well known restaurant Tao. It was so incredibly loud, for a sustained period of time, that it triggered the tinnitus in my left ear to such an extent that I was essentially deaf in my left ear the next morning. It gradually receded to loud ringing, and then was back to my "normal" level of tinnitus after a couple of weeks.
I guess I'll have to bring my own hearing protection to restaurants now. Kind of a sad state of affairs.
Wood floors, rugs, curtains, artwork, acoustic panels, etc.
It was really interesting but sometimes large scale interventions have potential tradeoffs with other desirable project criteria and goals. At the end of the day the client and the budget make a lot of decisions that seem careless but are just tradeoffs.
As you say - upscale and boutique businesses often have the desire and budget to differentiate themselves on sound design - having professionals consult is not inexpensive and often implementing their recommendations has a significant effect on budget.
Carpet the ceiling and the walls. Super cheap, super effective.
Ideally carpet the floor too - and if you use carpet tiles then when a customer spills something uncleanable on a tile it's a 5 minute non-expert job to pull up a tile and put in a new one.
I’ve had it since I was a kid because I always passed the hearing tests but every other kid had no trouble listening to music and understanding the words and so on so I put two and two together.
Anyway, I have never been able to understand anyone in any loud public space which absolutely blows when you’re not a home body.
See https://en.m.wikipedia.org/wiki/Auditory_processing_disorder
The Wikipedia page seems to really describe my condition, except for the potential overlap with ADHD. For example, the "difficulty following oral instructions". I can read something and retain it forever, but if someone speaks to me, it will often go in one ear and out the other.
I sometimes wonder if this could be a result of being a very introverted child who started reading at around 3 or 4 years old. (Because reading is so awesome, why bother listening to people -- and improving your auditory neurology -- after that?)
I think I've probably just adapted to the condition. It doesn't seem like any sort of problem or disability. But I suspect others around me find it much more annoying than I do. ;)
Theory: The bars and restaurants want young patrons, so the poor acoustics are a selection mechanism. Only young people can converse there, so older folks stay away. The place gets reputation as “young and hip.”
Whether by conscious design or “natural selection” for establishments, this seems to be the case.
I also have poor vision without glasses, and I’ve always found that when I go swimming (and can’t wear my glasses) my hearing also gets significantly worse. Or at least the cocktail party problem gets worse, as my brain seems to get overwhelmed by every single background noise. I think some of this is explained by many indoor pools being big echoey spaces, but it still happens at outdoor pools as well. I suspect that when one sense (sight) is degraded, my brain tries to compensate by focusing on another sense (hearing), and the end result is even worse due to APD.
You know, until you mentioned it, I'd never thought that that experience might have an ADHD-related component to it. Interesting.
My solution to this problem is to carry on my keychain a pair of concert rated hearing protection earbuds. It's not a perfect solution, but you'd be surprised at how helpful it is to have some highs and lows filtered out.
Only downside beyond not being perfect, is people tend to think I put them in to block sound entirely. So I usually end up fielding questions about what they are, with a resulting 'man that's a smart idea'
Ymmv
If the band is playing zip it!
I simply don't understand it, why the hell would I want a noisy place where I can barely hear anyone to have MORE noise!? it's not even good quality noise, it's usually top 40 from 10 years ago being blasted over shit speakers.
I also have trouble discerning sounds in crowded spaces. Thanks for sharing your diagnosis, really interesting to think about.
I actually find foam earplugs make voices easier to hear. I can have conversations during concerts with them somehow, in fact. So I figure if foam earplugs can do that for me, then earplugs designed to block only unwanted noise are probably even better.
Case in point, I was recently at a Swans concert—-they play very long sets, like 3 hours—-and my back got tight, so I started stretching. I heard someone 20 ft away talking shit! They said, “yeah, I guess you can do your Pilates here.” I finished the stretch and then heard, “oh, I guess he was just stretching.”
I’m not from the US but I visit often and could not understand this trend of party restaurants where you eat at club level sound blasting all around you. It’s the worst of both worlds, since you can’t talk with the loud music nor can dance since you’re sitting down.
Paraphrasing the great Donald Glover, I’m too old for this $hit!
That function of being able to mentally 'filter out' specific voices within a crowd is (semi?) common signal of autism. More accurately; I'm like that and I am autistic, I've read that it happens to a lot of others.
Not sure if audiologists phoning it in will know how to test for the latter
Is there a name for this condition?
It sounds like that audiologist still wasn't specialized enough - otherwise I feel like your story would have extended differently; high demand mainstream audiologists, including the mainstream audiologist profession, don't seem to have a certain lineage of knowledge that I dramatically benefit from first when I was 23 (now 41), when I was in essentially a high-functioning Asperger's state - to where my hearing had devolved into a hyperacusis state - a severe hypersensitivity to sound - but where prior to that I had the same symptoms of difficulty with conversation, and busy rooms with lots of noise was extremely mentally draining-fatiguing, not realizing it was putting my mind on overdrive drying to actively focus and hone in on sounds vs. it autonomously happening.
Have you ever heard of Berard AIT (Auditory Integration Training/Therapy) or the Tomatis method? There's a book on the sound therapies called "Hearing Equals Behavior: Updated and Expanded."
There's a non-standard audiogram to check for imbalances in the hearing. The standard audiogram is they just pick say 30 Dbs volume and check various frequencies in each ear at that frequency, and if you can hear that - then "great!" The non-standard audiogram checking for imbalances checks for HOW LOW-QUIET of a sound you can hear at different frequencies, and interestingly, with 100% accuracy you can predict a set of behaviours that a person will have if they have an imbalance at certain frequencies like 1000 Hz; not all frequencies have associated behaviours with them.
An example of an imbalance would be if at 1000 Hz in your left ear the lowest sound you could hear was 15 Dbs, but in right ear you could hear at 10 Dbs - an imbalance of 5 Dbs, but where the idea is that the body-brain-mind is a system of finding homeostasis and equilibrium, so it should be able to have it so sounds are heard at the same level - evenly, save for actual physical damage. This is just a simple example and there can be high peaks and valleys that show up in the non-standard audiogram.
I have a similar story as yours. My issues with sound were almost identified in Grade 2 when going from a kindergarten setting, where there were no real performance or attention expectations, to Grade 2. The teacher noted I was having trouble paying attention. They thought maybe I had hearing issues, it was a private school and so they brought in an audiologist. My hearing was fantastic! So indeed, unfortunately, 30+ years ago especially they didn't consider that my hearing could be "too good" - where sound was overwhelming me; so I was hypersensitive to sound, arguably hyposensitive to touch and other senses, and it was medications in my late teens and early 20s that caused my hearing to get super hypersensitive - to the point where I was in what I consider a torture state for 8 months - where even the sound of blinking was painful, or at least that was the sensory I was associating the pain with - until I was forced to do my own research and eventually stumbled into Berard AIT.
There's also pre-care questionnaires that some practitioners offer - a checklist for behaviour of a child, and also of an adult, where I had ~80% of the behaviours both as a child and as an adult, e.g. preferred to sit in the back or corner of a classroom, essentially as there'd be less directions noise would be coming from, had trouble relaying a story or following instructions, etc.
Adult checklist: https://www.aithelps.com/auditory_care_adults.php
Child checklist: https://www.aithelps.com/auditory_care.php
NOTE: these places started offering "home programs" to send some audio equipment and the specially modified music, however I won't personally trust them until I've had the chance to do what I'd consider to be thoroughly done research to compare the original high-quality sound equipment that would produce the absolute highest quality of frequencies vs. what the rentable-shippable equipment produces.
Someday I need to write a chapter of a book or perhaps a whole book on my experiences of it all - the before and after state, the blocked development process (e.g. autistic state) that got unblocked and the development process that began to unlock - essentially I was blocked from processing emotions properly, so I arguably had a lifelong backlog of unprocessed memories with emotions associated with them needing to be labeled to be organized, as well as PTSD from many very intense traumatic years.
> I don't get why people go there and leave money, cause I wouldn't go there even for that money.
That's a problem with lack of empathy and understanding on your part, not with the group dynanimcs of that person's friends.
They're not perfect, but the fact that I can move smoothly from "I need my ears open but still want to hear my headset" to "I need to block out sound and hear my headset perfectly clearly" with just a finger or a pair of earplugs is great.
Stick a shotgun mic (that's the term for a mic with tight directional cone, right? Not an audio guy) on the side and this would be really cool.
Every time I looked into this, it seemed to push the link with autism and/or adhd - back in 2008 I wasn't diagnosed so I poo-pooed the idea somewhat. Now I'm diagnosed as AuDHD.
1. Don't speak fast. Speak slow. Enunciate and articulate all the consonants. And do it very slowly. Give the vowels lots of room to be noticed.
2. Don't speak lightly.
3. Don't mumble.
Aayyeee hHHaaaVVE TTOOO GOOO NNAAAoooUUU
That's so you can lean in and get a little bit more friendly. Or go out for some fresh air together.
I thought I was the only one with this problem! Someone would make a joke and I would have to pretend to laugh because I didn't even catch what he/she was saying and asking them to repeat it the 3rd time would be awkward.
Even worse, it isn't always a joke, so even calibrating the laughter level is hard. Ugh.
Other person: *mumble mumble* SOMETHING CLEARLY SPOKEN
Me: I'm sorry, what?
Them: "clearly spoken?"
Me: No, the first part.
Them: "something?"
Me, giving up: *smiles and nods* Yeah!
(quietly hopes I didn't just agree that putting hamsters in blenders or something is a cool idea)
My Mother has had poor hearing for decades. She listened to a radio as a kid she held it next to her ear at a loud volume. Now she often says she "can't stand noise" but it's because she can't hear in loud environments anymore due to hear hearing problems. I've noticed she misses the start of a sentences like if I said "I'm going to get some milk" she heard "got some milk" (as in I just got it). Or she interrupts people because she can't hear the first part of someone starting to speak most people tend to speak low at the start of a sentence.
My family will often have the TV on, games playing on phones, and talking too - I just can't hack it.
Equally the option to work from home has revolutionised my productivity - without having 10 things to filter out, I can just focus on getting the task in hand done. In an office I often get lost and distracted, and have to power through the noise.
https://www.mayoclinic.org/diseases-conditions/auditory-proc...
A good hearing aid person or an audiologist can diagnose it. I found out that I likely have it, which explained alot of things in my life that I had experienced. In some scenarios, hearing aids can help, even though you don’t have a hearing problem per se.
Yes, the stem sticks out, and everyone asks about them at least once, but I usually say something about wanting to "prevent tinnitus" and that it helps me hear them even better and they usually don't ask about it again.
This is one of those frustrating gaslighting things that is half true in that half the time I also pretend to hear what someone else is saying even though I couldn’t just because it’s not really important and making a big deal about it (ie asking them to repeat it at continuously louder decibels) can get awkward.
I recently had a test at an ENT doctor who told me my hearing is fine and insinuated I was wasting his time. The test was listening to high-pitched beeps over white noise, which isn't representative of the problem. Distinguishing one particular tone over several similar ones would be more like it.
I wrote about my experience with this last year: https://news.ycombinator.com/item?id=35897515
I did exactly the type of diagnosis you're talking about. It was quite good at how it simulated a noisy environment with a bunch of background chatter and then a single voice you were meant to listen to that would repeat various patterns of words with various combinations of lower speaking volume and/or higher background noise.
One thing I wish I'd made a point of at the time was the fact that, despite being an apparently soundproof booth with headphones on, I could definitely hear people talking in the waiting room and another audiologist in an adjacent room. Though I'm not sure it would have materially changed their lack of diagnosis (they'd already detected I could hear into negative decibels).
I still don't have a diagnosis, but I'm increasingly coming around to the idea that maybe it's not that my hearing is bad but that I actually hear too much. What I'd previously thought was my unability to hear people speaking on the radio in the car when everyone else clearly could wasn't because I couldn't hear the radio, it's that I can't hear it over the top of all the tyre and wind noise I'm also hearing and trying to process out. I don't think the other passengers in the car hear the rest of the noise, they only hear the radio.
I bought various types of Loop earplugs and have found them fantastic for live music events. I can now hear my friends when they're talking to me! Unfortunately they greatly amplify my perception of the volume of my own voice when I talk which has the undesirable side-effect of making me talk even quieter so I feel like I'm having to yell when I want to talk to people. I've also not found them as useful as I'd hoped in restaurant-type settings.
go and look the up the price, they are deeply expensive, even for basic "make it louder" type aids.
Worse still, because they interfere with your ear, you tend to loose the ability to "steer" your hearing. This means that you can't tune out other conversations/noises or stuff.
The one good side effect of facebook spending billions on its (probably) futile search for practical and popular AR is https://www.projectaria.com/glasses/
Which is a (cheap) platform to do experimentation for AR type actions.
However it has eye tracking, microphone array and front facing cameras, so it can be fairly easily modified into being a steerable microphone.
My Phonaks have the ability to automatically switch programs to some extent, and to fine tune the program using a companion app.
I can function and even have conversations to some extent in noisy environments, something that would have been impossible for me with hearing aids from a decade or so ago. I'm very grateful for this of course.
The pair costs roughly $2000. Luckily, it's covered by the national healthcare system[0] (which I of course pay for through my taxes) so I end up paying $50 every five years or so.
I hope the advances in "AI" will make it possible to not just amplify and filter (even if it's in very clever ways) but to isolate and enhance/reconstruct voices in noisy environments.
Meanwhile, I hope the trend of playing louder and louder music in cafes, restaurants and bars dies out. It's an accessibility nightmare, especially (but not only) for people with hearing loss.[1]
0: https://www.socialstyrelsen.se/en/about-us/healthcare-for-vi...
1: https://www.vox.com/2018/4/18/17168504/restaurants-noise-lev...
I don't hold out much hope because as far as I can tell it's done to make everyone "shut up and drink". I could believe it adds at least 50% to sales because when you can't hear a word anyone says, all you can do is smile and nod and take a swig. And if the place is already full anyway, they don't care if you leave, you'll be replaced. What matters is whoever is in there taking up a space is drinking as fast as possible.
Of course, people who get substantially drunk (which is to say, customers who spend) also don't care because they're not really listening closely or making conversational sense anyway and their pain tolerance is way up, so it's just a good time to them.
Even more cynically, it also keeps the place "cool" because all the old, past-it fogeys like me don't even bother going in. From this sample of one, someone who thinks the music is too loud is un-hip, isn't adding to any hookup appeal (either not in the market or pushing the creepy end of the age range) and won't even spend much because they can't get really hammered because the hangover will take them out for two days and they can't afford to lose a weekend.
I don’t see any reason proper hearing aids can’t already be doing it now though I am sure some of them are but probably the even more ridiculously priced $8k+ models.
Nuheara is also in this space but marketed and designed more specifically to be a low budget hearing aid replacement. With a similar pride to AirPods Pro.
I thought it's well established that they're doing it entirely on purpose. Restaurants, cafes and bars don't make money on you chilling out and having a good time with friends; they make money on you buying food and drinks. They want you to order and consume ASAP until you're full and leave, freeing the table for the next group of customers. Loud music that prevents you from having conversations is how they make this happen.
I can't make out any conversation in a noisy environment so usually switch to try-to-filter-noise-and-fail plus some amateur form of lip reading which works ok for a casual conversation but not for a more serious one. Hearing is "ok" enough though when testing, so no clue what it is.
It helps a lot when the ambient noise level reduced by a few db and tuning down the music helps a lot.
It's funny how for some problems the path to the solution is blocked by deadlock
My current pair are about 6 years old, working fine still, thankfully... But in a recent visit to audiologist, they had me test out a newer pair... but they had a single button instead of rocker + button, and bluetooth/app was touted.
I dread the touchscreen phase of HA as a young person with functional fingers (vs elders with dexterity issues) and a preference for physical buttons (a la my 2009 car).
The idea of autoswitching the programs outside limited cases (direct audio input cables and increasingly-rare telecoil situations are the only things I would accept) also doesn't sound great! :)
Shes older so that might be part of the technical challenge with them but i would have expected better given the huge price tag. Feels a bit like a monopoly running the development but that is merely a hunch.
I often see AI hopes expressed in this format. I would put it another way:
> I hope the advances in "AI" will make it possible to restore hearing to baseline average human level
Wishful thinking would be to enhance it beyond baseline. It's perfectly reasonable to think AI-advances can help researchers restore hearing in most cases, and reasonably within 10 years or so.
(0) https://www.thebignewsletter.com/p/silencing-the-competition...
Otherwise the stuff you described in your comment around attention filtering starts to happen because of the sensory loss. Therefore the longer you avoid hearing correction once your hearing starts to become impaired, the more complex and expensive a hearing aid you need to re-do. This is because these expensive hearing aids do a poor approximation of the things your brain/auditory cortex was doing prior to the sensory loss.
BRB - better go book a hearing test.
I guess they're expensive because of relations with medical / health companies being complaint makes things expensive (eg. the same display but with certification to use in a medical facility would cost many times more).
The device is relatively simple to make so I asked my teacher why were they so expensive. He said that yeah, the engineering/manufacturing side of it is about 200 EUR, the remaining 9.8k EUR is spent on certifications/paperwork.
Obviously, wages factor into this but over time I've come to see how paperwork and paying lawyers do in fact account for the majority of the cost.
Expensive, yes. My hearing aids cost ~$2500CAD each but "how shit hearing aids are" is not my experience at all. My hearing aids (Widex) are awesome! The quality of audio in normal situations is fantastic. My only real complaint is that they're not completely waterproof so I have to plan ahead a bit if I'm going to be outdoors in the rain.
Do people genuinely have that ability, to listen to a specific person and ignore the rest?
Asking because no matter how hard I try I can never understand a fucking thing if there's many people talking loudly in the background, it all mixes together into an incoherent whole.
https://en.m.wikipedia.org/wiki/Selective_auditory_attention
Indeed! however like visual depth perception, not everyone has it. The human ear has a load of bits that allow removing noise from the signal. (I don't have a block diagram, sorry!)
In theory one should be able to locate a noise in 3D. You can test this by getting someone to hide your phone and then ring it. if you have 3d sound perception you should be able to work out if the phone is behind/front/up/down.That forms part of the "steering" ability.
Then there is filtering the noises that you don't want. Music is can be a good test for this, how many instruments played on this track, what instruments were they, what were the lyrics, etc. Being able to do this requires that you be able to filter out noises that you don't like.
Again like all human senses, there are levels of ability, and in some cases can be improved with "exercise"
But, hearing what people are saying against a loud background is really really hard, so don't worry too much. Plus voices have specific human social encoding, so they can be affected disproportionally
I don't think anyone would suggest them as a realistic choice today but I could see Apple going after that market and where Apple goes others follow.
Much like the market for prescription reading glasses has been eroded by off the self glasses.
I wonder if some constructors could target this use-case by making aids that are very easy to self-calibrate.
On the other hand, the median hearing aid user is pretty old, has lots of disposable money and has never watched a youtube tutorial in their life, so it might be a small market.
Not to mention, it probably takes a couple of years to get the certification for a device. So, any device to market is easily 2-3 years old tech by definition.
You have two microphones already, spaced about 30 cm apart...
but you don't need eye tracking all the time, as most you can latch on to the location of the speaker without looking at them continuously.
The possibilities are rather good, but it needs someone willing to fund the research
Like, $50,000. I'm hoping that removing the need for prescription ones, will allow the price to go down significantly.
As an older person, I have noticed that my hearing has gotten "louder," over the years.
I still hear dB levels fine. The problem is that I hear all the noise. I used to be able to hold conversations in loud environments (like bar/restaurants), being able to hear the other person, despite the background noise.
Not that long ago, I was at dinner in a noisy restaurant. I was sitting directly across a narrow table from someone (about 30 inches -max).
I couldn't hear a word they said. They could hear me fine (they were younger).
If this works out, it might give the folks currently collecting $50K a pop, another way to charge eye-watering money.
I've noticed Americans like to bs their medical costs as bad as the system is, you can't compare some newest luxury devices to what an average person is using all around the world.
Looks like I’m wrong, making a general statement, based on anecdotal information.
We’ll have to see what the future holds for us.
Apple was working taking the platform even further at one point and I would not be surprised if we see some new announcements eventually.
Imagine if a $200 set of airpod pros outperformed top hearing aids.
https://www.apple.com/newsroom/2024/05/apple-hearing-study-s...
If you can pick out audio from individuals, you could also send it through speech recognition and subtitle real world conversations for when hearing is worse or not there at all.
But it really needs mobile devices capable of doing the processing locally, as I think round trips to the cloud would make it less useful or potentially useless.
Still, the transcription part is already here today. The Google Translate app has a transcribe app that does this (runs locally; does not do magic AI "pick voice out from crowd"). My father-in-law has been using it for years. When I'm in a loud environment, the app I use on iOS is called Big, which just displays large text on the screen.
I'm in one of the biggest cities of my current country, and the RTT to google from me is 87-91ms. Well over 4 million people live within 100km of me, so I suspect they see similar latencies. On my cell, I see 191-207ms.
I would think this shouldn't be a problem as the correct hardware gets adopted in phones. As it stands now, you could probably run it on a Coral USB accelerator and battery run Pi (just an example of hardware, obviously we don't have the code).
Or for when you don't speak that language.
And given the US healthcare system, somebody is gonna take all our money too, one way or another. :P
Has a problem that I think the AI headphones wouldn't solve either: in a (non-quiet) group setting you still need to anticipate who's going to speak when and look at them for best results.
The direction bit is just biasing to preferring forward stuff (via two mics on each ear's HA).
Sadly, no backwards bias option for overhearing people behind you ;)
Put them on "backwards"; left cup on right ear and vice versa: forward facing mics now face backwards ;)
>> https://www.cnet.com/health/medical/what-did-you-say-these-e...
So something that would enhance the speech of whomever I'm looking at would be super cool. Apple AirPods already have some sound shaping abilities to react to environment and mode. They also support specific voice enhancement if you put your phone down in front of the person speaking. If they ever support directional voice enhancement, like in this research, directly in the AirPods it would help me so much with social interactions in loud places.
It's common to wear ear plugs at concerts, to avoid destroying your ears. Not imagine replacing those ear plugs with in-ear headphones that filter everything except your family/friends and the concert, while regulating the volume (if your SO talks to you make that one voice loud over the rest), maybe keep the "crowd noise" going with the flow but remove normal conversations for people around, etc ...
I'm not in ML/AI/etc ... At all but my understanding is that none of that is actually impossible with current tech ? Sure the battery and power limits exists, but this is a concert those headphone with a "band" going behind your head to keep them in place / not lose them if it falls makes sense. Would need some training for "your voice" but if alexa can do it in 10 seconds then a phone app can do that too.
Hell, if it existed for movies theater at below 200€ I would probably buy one right now and maybe go to the movies again.
Over 4 years ago nvidia released a feature that lets you remove arbitrary background noise in real-time.
Here's a video where a guy put a fan, vacuum cleaner and leaf blower right next to his microphone: https://youtu.be/Q-mETIjcIV0?t=535
It definitely chopped out a bunch of his natural frequency but it was clear enough to hear him without issues. Earlier in the video he did more normal tests like removing the sound of his keyboard in which case his voice's frequencies were mostly left untouched. He also banged a hammer on his desk while talking.
I don't sound fantastic on it, but i sound better than people using cellphone microphones and thrift store microphones. It also works if i talk loudly from another room, but won't pick up normal volume conversations in the same room, which means there's a noise gate in there, too.
I've heard a very abrasive sneeze sounds like "chew!", like a cartoon sneeze or something. I couldn't tell the difference in a blind test between a cellphone's noise cancelling with the sound recorder and the asus device vis a vis overall quality, but the gating on the asus is more aggressive. It also works better than the default discord noise reduction, but is about equal to the Krisp (iirc) implementation. Its gate is faster than discord if you have both krisp and the normal noise cancelling on.
I think they're discontinued. If i ever see one in the wild i'll be sure and buy it. I have never tried it with a decent microphone - and i do have a couple, including shure and marantz - because there's no need. I wouldn't use it for podcasting or doing anything where the overall quality would be noticed; but for discord / in game / PC telephony it works great.
My understanding is that to "mute" a sound, you need to inject another wave that is exactly the opposite, with the exact same volume and in perfect sync, so that the two waves interfere destructively. However, in general but especially in AI, you can never guarantee 100% accuracy. If you use this technology to "silence" a background fountain, and something goes wrong, at worst you get a lot of noise that make you grimace and remove them. If at a concert with 100+ dB of music you get an error and your headphones start producing a similarly loud, but not perfectly aligned noise right into your ears, you probably won't have the time to remove them before damaging your hearing system.
In general, I think that having a tool that drives 100+ dB straight into your head is probably not a wise idea :-)
Like others in this thread are saying, though, it would be an incredible boon for people who use hearing aids.
During the first aborted product effort to develop headphones, we were looking at a conceptual feature similar to this - selectively allowing people’s voices through the ANC chipset.
I don’t recall the exact approach the DSP folks were using (I was closer to the hardware for ANC) but they were really only able to figure out how to isolate the wearer’s voice by virtue of that signal having more power than all the others.
This is terribly cool. I wonder what other kinds of fun you could have with headphones. ANC chipsets are incredibly powerful and I’d wager their capabilities are not even close to fully tapped.
Seems like noise cancelling has been solved for the listener (isolation + ANC) but I would sure love a hardware/software combo to come along and allow me to work truly remotely by blocking out noise/isolating my voice to the recipient.
Apparently the microphone has noise cancelling [1], and it seems to work pretty well. Nowadays at home, I've been in calls where my wife is using the blender at high speed, I get annoyed by it, but none of my colleagues hear it (between the headset NC and Google Meet's NC apparently it works pretty well)
[1] https://hyperx.com/products/hyperx-cloud-iii-wired-gaming-he... Crystal-Clear 10mm microphone, noise-cancelling, with LED mic-mute indicator
Unfortunately, they don’t do active noise canceling for output sound, but their mic is incredible at canceling out noise. I’m not sure how they do it but have heard there is a contact mic in the bridge. Just got mine, have gotten positive feedback from others when making calls in busy environments.
The "technical" solution is to pull the plug and reboot it (which you can't even do remotely, even if it's connected to Wifi and you want to reboot because Spotify connect on Sonos can be buggy as hell).
I can keep a wifi connection up myself and always reconnect using an esp or similar TI etc module...is it so hard for the Sonos firmware devs to do something so basic?
Does that mean sound coming in the same direction as the person they are targeting? Does it mean sound near the person /object they are targeting?
So perhaps this is not as out of reach as many pop-science articles. I’d love to hear if anyone is able to get this working independently.
The only case ANC would block your partners voice would be if it is about as high/low level as the background noise/sound so that it is all mixed into a white or colored noise which ANC can suppress.
This is one of the safest technologies I can imagine.
It's more likely that the radio waves from wireless communication (phones, Bluetooth headphones etc) will have negative impact, but even that's unlikely at this point, considering how widespread their use is and no statistically significant link exists.
TBH at this point, I wouldn't even object to losing my hearing to have forever ANC (hearing loss) and turning up the hearing aid.
E: no offense to those with hearing loss in this thread
I use a lot of curse words. ;)
I’m pretty sure a human guitar teacher would have not problem with helping me with that.
[1] https://cdn.openai.com/spec/model-spec-2024-05-08.html#excep...
EDIT: I got a translation but it lectured me first (and I'm not convinced the translation is accurate) https://chatgpt.com/share/10c63fad-4716-4cde-885d-a681c7cb78...
Prior to what i would describe as idiotic practices in my 20s and 30s, i could focus rapt attention on as few or as many people as required in nearly any situation, parties, bars, whatever. Music venues obviously not, but looking at someone and talking loudly near their ear (vice versa) worked fine.
My hearing issues are similar to my FIL's, who is 50% deaf in one ear and 95%+ in the other. if you're sittin on the wrong side, you're getting a lot of smiles and nods, because "eh?" gets old. real. fast. However, i can hear a raccoon messing around outside, and no one else seems to hear it; also a phone notification going off in the next room, as examples. my hearing "feels" fine, except in the very specific circumstance of human speech recognition. I also have tinnitus sometimes - it comes and goes, and if i concentrate i can focus it to the point of wincing, sometimes.
interestingly i have to turn down the master volume of games when i am voice chatting. any amount of noise from the game will interfere with my ability to process people speaking, especially if the game has a lot of talking. Additionally, i cannot watch any tv or film produced after about 2000 or so without subtitles, regardless of how fancy the "5.1 surround" center channel DSP is.
sorry for rambling, i suppose i don't think ADHD has anything to do with my hearing loss or issue; I should have worn hearing protection more when i was younger and i'm not regretting it much now, but when i can't hear music i enjoy i'll be pretty sad about that.
I wouldn't be so quick to dismiss that as a potential factor, auditory processing issues are known to something that can be connected to ADHD in some people. (Also mentioned by a few other people in the comments.)
Couple of links that go into more details:
* https://en.wikipedia.org/wiki/Selective_auditory_attention
* https://en.wikipedia.org/wiki/Auditory_processing_disorder
* https://en.wikipedia.org/wiki/Cocktail_party_effect
> Additionally, i cannot watch any tv or film produced after about 2000 or so without subtitles, regardless of how fancy the "5.1 surround" center channel DSP is.
IMO that's potentially more because audio engineers for home movie releases are... ok, well, let's just say, "have different opinions on how to mix audio with dialogue than me". :) It's a very common compliant.
> I should have worn hearing protection more when i was younger
Shouldn't we all? :) Still worth trying to protect what you have now--I've used the "Etymotic" brand ear plugs mentioned elsewhere in the comments which are intended to more evenly reduce sound levels without just "muffling" everything.
> but when i can't hear music i enjoy i'll be pretty sad about that.
Indeed. Understandably.
This looks like it could do just that with the headphones feeding directly into the mixer and behaving like a focused mic.
I wonder if the problem maps easily from "select this source" to "select everything but that source"
> The headphones send that signal to an on-board embedded computer, where the team’s machine learning software learns the desired speaker’s vocal patterns
Their "AI" is good ol dumb machine learning
If you have good eyetracking, a microphone array and decent object tracking on your AR glasses, then you don't really need much "AI" (ie you have access to https://facebookresearch.github.io/projectaria_tools/docs/AR...)
but its not quite possible to do it all on device yet. However its not far off.
Edit: I found the link to the paper. It isn't stream splitting so much as it is GPT-assisted beamforming estimation. Good stuff for sure.
But this is super expensive since you need calibrated mics etc.
The biggest advantage of neural nets in this field is that you can use a dirt cheap microphone and postprocess it so good that it is good enough or even very good for humans.
I remember learning about it in the early 2000s. It was considered a very challenging problem, with very important applications, most notably speech recognition in natural settings.
I wonder what is the current status on this. Is this considered solved nowadays?
However, I think in the next 50 years, headphones will disappear or.. should I say evolve as part of the human anatomy. Same thing for screen monitors, mouse/keyboard, smartphones, etc.
Think about it. The way things are going, along with "AI" (sure buzzword in a number of ways but something that will change our way of living) many of things we use will be replaced and, likely, be simple extensions or, dare I say, be implanted.
Hard drives will be a thing of the past. Everything will be (as we call today) "cloud-based" and we will be more cybernetic than we think. Of course, someone today will fear such an idea. As we slowly accept the little changes.. in 50 years we will look back and think "how did they cope without it"... a bit like how someone today look back and think "how did they cope without the internet"
Many fear what they dont understand. It is the unknown. AI is a fear factor for many. For me, I accept it for what it is and the changes it will impact our lives and our careers.
All I will say is -- strap yourselves in.. it will be a bumpy ride. I hope we make it through without destroying ourselves. Once we past it, the world could (finally) be at peace and to quote a famous TV show --- "to bondly go where no man has gone before!"
no, i don't, because of your kenny G knockoff playing at low volume.
Researchers are getting closer. Dr. Chen from Harvard was able to regenerate hair cells in mature mice last year.
The problem is also becoming more widespread. 30 Mio people in the US and 400 Mio people worldwide have disabling hearing loss. Regenerating hair cells and the synapses around them would also cure Tinnitus. 30 Mio x $5k for a treatment = $60B market (probably way bigger with aging population)
I think we probably need more rich tech billionaires to get affected to attract large funding.
What billionaires that you know are affected besides:
- Brad Jacobs
- Ryan from Flexport/Founders Fund
And make sure if you're going to do bluetooth + wireless, remember that both bluetooth and wifi transmit on 2.4 GHz, and need to coordinate in order to coexist in the same IoT device. There are interconnects and wire protocols to connect the bluetooth and wifi chips together - or, preferably, you buy a chip that does both.
There's an episode of the sci-fi show Black Mirror (White Christmas), where a person is convicted of some hideous crime and permanently blocked/made invisible and inaudible from everyone (the entire population has embedded audio/video processing enhancements by then).
You can imagine future headphones where you could block out the guy in your office with the annoying laugh our download 'blocks' from the headphone appstore - no more Rick Astley or the politician you don't like etc.
Perhaps the accuracy of identifying the correct voice could be vastly increased by adding video input. The AI can then try to match the various voices with the lip movements in the center of the video, basically lip reading.
But this is very useful for people like me who don't hear well in the high frequencies.
Oh, so it barely works and its a proof of concept.
What is the interesting thing here? We all know how sound waves work. Pretty sure this technology is old. Until there is a product here, it just sounds like you are rehashing noise cancellation.
Academia has dug this grave of skepticism. I just have 0 faith this will get to market through University of Washington. Maybe it will be patented and used even less!
I see a Python script I can run on my computer, I haven't tried it yet, but I think I could connect a microphone and process real-time audio and output it in real time, but I don’t know how to detect the user looking at someone. Could you tell me how that works?
I guess by reading the article?
Basically, her hearing is perfect but her brain struggles to process sound in a noisy environment; she can't single out what she is listening to.
This sounds perfect for her!
* https://en.wikipedia.org/wiki/Selective_auditory_attention
* https://en.wikipedia.org/wiki/Auditory_processing_disorder
Actually from the description it sounds (no pun intended!) like you wouldn’t even have to look at them. You just would need to be facing their direction. You could be looking at something else in the direction, like the ground in front of you.
When you tell it to start listening to the person you are looking at what it really does is start listening to the person whose sound is coming from that direction, which it figures out from arrival times at both ears.
The only behavior close to that I can think of is when someone is looking at someone expectantly, and trying to break in.
I realize this is slightly tangential, but please don't replace customer support with chatbots or whatever you want to call them. It's a freaking horrible experience.
Turn anything into a mirror, or something like that.
With this you look at someone, signal that you want to keep listening to them, and the AI learns their voice. It then lets that voice through the noise cancelling system even as you or they move around the room or you look elsewhere.
I know this is just the beginning and the tech and UX will mature a lot - but being able to consciously choose what we allow into our sensory world would be a great superpower to have.
In the distant future this will all be embedded inside a cochlear (neural?) implant.
You can "save" known voices, prioritize them, identify various scenes/modes automatically like meetings/parties/concerts/driving/walking etc, know when to allow external sounds in (alarms, honks, someone calling your attention, etc)
And with great power yada yada.
I can already imagine a few ways this can be misused / abused / create non-existent challenges and problems too. But I am (cautiously) optimistic that we as human race will collectively figure out how to steer these new technology applications into net positive territory.
2040: iAudio and xSmell blamed for people losing connect with nature's sounds (like bird chirps and flowing streams) and smells (petrichor) - things that inspire us, make us creative, make life worthwile, and make us humans.