Enhance Speech from Adobe – Free AI filter for cleaning up spoken audio
podcast.adobe.com
podcast.adobe.com
[0, original] https://youtu.be/gwkCIdwHRhc
[1, enhanced] https://youtu.be/RPnUqmSyZ6Q
Also, very enjoyable & clear presentation.
This would be wilder if I were in an unconstrained environment with similar enhanced audio quality. It would seem like a voice-over.
Is there any way one could easily train it on my own voice to make sure it isolates my or any other trained voice from noisy environments?
However, when I ripped that video's audio and put it through the linked AI filter (Adobe Podcast's Enhance Speech), it went from being unpleasant-but-understandable to perfectly clear gibberish [1]. For example, about 15 seconds in I said:
> [...] at Second Street, the company whose office you are currently sitting in. Uh, you can find me on the web on GitHub, on Twitter, or my own site at Kerrick Long Dot Com.
In the "enhanced" version, it sounds like a person who isn't me said the following in soundproofed studio (with a head cold):
> [...] at Second Street, the cubbany od office you are currently city-aly. You gidhi me on the web on GitDub, on Twitter, or my own side at Kerri Flow Dot Cowf.
[0]: https://youtube.com/watch?v=wc72cyYt8-c
[1]: https://soundcloud.com/kerricklong/javascript-promises-think...
https://youtu.be/1LDlOmKtfeQ?t=60
This is a 20 yr old vid shot at Pearl Harbor on the deck of USS Missouri with a low-end consumer Hi-8 camera using built-in mics. Notice flags and clothes rippling in the wind. This filter works pretty well in this case. Maybe it was trained on a New York accent? Mixing in the enhanced track as needed works well. Using only the enhanced track may sound artificial at times.
Here's another vid shot at a brewery where this filter helped clarify the brewmaster's voice over a noisy restaurant and outdoor machinery (again, built-in mic):
I found this filter useful.
Funnily, in order to use this sort of model for tasks involving speech recognition it's often recommended in the literature to mix back in some of the original noisy audio. This reduces the impact of artifacts introduced by the enhancement which would otherwise reduce ASR quality due to domain shift in the data.
Guess humans and computers have similar needs in this case. :)
These are impressive results, the audio mostly sounds like you gave the guy a lapel mic. :P
- If the microphone quality itself is bad, the enhanced audio is still pretty horrible. - It does clean up echo but with there's some pretty aggressive EQ that doesn't sound nice, and the noise gate is pretty severe - Compared to my XLR shotgun, the quality of the phone was pretty horrible
What we can conclude is that if you already have a good recording but with some problems, you might be able to use this to remove those problems. However, don't expect a crappy microphone to turn into a good microphone, or a crappy recording to turn into "studio quality".
The bottom line is that there's no substitute for a decent microphone in a decent space. (At the very minimum, small room without echo.)
Both microphones need to be cleaned of any blockages so the hardware echo and noise cancellation on a given phone works well. Otherwise you've got distorted audio getting processed as if it's not distorted...
Look at all the money that's been going into making phones compelling substitutes for professional cameras. And camera tech is still evolving. Consumer audio devices will get the same attention and investment.
We'll eventually have models for audio signals in all sorts of distorted and noisy environments. I'd bet that in ten years a cellphone microphone can duplicate a professional audio setup in 90% of circumstances.
It's the same as with photos. If your raw material is bad, no tool on earth can make it good.
Why small room? Does that reduce echo?
In theory, would a gigantic room that was miles to the nearest wall be even better?
(and of the rain too)
Yes, as effectively it's open space. In practice though, to record in high quality you would rather build an anechoic environment, as small as possible (preferably a booth).
This is where the Startup Garage analogous cliche for musicians comes from: recorded in the closet
You need to absorb the sound of your voice so there is less echo by baffling material on the back of the mic, and absorb room tone and echoes in the area the directional microphone is pointed, generally behind your head. Small spaces have only disadvantages as studios.
To take advantage of a reach-in closet full of clothes, put some pillows on the shelf over the clothes, take the closet doors off, and back into the closet as much as you can. In this way the microphone is primarily listening to the baffled sound inside the closet, and you can avoid bass pooling by speaking into the room—ideally with baffling material (e.g. see http://PillowFortStudios.com/ ) ON the back of a LDC microphone.
[0] https://en.wikipedia.org/wiki/I_Am_Sitting_in_a_Room [1] https://www.youtube.com/watch?v=fAxHlLK3Oyk
I would definitely say this is "similar quality" to adobe's. Nice work.
I’ve checked several similar services and it really frustrates me that all of them price by the hour the bill is a monthly plan, but in actual fact the real cost is $/hour of processed data. The these plans inevitably end up as a series of x hours per month + some features, and all features of the previous plan teir… with many companies using larger hour requirements as a way to force you to pay more per month.
Also why the hell is it so hard to find the equivalent of this technology but for realtime as in streaming my microphone audio live … i would happily pay $100 -> $250 for a AudioUnit plugin (or other equivalent audio pipeline plugin formats) for this kind of real-time voice cleanup… but it doesn’t exist. So can you help explain why? Since you’ve built something like this I’m hoping you have more insight into why it’s harder for real-time processing…
Results:
- background noise was reduced
- some previously clear words are turned into garbled non-words
- some parts are replaced by a different Indian voice, I assume AI, so it sounds like multiple people talking
All in all, the results are not anywhere near what the sample shows.
https://www.theverge.com/2013/8/6/4594482/xerox-copiers-rand...
Bad news: While the 5% was a minor inconvenience for customers, the 1% is bad enough to end your company
Just curious how the performance differs between PCMU @ 8khz compared to Opus @ 48k or IMBE and AMBE+2 (Project 25 Public Safety audio codecs) :D
My dream would be doing audio processing in real time to clean up the audio of phone calls
The tool that now comes with the latest updates to FCPX handled them without a problem. (Still some background noise, but you can clearly hear every word.) I think Adobe has a long way to go on this.
Gobsmack all around. Mine, that it did better in 10 seconds than I think nearly anyone could have done in 10 hours. Theirs, that they thought I performed a literal miracle.
Life comes at you pretty fast sometimes.
Many other professors (as well as me, occasionally) also use to involuntarely (and often unaware of that) say something like eeeeeh when they strugle to recall the right word. Would be great if this could be removed as well.
I'm not sure I understand what you're saying. Do you mean that the talk was in Japanese but the output of this service somehow screwed it up in a way that it sounds like spanish?
Or are you just mentioning that the thing you uploaded was spanish and not describing the quality of the output?
Also would be great if used in Zoom recordings of podcasts
However, a few seconds into my recording there is a part where there is someone else's dialogue for a few seconds. I can't make out what they're saying but it definitely sounds like a man, with a Latin-American accent, speaking English for a second.
Could that be a hallucination or somehow they mixed audio from another recording? It only last for about a second, but it's so strange.
It will really get freaky when there an ambient noise resembling a human voice. I'm thinking the Bear scene from the movie Annihilation.
It can be a useful technique for learning how slightly different prompts affect things
It seems Adobe never advanced Voco past the research project stage ([1] via [2]). I'm guessing they had trouble getting it to work reliably on a wide-enough range of real-world audio.
[1] https://community.adobe.com/t5/audition-discussions/beta-tes...
[0] https://fsi-languages.yojik.eu/languages/FSI/fsi-french-basi...
I feel like we’re real close
People do this regularly today using Descript. https://www.descript.com/overdub
Through answering this I also found Respeecher, which seems interesting too. https://www.respeecher.com/
If you'd like to run something locally, there's also https://www.nvidia.com/en-us/geforce/guides/nvidia-rtx-voice....