Voice Isolator: Strip background noise for film, podcast, interview production
elevenlabs.io
elevenlabs.io
I have a recording I've been sitting on for 2 years(a guest lecture which a friend recorded) which contains a very heavy amount of background noise, where you can just barely make out what is being said by the lecturer. I wonder if there is any hope I will ever be able to read a transcript from it.
I can figure out what the lecturer is saying (maybe only because I have some context about what he is talking about), but it is too painful to sit through 2 hours of it and try to transcribe it.
I tried uploading the audio file to this service, but basically get nothing useful returned to me.
Disclosure: I work for Deepgram
I get the appeal of automating this task but the SOTA is not to automate it at all.
by the way, even once we get to a sufficient AI, how do you verify it without listening to the whole thing anyway?
it's only 2 hours, if you're a fast typer at max it would take you 1 work day to transcribe yourself, or <$200 by a professional
Do you know if there are any license issues with this? I don't see any license page--will they train/retain the recording?
TTS: GPTSOVITS / StyleTTS2
VTV: RVCv2
Open source isn't really doing a great job at voice, music, or video. It's managing to keep up in LLM and image spaces, but it's falling far behind in the multimedia department.
I just started trying this today with my wife's hour long interviews in Latvian language and it is extremely good, far better than any transcription model. This is a huge SoTA LLM with audio tokens, so it just has vastly more capability than Whisper or whatever. In my case it nails all kinds of brand names, weird neologisms and loan words, it writes inline quotation marks when the speaker is quoting someone, and so on.
This is what GPT-4o supposedly can do in the version OpenAI has postponed rolling out.
If you want, I can try to do it for you if you send the audio.
Set up an audio source, for example, your phone, playing a reasonable length of talking, for example a youtube video, or a podcast on spotify. Then record from your computer or other recording device, and test with that?
If such process can output a clear sound, I could chain it with Blackhole and have it and use the processed clear signal as an input for the call.
To do it right you really want digital analysis.
It does take some technical elbow grease to integrate but I've used it in calls and while gaming on Linux via Pipewire to great effect.
I suspect that this is noise cancellation that's failing because they keep their phone far away from themselves, to fit two people in the shot; and audio is bouncing off the walls or otherwise suffering enough delay to mess it up.
I have old mono records that I wanted to clean up. In that case, any stereo content is obviously scratches and surface noise, so removing it would be most of the job. But nope... not one DAW offered this filter, despite offering the opposite (removing mono content and keeping the stereo).
And yes I did try removing the mono content and then subtracting the result from the full source, but this didn't work; I don't remember (or know) why.
Second on the list - https://www.musicradar.com/news/the-best-daws-the-best-music...
Fifth on the list - https://producerhive.com/buyer-guides/daw/best-daws/
Second on the list - https://mixingmonster.com/best-daws/
It's pretty popular and I'm sure it has the most tutorials on youtube
Judy Garland, Burt Reynolds, Laurence Olivier, and James Dean are the first ones.
You used to be able to pull out your phone and play Disney soundtracks or Taylor Swift music which would result in the video being non-monetizable. But improvements in audio isolation techniques have now defeated this countermeasure. Being a professional annoyance is once again a career choice.
Edit: this is one instance I've personally seen: https://www.instagram.com/p/C7IEFxQSJQw/?hl=en&img_index=1
>Provoke a response
They mostly do it to cops and people in authority. It's their right to do so, they should be able to. They expose so many cops and authoritarians who blatantly do not respect citizens' civil rights. Good for them.
The fact that "oh no you're annoying me I'm going to arrest you because you're annoying" is even a talking point from you is baffling.
In Santa Barbara there is a group that targets random businesses; random shops and restaurants with outdoor eating.
It sucks for the business, it sucks for their clients, it sucks for random people walking by on the street.
I'm all for limiting the unchecked authority we give police, we need to end qualified immunity, etc. But we should take the problem on directly. And I'm all for filming cops who abuse their privilege. But the reality I've seen in person is this is sucky.
> The fact that "oh no you're annoying me I'm going to arrest you because you're annoying" is even a talking point from you is baffling.
Who are you replying to? What did I say that's even close to this. Talk about baffling.
Good point, I may have misread part of what you said.
Yes, people who are assholes in public are annoying. Shoplifting and bank robbing are probably also career choices. Don't rely on a side effect of "big copyright" systems to save us.
Are there really that many people who 1) are aware that this could be effective, and 2) are quick witted enough to pull their phone out and play music in response to being harassed?
Perhaps we should demonitize every form of journalism and media that annoys this guy!
Can't really imagine why you would both give a response to an interview-style question, while being recorded, and simultaneously not want that response to be public. Or are they doing it secretly?
In my opinion, this is a bug, not a feature. If you pull out your phone and play Taylor Swift, you are in fact making a public performance without permission. Even if you had permission (as some cops allegedly do to use some bands music for this purpose), this is not the correct method to deal with professional annoyances.
As a police officer, your job is to be the adult in the room. Society is trusting you with a tremendous amount of power. If you can't handle some annoying whiny YouTubers professionally without using "countermeasures", you should hang up your badge and get another job.
This is actually illegal for you to do.
It didn't seem to do much better than audio filters for ffmpeg that have been tuned for removing background noise and enhancing voice. Maybe I'm missing something or using the wrong source data.
Elevenlabs is the only model company I can think of that is ahead of everyone else in their category. Video and LLMs are hyper competitive, but voice is a one-company game. Elevenlabs hired up everyone in the space and utterly dominates.
I'm hoping this changes. They've been in pole position for over a year and a half now with nobody even coming close.
There's probably a reason why they're so research-oriented. The minute an open source model is released that rivals Elevenlabs in quality, they're in big trouble. There's absolutely zero moat for their current products and there are fifty companies nipping at their heels that want to be in the same spot. Elevenlabs' current margins are juicy.
Since when are characters a currency?
Actually other major cloud providers, including AWS, Azure and GCP, have similar character / token / word count based pricing as well.
I did end up clicking thru to get the full story. On their pricing page (https://elevenlabs.io/pricing) they are up front, with several monthly tiers; their "most popular" $11/month tier says "100k Characters/mo (~120 mins audio)"
Nowhere is this novel usage of "character" defined. I know about text characters, and story characters. But this seems to be different. It's hard to imagine why they didn't define what they mean by "character" or just make the pricing model more straight-forward.
The definitely-most-popular Creator price point is 100,000 "characters" for $22, meaning, according to the FAQ, 100 minutes of audio listening costs $22. Not sure why they can't just say "100 listening minutes" or whatever.
Though I just noticed they also claim that the 100k characters is ~120 minutes of audio but 30k is 30 minutes of audio. I'm not sure where they're getting their numbers but it looks like they're either being dishonest about the former or underselling the latter.
Because this part of the product is per minute, but the other (and earlier) one is charged per character of text.
> Though I just noticed they also claim that the 100k characters is ~120 minutes of audio but 30k is 30 minutes of audio. I'm not sure where they're getting their numbers but it looks like they're either being dishonest about the former or underselling the latter.
It's just a very rough estimate, the comparison points are "about 10 minutes, about half an hour, about 2 hours". I think that would be clearer if it said ~2 hours.
Everything else they bill in characters, so that's the "currency" customers have. It works out I think as costing roughly the same per minute as generating audio.
many many ppl are complaining that they have to spend quite a bit of credit to get the desired effect
so likely this is just another "pay-to-fine-tune" not unlike "pay-to-play" schemes in online games--the hook is to get you in to buy credits which you will use to chase the desired quality.
besides there are local TTS models now that rivals Elevenlabs. Their pricing is ridiculous $200/1M is way too expensive.