Show HN: Cleanvoice – Automated Podcast Editing
cleanvoice.ai
cleanvoice.ai
It really is the difference between being able to edit a 1 hour episode in 1 real life hour (editing at 2x speed) vs literally spending 5 hours to edit 1 hour when there's a lot of filler words or ums. That's due to having to stop every few seconds, think about when to cut it and perform the cut. This is using a heavily optimized keyboard shortcut focused workflow too.
I hope you don't mind constructive criticism but in my opinion your "after" version doesn't sound natural. This isn't an attack on your service specifically, because the outcome is the same with all of the automated tools I've tried. I haven't tried them all but I did play with a few of them.
For example in your case the pause between "Removing" and "filler" doesn't match the pace of the rest of the sentence and the transition from "very" to "time" has a very hard cut. This is also a 10 word clip that's about 6 seconds. If you listened to a 1 hour podcast episode that was edited like this it would be much more noticeable.
There's so many intricate and subtle details around when and what to cut to remove these things in a way where it's not noticeable. Are there any paths moving forward in AI / ML that can lead to this being indistinguishable from being humanly edited?
I debated deleting this comment before posting it because it's a combination of feedback but also saying the service isn't something I would buy in its current state but I'd like to think it's more beneficial to post this to show there is a real demand for this service if it can be executed flawlessly.
Regarding if ML would be indistinguishable from humanly edit. Hard to tell. I think it will be like self-driving cars in the future. 98% edits good 2% bad edits.
My first impression of the unnatural recording was that it must be that way to make it easier to get a good result, but then the result doesn't sound natural either. I think a lot of this is the drawn out uterrances made the speaker vary their pitch/cadence a lot more than usual. Once edited to remove the gap, the sudden change is very noticeable.
I don't think that's due to your software, but just a fact of the unnatural source audio. I think a different, more realistic source audio could let you have a really awesome example, without it being disingenuous or not representative of real-world results.
Thanks for jumping into the ring and answering questions in here!
Using RiversideFM to get two locally recordings is also a big help.
I was sat next to an audio editor and producer at a wedding recently and we got on to this topic and he said “your number one job when editing an interview is to make the host sound good and then just do the minimum on the guest, otherwise you’ll waste too much time”.
Doing the kind of editing 8 hours a day I can see why he says that.
I'm kind of surprised that wedding producer openly said that. My philosophy has always been the opposite. One of my main goals of the show is to make the guest walk away thinking this was the best podcast experience they ever had from start to finish as well as do everything I can to make them come off as good as possible.
I rarely cut content but most episodes have hundreds of manual edits to remove filler content and create a more concise flow by removing long pauses because my 2nd main goal is to optimize for the listener. I keep the edits organic at the same time by leaving in some filler content and subtle things like a deep inhale or a sigh because there's a lot of meaning around that when it comes to sentiment and tone, the same can be said for sometimes leaving in an extra 500ms pause to amplify the meaning behind something. At the same time, sometimes filler content gets left in because it flowed too quickly into the next word so cutting it sounds too unnatural as if it clipped.
This is why I think it's a crazy hard problem to get a machine to be able to make decisions like this.
I do use separate recordings (we each record our track locally), it definitely helps eliminate the few cases where we talk over each other or being able to lower the volume of a laugh so it doesn't overpower what the other person said while still keeping it in because it's a good part of a conversation and a snort or laugh can easily be the difference between a listener wondering if the guest was offended or happily agreeing with something.
None, everything is manual.
I use DaVinci Resolve to do the editing where both the guest and myself have separate tracks. Then I line up the tracks (only takes a few seconds) and start playing things from the beginning at 2x speed. I stop to make cuts mostly to remove filler content.
Through out this process of editing I'm also creating show notes as I go. An example of the end result is here https://runninginproduction.com/podcast/103-great-question-m.... Basically every few minutes I recap what was said into a 1 sentence bullet point with a timestamp. Along the way I list out techs used as tags and list out reference links / libraries into a Markdown document. Then once I'm done editing the show I write a few paragraphs which is a TL;DR of the episode.
All in all if the guest uses minimal filler words or noises it takes about 1 real life hour per 1 hour of recorded content to do all of the above. For context, the episode I linked has someone who I would bucket into a category of speaking very fluently with minimal filler content. I was able to blaze through that one.
I also have a 2560x1440 display and use the "always on top" feature of most window managers to layer the Markdown document and a preview of the page just above the waveform in DaVinci Resolve so I can quickly make cuts and update the notes with minimal mouse movement. Almost everything is keyboard driven.
What tools can be used to speed up that process?
It is mentally taxing though, it means during the whole editing process my brain is constantly identifying and removing filler content, listening for specific tech choices to tag, listening for specific references that could be interesting to link, listening for mentions of libraries to link and also digesting the main takeaway of what's being said to sum it up into a note. All of this happens in 1 pass during the editing process. I tried doing it in 2 passes where I only focused on mechanical editing the first time around and doing the show notes on the 2nd but it took longer in the end.
I’m not suggesting no edits, just relaxing things a bit so the burden doesn’t become an existential threat to the podcast.
Probably but I have no way to turn this off and be happy with myself.
I try to approach everything I do from the angle of "what needs to be done to make this as good as it can be with my current skill set?". From a listener's perspective if I had to listen to something with a bunch of mouth noises, ums every 3 seconds or long pauses I would end up focusing on that instead of the topics being covered. It would give off a wrong impression that conflicts with my core values.
I can appreciate this, but...
>It has gotten to the point where after 2 years of running my podcast I'm seriously considering *stopping the show* because I'm getting burnt out from editing and without sponsors it's not feasible to hire an editor, but even with the show making no money I would happily pay triple your asking price if I could click a button and have the problem solved in a way that matched a human's ability to edit out filler words.
(emphasis mine)
I don't think it's actually the case, but extrapolating a list of priorities from this, I can only arrive at the following:
Priority #1 - no aahs, umms, slurps or smacks
Priority #2 - no ads or obvious sponsors
Priority #3 - surfacing hard-won lessons from experienced folks for the world to learn from
Maybe that resonates, maybe it doesn't, but to me it seems upside down.
I'm only commenting because what you're describing used to be me. I used to do this type of editing for recordings of live audio production and I've gone down the rabbit hole you're describing above. The problem is there's no obvious point of 'done', and chasing perfection in the output can become a pathological obsession. You can get so lost in mating phase angle at each end of a trim or taking an eraser to get rid of a sleeve drag across the desk that you lose sight of the totality of it. Ultimately you end up in a weird uncanny valley, like those folks that keep 'fixing' their face with plastic surgery. Once you get to that point, you can no longer identify specific issues to correct, you just fall into a diffuse unease.
For me podcasts are a way to join a conversation that I wouldn't otherwise have an opportunity to listen to. I don't see them as a show or corporate media product, and the more they start moving that direction the less inclined I am to listen to them. Julia Childs had a quote that I've found oddly applicable in this context: 'It's so beautifully arranged on the plate, you know someone's fingers have been all over it.'
Hope this doesn't come across as negative. Good luck!
> I don't think it's actually the case
Do you mean the editing process isn't what's making me want to stop the show?
For perspective, phrases like sleeve drag aren't even in my vocabulary. I mainly do my best to quickly get rid of filler content without it sounding like there's hard cuts. It's not chasing absolute perfection where I'm zoomed into the waveform so much it looks like an oscilloscope while I hem and haw about there being a 35ms or 50ms pause between 2 words, or agonizing if I should leave an um in there so things don't sound over processed.
Here's a screenshot while editing an episode where the guest was extremely fluent and I didn't have to edit much filler content: https://i.imgur.com/7CBZ1yc.jpg, for context the episode was 90 minutes long but I zoomed into the point where you can see a ~10 minute chunk (normally I'm zoomed in much more while actively editing). This is a best case scenario where I "only" had to do 305 cuts for a 90 minute show. In the worst case scenario it's gone as high as 1,800 cuts for 90 minutes.
I try to keep things organic while being respectful to listeners. All of the cuts you see there are related to removing filler content (umms, ahhs, mouth noises and long pauses). I also remove their dead air when I talk to avoid any of their mic's background noise overlapping my voice since it's all recorded in an uncontrolled environment.
The before and after is pretty staggering even with a fairly minimal amount of filler editing. To be honest I would feel embarrassed posting the unedited version of most episodes.
It's also very interesting because in a way I think posting a much less edited version where I kept all of the filler content in wouldn't save me much time in the end. Not to sound too over confident but I'm really confident in my ability to perform quality assurance of each episode while I'm doing the editing. I haven't listened to a single episode in its final form because I've gone through each sentence and phrase multiple times during the editing process. For example I'll start playing it, hit a cut point, make the cut, rewind a bit and ensure things flow smoothly, then continue onwards.
If I did a much less edited approach I would still need to listen to the show at 2x speed, so no matter what I'm spending 30 minutes listening to 1 raw hour. However I'm also creating timestamped show notes like you see here https://runninginproduction.com/podcast/99-a-custom-electron... along the way while editing so I have to pause to write these down.
Basically I would still be spending quite a lot of time to produce things and I don't think I can outsource that because it would involve finding someone who is not just an audio editor but they would need a ton of domain knowledge around 100 different assorted technologies. A lot of those timestamped notes aren't verbatim quotes. I'm mixing quotes with trying to keep it concise to fit into 1 line. I'm also making judgment calls on what to include because not everything is worth making a note over, otherwise there would be one every 30 seconds (I used to do this in earlier episodes).
Personally I would rather have a transcript with timestamped links where each guest is broken up into their own paragraphs but to have them done right costs a lot of money. Every machine generated transcript service I used had really bad grammar issues and mistakes. A human reviewed one would be well over $100 per episode to make which is a lot when the show already has a net loss on every episode (hosting).
That quote you mentioned was really good by the way. I'd like to think my editing style is more on the side of someone occasionally using their hand to make sure the food doesn't slide off the plate while you run the plate over from the kitchen to the customer. That's how I feel during the editing process. I'm trying to get through it as fast as possible but taking great care to ensure a high quality meal arrives to the customer. I'm optimizing for folks wanting to come back to their favorite restaurant on a regular basis, not serve an artificial feeling $10,000 plate to a king.
> Do you mean the editing process isn't what's making me want to stop the show?
No this is just confusing language on my part. What I meant was that I don't actually think those are your list of priorities in order, but that is how they could be extrapolated based on which part has to give.
OK so after your description of your workflow I think I was reading too much into where you were at specifically with regards to the content clean-up. I was worried that you were hovering over every sentence trying to optimize it and was just trying to talk you down off the ledge. :) For some reason I tend to gravitate towards jobs where I'm at my best when nobody knows I did anything at all. Editing is probably one of the best examples of this and, as a result, it's hard for anyone that hasn't done it to truly appreciate how much work there is behind it.
(Some of this is selfishly motivated btw, I've been following your podcast since the spring and don't want it to go offline lol. If i have to listen to some CTO's lips smack every time he gets ready to talk I'll allow it. :) )
Yes, this is perfectly said. It's exactly how I feel and what I strive for. I think most folks would be surprised if they listened to a before / after even if all that was done was occasionally remove filler content and mouth noises. It's like that one business analogy iceberg picture with "success" being the 10% that's above water and the other 90% is buried with all sorts of things you never hear about.
> Some of this is selfishly motivated btw, I've been following your podcast since the spring and don't want it to go offline lol
That really means a lot and I'm happy to hear you like the show but unless a big pile of money falls from the sky to afford hiring a dedicated editor and human reviewed transcripts then I have to pull the plug. I've already been feeling this way for 3-4 months but tried to power through it. I've reached the point of feeling resentment and disgust just thinking about opening my editing tool of choice and it's taking its toll. It sucks because I would love to record the show until the day I die but these are the cards I'm dealt and I have to choose sanity over suffering at this point.
There's no middle ground due to the last half of my previous reply.
I posted a new episode today since I still have 6 unedited episodes left, I figured I would release them once a month until they run out. I'll also be posting a "what's happening with the podcast?" video on YouTube tomorrow.
I like podcasting, but I hate editing them. I tend to stutter and have a lot of filler words in my podcast. That's why I created Cleanvoice, in order to spend less time editing them. Cleanvoice is an ML tool which removes filler words, mouth sounds, stuttering and dead air from your podcast. To use it, just upload your podcast - wait some minutes - download the cleaned audio.
It's still not perfect, but it's at a stage where I can blindly use it on every single one of my podcast.
I would love to hear your feedback!
That scenario seems really appealing for conferences, even if it just quietens down the verbal ticks, but I'm guessing if the lag is too great it would get like a bad lip sync issue
What it sounds like the GP is after is something more like hiss and pop removal (to use an only vinyl analogy) and that’s a different and also simpler problem to solve. I’d wager there are already tools on the market for that.
To me this seems like it could be worthwhile even if it results in silence or less prominent umms and other filler - I've sat through enough conference talks by technically gifted people who I very much wanted to hear but who unfortunately make their talks much harder to follow due to the ticks. It might even help relax some nervous speakers if they knew any of these that creep in were being suppressed.
With an EEG cap, though, I bet a smart person familiar with the methods could bash something together in a day that would work.
1) It doesn't work well if you have a strong accent. As an non-native speaker, the transcription were quite bad, making the editing quite bad.
2) Cleanvoice works with multiple languages, descript doesn't.
3) Cleanvoice can remove stutters (not always, but it tries) and mouth sounds like lip smacking, teeth clicking. Descript can't. This is not a big deal for most, but since I stutter alot this was essential.
My approach is different from Descript. They use a transcription service, and then they edit the audio based on the text. I work directly on the phonetics level. Allowing me to have more control over audio.
Depending on the needs, either one is better. I guess you should try it for yourself and compare.
While Cleanvoice has some niche features that Descript doesn't offer I would not be surprised to find them rolling these features out in the next major release they're doing. IMO the founder of Cleanvoice should sell/join Descript.
Is it possible for you to do a live, personal demo? No logins or anything. I'm thinking something where you tell people to start up their audio and then give them a quick prompt like "Describe your breakfast yesterday." Record for 30 seconds, and then let them play back the original and cleaned versions. You could limit them to, say, 5 goes, with a different prompt each time.
I suggest it because a) a little personal investment makes it more likely they'll give you their email address for signing up, and b) many potential customers underestimate how much they need something like this.
My biggest fear is that without login, people will start abusing it in ways that I don't expect. Definitely considering it. Thanks you!
Thanks for listening, and good luck with your project!
Can I suggest the ability to export as project files for popular editors for your roadmap? It'd cut professional workflows down substantially, which would be worth an (even higher) upcharge.
(It wasn't immediately obvious to me if you already did this)
Edit: https://cleanvoice.ai/integrations seems pretty close. I'd honestly charge more for integrations and provide a base tier for just exporting sound. I imagine most indie users would benefit from finished exports enough to pay, while project files would command a higher fee from editors looking to speed up their workflow to take more clients. That's where I'm coming from on pricing tiers and upcharging for professional features.
Regarding Pricing, that's a good point. I will definitely consider it, thank you!
https://www.rev.com/blog/how-to-import-an-edit-decision-list...
This could help when you have for example multiple camera angles, to switch between/do morph cuts (https://helpx.adobe.com/premiere-pro/using/morph-cut.html) for video interviews.
It seems like this particular product might do a better job of the automated editing specifically, but Descript has a ton of other features (speaker identification, transcription, real-time editing based on text edits, asset management, and uploads), and I definitely wouldn't trade them for marginally better auto-removal of noise and filler.
Does anybody who develops Cleanvoice have any commentary here?
Before building Cleanvoice, I tried to use Descript for my podcasts.
1) It doesn't work well if you have a strong accent. As an non-native speaker, the transcription were quite bad, making the editing quite bad.
2) Cleanvoice works with multiple languages, descript doesn't.
3) Cleanvoice can remove stutters (not always, but it tries) and mouth sounds like lip smacking, teeth clicking. Descript can't. This is not a big deal for most, but since I stutter alot this was essential.
However, if none of these apply for you. There is no reason to change from Descript.
This is not the case with my app. I keep the edits longer than shorter, since I also find that unlistenable.
Their pricing is also similar, but Auphonic allows both subscription and prepaid "credits".
Auphonic and Cleanvoice go well together.
I guess the idea is to have your podcast edited by Cleanvoice and then the audio post-processing with Auphonic.
I definitely prefer pre-paid credits to a subscription given my podcast production varies a lot.
Is it possible to keep some filler words? I make something similar (but not professionally), and sometimes I like too keep a few of them.
Currently, there is no way to set it for now. But customization is planned for Q2 next year.
>Is it possible to keep some filler words? For now no, but keeping some filler sounds to keep it authentic is something which I plan.
In other comment, eganist posted a link to https://cleanvoice.ai/integrations It looks interesting because I can choose which to keep and even use it to sink with video [with some additional work]. I didn't see it the first time in the page.
That will change in Q2, when I add support for video.
For single camera floating head style videos where you're continuously talking about 1 topic it's going to be very jarring if you start cutting out filler words. You'll end up with a bunch of jump cuts where it looks like video frames are dropped.
If yes, you might consider a page or callout about that use-case, as it might attract some additional users. Just a thought.
Out of curiosity: Which ai-technology did you use? OpenAI? Google API? Or did you train the models yourself with Python (sth. like Tensorflow)?
Cheers, Mike
I trained my own models. No OpenAI/Google API.
Liebe Grüße, Adrian
Other than which country, though? Presumably an English speaking one - UK? New Zealand? Canada? US?
1. Any plans for an API and bulk pricing?
2. Any plans to add loudness normalization, balancing, etc to the processing?
1) API Access will come end of Q1.
2) In the next 6 months, No. However, Auphonic would be a good fit for you.
Zencastr RECORD
Transistor.fm HOST
but looks like I'll be adding
Cleanvoice AI CLEAN UP
Auphonic AI EQ
any other suggestions for optimum output from multiple inputs all over the world?
Does it involve relying on speech to text with timestamps and then a series of cuts based on that?
Additionally, if you are located in a country that (like Germany for example) has regulations on the necessity of an imprint, this might also be missing.
Or do I misunderstand the law?
[1] Strictly necessary cookies — These cookies are essential for you to browse the website and use its features, such as accessing secure areas of the site. Cookies that allow web shops to hold your items in your cart while you are shopping online are an example of strictly necessary cookies. These cookies will generally be first-party session cookies. While it is not required to obtain consent for these cookies, what they do and why they are necessary should be explained to the user.
[1] - https://gdpr.eu/cookies/
> By posting your Contributions to any part of the Site or making Contributions accessible to the Site by linking your account from the Site to any of your social networking accounts, you automatically grant, and you represent and warrant that you have the right to grant, to us an unrestricted, unlimited, irrevocable, perpetual, non-exclusive, transferable, royalty-free, fully-paid, worldwide right, and license to host, use, copy, reproduce, disclose, sell, resell, publish, broadcast, retitle, archive, store, cache, publicly perform, publicly display, reformat, translate, transmit, excerpt (in whole or in part), and distribute such Contributions (including, without limitation, your image and voice) for any purpose, commercial, advertising, or otherwise, and to prepare derivative works of, or incorporate into other works, such Contributions, and grant and authorize sublicenses of the foregoing.
It sounds an awful lot like "we are allowed to do anything and everything we want with the content you upload to us". Maybe I'm misunderstanding something, but I'd be extremely hesitant to upload any content I create to a service with those kinds of terms.
I would like to rewrite it.
What I do is just keep your files on the server for a week. In case you have an issue, I will look into your file to fix your issue. And if you want, you can give consent for me to further improve the service. (Say you have an accent which the AI is bad and I can use your audio file to understand why it failed.)
I would pay for a piece of software that does that job on my computer with no Internet.
This way? I may even end up in court for saying something “improper”…
Edi. OK: I’ve just read the developer’s reply below.
Honestly: you need to fix this because right now it is more scary than not.
Congratulations for the project but please do fix this.
I will get proper terms soon as possible. Especially, since now people have mentioned it.
In your FAQ, you say: "Currently we remove lip smacks, saliva crackle, mouth clicks and harsh parts of breathing (not the whole breath). If you want to remove a particular mouth sound (ex. Chewing), write us in the chat as a feature request." I don't think most English speakers would understand what "harsh parts of breathing" are. Typically a parenthetical example in English would be written "(e.g. chewing)" not "(ex. Chewing")".
Your question "What filetype and sizes do you support?" doesn't answer what filetypes you support, and I suspect the singular "filetype" was a grammar error. You also write "We have an audio file size limit of 1.5G per file or in case you are uploading multi-track and a total file size of 2 GB. ". The part that says "or in case you are uploading multi-track and" doesn't make any sense in English. I think you mean "We support file sizes up to 1.5GB per file for single-track files, or 2GB if you are uploading a multi-track file as separate files." but I'm not sure.
In general I don't understand why each selling point has a separate FAQ page but the FAQs are often not related to the selling point. I don't think people think the "Mouth Sound Remover" page is the one that lists file size support, while the "Stutter Remover" page is the one that lists the maximum number of tracks per project.
Your integrations page lowercases "cleanvoice" whereas other pages write it as "Cleanvoice".
Under integrations, you have a section called "Markers Export". This should probably be "Export Markers" or "Marker Export".
Under "How to Export Edits", you probably don't want to capitalize "Results" or "Editor" unless these are supposed to be title cased, in which case you probably want to title case all of them.
Under your pricing FAQ you have "Does my credit expire at end of the month? Your credit will reset every billing month. Unused credit will be lost." This is needlessly confusing. You use the verbs "expire", "reset", and "be lost" to describe the same thing, and you don't actually answer the question. Also you don't want "at end of the month", you want "at month's end" or "at the end of the month". I would rewrite as "Does my credit expire at the end of each month? Yes. Credit resets every month and cannot be carried over to future months. Unused credit will be lost." This is a terrible business model, though, and so I suggest you not do this. Either sell as a subscription or sell as a credit model, not both, this is gross.
In general I think you want to pay someone who is a professional English copywriter to fix your website. Cheers.
Edit: I just noticed your changelog is powered by a service called Headway. I am not sure if you also made Headway, but Headway's website is also in need of English copyediting.
I'm curious why the Subscription + Onetime Credit is bad. But I agree it is confusing.
My understanding is that not every customer wants or needs a subscription, since they upload podcasts irregularly.
This business model is seen in other AI products:
https://www.remove.bg/pricing https://auphonic.com/pricing
I am very grateful, you took the time to help out. Really appreciate it!
They are great. As a German native speaker I came a long way with using them when I needed valid translations.
Perhaps you could have a mode to detect how much one stutters, and parts worth redoing without spending as much time combing the whole thing.
I haven't had a lot of luck using teleprompters but maybe I just haven't hit of the right setup.
Something else someone told me recently was to try to work in short segments that you redo until you get right and then do a cut to the next segment somewhere that it's natural.
The result feels something like pixel art: Clearly not the closest possible imitation of conversational speaking, but something else. A style in its own right with different considerations.
Now that it's par for the course to have jump cuts, I see them used more sloppily everywhere, where it's clear the narrator decided where to do the cuts after the fact. Cutting off the beginning or end of a phoneme, missing or repeating bits of a thought because they they liked one phrasing in recording but opted for another one in post, misordered cuts where something which moved in the background moves back to its old place, etc. Phillip's style looked lazy but it can't really be imitated with actual laziness.
These days I look back and really cringe at the substance of his show. But I still see the style as professional.
I remember fondly a student of mine who seemed unable to express himself properly. I told him to memorize his final project dissertation because otherwise it would be a wreck (OK, I did not say this last part, it was more of a suggestion).
BOY: did he memorize it. He got an honors and I did think “this guy has really done it, and it sounds like music!”
When you do it well, it tells.
I find the cadence very unnatural when all the spaces between phonemes are removed.
Any editing can be overdone and, while I do a modicum of editing out umms, you knows, and other verbal ticks when I'm putting together a podcast interview, I'm not fanatical about it.
You do occasionally get someone who just speaks quite slowly and it is sort of annoying to listen to as audio. So I've done some automated gap reduction is a couple cases.