I have thought about this a lot. It might make us more considerate of what we're building, i think the speed at which we can mvp has made us less critical of what we output.
great suggestion - the transcription tab is already quite limited for real estate. i recommend trying out the free trial to get a feel for the workflow.
Descript is a timeline editor, so you delete text and the audio goes with it, and you come out with one finished piece. Vocal Slice goes the other way. You highlight a phrase and that span exports as its own file. Vocal Slice is also local rather than cloud based like Descript.
the truly important thing from discussion with users is the actual workflow optimisation that the app provides, but i totally agree some of the copy on the website could use improvement and the comments on this show hn have made it clear that i'm not communicating things as well as i could
Really appreciate your thoughts on this, i totally agree - it is food for thought and as i monitor things it may make sense to tweak my approach. Bloating the tool would be counter productive for the sake of validating its cost, the idea long term is to improve, support and develop smart. Any financial model's goal in my eyes should be the means to allow people access, while justifying the product's existence.
the initial idea came from exporting lines and multiple takes for video game voice over delivery, but the functionality is there for doing longer clips too by making longer text selections. i'd love to get some feedback from this group if you'd like to discuss further? wesley@vocalslice.com
Vocal Slice uses whisper for the speech-to-text step, then performs its own logic on the transcription and audio, search, take matching, audio slicing, etc.
Thank you! speech to text models can be extremely accurate, but it really depends on the audio you feed it. the beauty here with Vocal Slice is that it doesn't need to be perfect because the user is given the tools to fine-tine the selection via the waveform controls and listening to playback. over time i would like to solve the purely automated workflow though, and make the manual step less necessary.
currently no, the app isn't open source - however the app itself announces and is quite transparent on how it functions. it takes your audio, transcribes it locally using speech to text, recording the time stamps of each word. when you select text to make slices, Vocal Slice makes use of the time stamps to align the text selection with the correct audio in and out points, matching the text you have selected.
Thanks tene, from your comment i believe this functionality does exist, and i realise that i don't express it very clearly on my website. Essentially when you select text, it tells you how many 'matches' that text has. So if the same line is repeated, you can quickly seek between the matches to cut/AB multiple takes very quickly!
Thanks for your question - currently Vocal Slice only supports audio files, so you could download your youtube clip, extract the audio and run it through Vocal Slice. I have plans to natively support video in the future.
great question - I had an older iteration of this project which was a much smaller scoped utility (without the waveform functionality) that was a one time purchase on itch. The idea with this project is to add additional features over time (um and ah removal, the ability to actually remove parts of the transcription and output audio).
This is a cool idea but the website was not very performant. Having this as a desktop tool that can be tweaked/driven via claude could be a good approach
Thank you for the amazing support so far, a few of you have reached out via email to share use cases I hadn't thought of. This tool was originally conceived around long voice acting recording sessions which were a pain to sift through, but the privacy aspect has implications in legal and NDA scenarios too.