Show HN: SpeechBoard – Edit Podcasts from the Transcript
speechboard.co
speechboard.co
Craig from YC here. Ramon Recuero and I built SpeechBoard.
Here's how it works: you record a podcast and upload it to SpeechBoard. We run it through a few speech to text APIs to generate a transcription for you. From there you can delete words from the transcript and we cut them from the audio.
Then you can download three files: your edited audio, your original file with cuts marked in metadata for importing to Audition/Audacity, and labels for importing into Audacity.
Nailing the in/out points of words was the hardest part, which led us to create the Audition/Audacity import feature and now I think that's the best part :)
Let us know what you think!
So it's written in Python and the main API we rely on is IBM's Speech to Text.
One thing that became an issue was deleting half of a word or inserting characters. Now we handle that with JavaScript. In the future we'd like to get some Lyrebird in there :)
If you have specific questions just lmk.
Are you storing the audio and/or edited transcripts?
But we don't have a formal policy. We'll work on that.
One suggestion is to be able to add extra silence (whitenoise?) so we can retain a natural flow after cutting aggressively.
Maybe recognize an ellipses ... and add a second or two?
It's definitely something we're interested in though.
Check out ScriptSync too - http://www.avid.com/products/media-composer-scriptsync-optio...
I also talked about this project in depth on Spencer Wright's podcast - https://theprepared.org/podcast-feed/2017/10/15/craig-cannon...
We do it with a few speech to text APIs.
We've only tested with English so far but it should be able to handle a bunch: Arabic, English, Spanish, French, Portuguese, Japanese, and Mandarin.
The other major pain-point that I've identified after spending 10 years in the podcast and audiobook industries is audio mastering. As a mastering engineer, I know how troublesome it can be to master audio so that it is at a competitive dynamic range.
Would you be interested in implementing an "auto-master" feature that I've developed? If so, please reach out! My email is in my HN profile.
Of course, the focal point of the episode isn't the technology itself, but the implication it could have on society once it gets good enough that its output is virtually indistinguishable from real video (i.e. fake news in the form of convincing-looking videos).
Highly recommended if you have the time.
[0] https://www.radiolab.org/story/breaking-news/ [1] https://www.futureoffakenews.com/
Sample:
https://reader.listensynced.com/ycombinator-jessica-livingst...
As a podcast junkie - I've often run into issues with searching and sharing. Linking transcript and audio is first step to solving this...
I think the transcript creation service in itself must be worth something for these guys.
YC does :)
For the YC podcast we chose to host with Backtracks because they offer transcription and allow you to link back to exact times in the episode, sort of like YouTube.
Since August of 2016 I've listened to 30 days worth of podcasts and saved another 22 days worth of time by skipping introductions, listening faster than 1x, etc. The most annoying thing I'm facing is that I've heard hundreds of stories and the audio cannot be indexed easily. If I want to send a friend to a specific episode for a certain story, I have no good way to remember if it was the Freakonomics podcast, This American Life, Story Collider, Planet Money, or one of the 20+ other podcasts I listen to.
I'd love a system which would make available a searchable transcript of every podcast. I couldn't pay for transcribing all of them, but I'd pay 50 cents/podcast. Google tells me 1.75/minute will get a transcription from the top listed service, so if we had 210 people like me, we could transcribe an hour of audio.
I speed-read. I don't speed-listen. And my lifestyle has zero podcast-compatible travel time. So there's a whole world of great content in a very user-unfriendly format for me currently.
Great job!
human@speechboard.co
We're looking to chat with hobbyists to see what you'd like out of it.
I hit an error when I trimmed the text down to:
> Hey this is a different original.
> The text cuts into
Maybe I was too aggressive?
Some undo support would also be really helpful. Have you considered just having a free-form text field, then using something like wdiff to produce the edits? That might make the UX easier, since you wouldn't have to manually reinvent the text editing tools people expect (although you'd have to handle invalid edits, like people adding new words).
Yeah, without looking at your logs I suspect you cut too much. :)
We were using a free-form text field before and it led to a couple issues: cutting words in half + inputs. Both of those basically break it now so we went for a slightly less convenient but mostly functional demo.
I totally agree though, this needs a lot of polish on the UX side.
This looks very interesting. You might be the right person to ask about something related that I am currently working on: do you know of any app that would extract keyword / name based parts of audio? For example, extract only the parts where Elon Musk speaks given audio input (podcast, YouTube etc.)? Alternatively, extract only the parts (-30 and +30 seconds) when a specific word is mentioned.
Thanks!
Audiogrep may be able to do that for ya - https://github.com/antiboredom/audiogrep
Hi Craig, do you know an app/code that can split the audio/transcript based on persons? Detect different persons in a podcast and group the transcript by person. Thanks!
That's something we're also interested in.
You can read up on the subject and see a few projects here - https://en.wikipedia.org/wiki/Speaker_diarisation
https://dsp.stackexchange.com/questions/3119/library-to-diff...
But to answer your question, I have yet to try an app that can do it well.
https://sixcolors.com/post/2017/03/the-dream-of-converting-p...
Similarly, I’d like automated data on what ads are run/read.
Essentially, I’d like rich automated metadata, in addition to timestamped transcripts.
Like the demo. Wish I could try it out on some other audio.
Side question: I just need (good) transcription of audio. I've never been able to find a good service for the price.
Does anyone have any recommendations?
Any services doing this automated?
No idea if they're any good, but they're certainly cheaper than Rev.
If you test another language out definitely let us know how it performs.