Subtitle is now open-source
vedgupta.in
vedgupta.in
[0]: https://github.com/openai/whisper/blob/e58f28804528831904c3b...
I wonder if the whole thing is just an AI-generated project. The "About Me" section is pretty illuminating (unabridged):
> I'm a Developer i will feel the code then write.
* What languages are supported? Is there a list?
* What does 'subtitle' do, which 'whisper' doesn't?
* How do I install this system-wide on an apt-based system (in which pip install --system doesn't work)?
What I'd like to do is create a type of ebook of all these transcripts, where if I click on a word, then the corresponding audio will start playing from roughly the same point in time within the interview.
Otter can do this already (if I'm online and logged in to their website), but I don't want to be tied to their website forever. I'd like to have a local copy that can perform similarly. Amazon ebooks can do this as well, I believe, where there is a corresponding verbatim audiobook. However, this project of mine is purely personal. I won't be selling my audio recordings or transcripts.
Any advice? Could software discussed here be helpful in what I'm trying to accomplish?
If you already have a .vtt, this is not a hard exercise to do e.g. entirely in a browser: parse the .vtt (they're simple text), lay out the text as you like with each segment being a clickable element (e.g. a link), and hook that up to seek an `<audio>` element to where you like.
So, the value proposition of a subtitle-generating wrapper for Whisper would be to have an option to split audio into ~1 minute segments, transcribe them separately, and to somehow accurately join them. And I don’t think this one does such a thing.
It gives pretty good subtitles.
I'm not a native English speaker and I tend to use the LiveCaption application in Linux when I attend English speaking online meetings. Would love to have the opportunity to have subtitles in my native language (Greek) too while doing so.
There is currently a problem with diarization, but otherwise, it is SOTA.
I hope Siri does something to improve. It’s voice-to-text for me is still horrible.