To me, STT should take a continuous audio stream and output a continuous text stream.
To me, STT should take a continuous audio stream and output a continuous text stream.
Whisper and Moonshine both works in a chunk, but for moonshine:
> Moonshine's compute requirements scale with the length of input audio. This means that shorter input audio is processed faster, unlike existing Whisper models that process everything as 30-second chunks. To give you an idea of the benefits: Moonshine processes 10-second audio segments 5x faster than Whisper while maintaining the same (or better!) WER.
Also for kyutai, we can input continuous audio in and get continuous text out.
- https://github.com/moonshine-ai/moonshine - https://docs.hyprnote.com/owhisper/configuration/providers/k...
(maybe with an `owhisper serve` somewhere else to start the model running or whatever.)
For just transcribing file/audio,
`owhisper run <MODEL> --file a.wav` or
`curl httpsL//something.com/audio.wav | owhisper run <MODEL>`
might makes sense.
https://github.com/fastrepl/hyprnote/blob/8bc7a5eeae0fe58625...
https://github.com/bikemazzell/skald-go/
Just speech to text, CLI only, and it can paste into whatever app you have open.
What exactly does the silence detection mean? does that mean it'll wait until a pause, and then send the audio off to whisper, and return the output (and stop the process)? Same question with continuous. Does that just mean it continues going until CTRL+C?
Nvm, answered my own question, looks like yes for both[0][1]. Cool this seems pretty great actually.
[0] https://github.com/bikemazzell/skald-go/blob/main/pkg/skald/...
[1] https://github.com/bikemazzell/skald-go/blob/main/pkg/skald/...
The short duration effectively means that the transcription will start producing nonsense as soon as a sentence is cut up in the middle.