LeMUR: LLMs for Audio and Speech
assemblyai.com
assemblyai.com
Really not trying to be a jerk -- I think this is a neat project.
If someone wants to compare, Universal Summarizer [1] can do real time summarization of audio/speech, with unlimited input token length and is free (for Kagi members or can be tried with a trial account). Just point it to the URL of the podcast/audio/speech file.
Paid API is also available [2].
with a few clicks and decide if you really want to watch
https://www.assemblyai.com/playground/v2/transcript/6lu93wlw...
I'm happy to answer questions about the API as well
EDIT: Also it is unclear if you support other languages than English. Whisper does, so in theory you should. There are companies out there where English is not the work language.
It looks like their synchronous transcribe is much slower than whisper, but if you need it fast, you need their realtime ASR (or amazon or google's).
[0] Conformer-2 is trained on 1.1M hours of English https://www.assemblyai.com/blog/conformer-2/ [1] https://www.assemblyai.com/docs/Concepts/supported_languages
- https://deepgram.com/learn/nova-speech-to-text-whisper-api
Not surprising though as at this level all these options are starting to be leveled by inconsistencies in manual groundtruth. Conformer alone also isn’t the most powerful architecture out there for speech. This is also slower than, say running a large k2 zipformer via onnx on cpu.
Also if you have a small shop at this point you can do all of this yourself with whisper large v2 on a single 16gb gpu via some tweaking of https://github.com/guillaumekln/faster-whisper and an OSS LLM.
Interesting stuff but I think margins in this space are getting ready to simply vanish.
It seems like this is cheaper than a full transcript. Is it because it skips stuff like diarization and aligning time stamps?
The floating icon for cookie settings in the bottom left obscures the play/pause button for the audio track in playground.