Cohere Transcribe: Speech Recognition
cohere.com
cohere.com
In OCR, even when the characters are poorly scanned, the deep domain understanding these large multi modal AIs have allows it to understand what the document actually meant - this is going to be order id because in the million invoices I have seen before order id is normally below order date - etc. The same issue is going to be there in ASR also is my worry.
With OCR the risk is you get another xerox[1] incident where all your data looks plausible but is incorrect. Hope you kept the originals!
(This is why for my personal doc scans, I use OCR only for full text search, but retain the original raw scans forever)
[1] https://www.dkriesel.com/en/blog/2013/0802_xerox-workcentres...
Probably the answer is simply to tweak the metric so it's a bit more smart than WER - allow "unclear" output which is penalised less than actually incorrect answers. I'd be surprised if nobody has done that.
For example, if the prompt includes that Caitlin is an accountant and Kaitlyn is an engineer, if you transcribe "Tell Kaitlyn to review my PR" it will know who you're referring to. That's something WER doesn't really capture.
BTW, I built an open-source Mac tool for using gpt-4o-transcribe with an OpenAI API key and custom prompts: https://github.com/corlinp/voibe
>Timestamps/Speaker diarization. The model does not feature either of these.
What a shame. Is whisperx still the best choice if you want timestamps/diarization?
It doesn't use an extra model (so it supports every language that works with Whisper out of the box and use less memory), it works by applying Dynamic Time Warping to cross-attention weights.
My experiences with Google’s Chirp have been horrendous, with it sometimes skipping sections of speech entirely, hallucinating speech where the audio contains noise, and unreliable word level timestamps. And this all is even with using their new audio prefiltering feature.
AWS works slightly better, but also has trouble with keeping word level timestamps in sync.
Whisper is nice but hallucinates regularly.
OpenAI’s new transcription models are delivering accurate output but do not support word level timestamps…
A lot of this could be worked around by sending the resulting transcripts through a few layers of post processing, but… I just want to pay for an API that is reliable and saves me from doing all that work.
See the very bottom of the page for a transcription with timestamps.
I upload edited gameplay vods of twitch streams on youtube, and use whisper-large-v3 to provide subtitles for accessibility reasons (youtube's own auto-subtitles suck, tho they've been getting better).
My checklist for a good ASR model for my use case is:
1. Have timestamp support.
2. Support overlapping speakers.
3. Accurate transcripts that don't coalesce half words/interrupted sentences.
4. Support non verbal stuff like [coughs], [groans], [laughs], [sighs], etc.
5. Allow context injection of non-trivial sizes (10k+ words)
1 is obvious because without it we can't have subtitles. Force alignment fails too often.
2 is crucial for real world scenarios because in the real world people talk over each other all the time, in my case it's a streamer talking over gameplay audio, or when the streamer has guests over. When 2 people speak the transcript either ignores one of them, or in the worst case, both of them.
3 and 4 are an accessibility thing, if you're deaf or hard of hearing having a more literal transcript of what's being said conveys better how the speaker is speaking. If all subtitles are properly "spell-checked" then it's clear your model is overfit to the benchmarks.
5 Is not a requirement per se, but more of a nice to have. In my use cause the streamer is often reading stream chat so feeding the model the list of users that recently talked, recent chat messages, text on screen, etc. Would make for more accurate transcripts.
I've tried many models, and the closest that fulfill my needs are LLM style models on top of forced alignment. It's too slow, so I've been sticky with whisper because with whisperx I can get a transcript in 5 minutes with just a single command.
One thing all these models do (including whisper) is just omit full sentences, it's the worst thing a model can do.
It has the most crisp, steady P50 of any external service I've used in a long time.
My experience with Cohere and interacting with their sales engineers has been boring, I say that is the most flattering way possible. Embeddings are a core service at this point like VMs and DBs. They just need to work and work well and thats what they're selling.
So far, the best I have found while testing models for my language learning app (Copycat Cafe) is Soniox. All others performed badly for non native accents. The worst were whisper-based models because they hallucinate when they misunderstand and tend to come up with random phrases that have nothing to do with the topic.
Soniox (stt-async-v4): 176/248 (71.0%) ElevenLabs (scribe_v2): 170/248 (68.5%) AssemblyAI (universal-3-pro): 166/248 (66.9%) Deepgram (nova-3): 158/248 (63.7%) AssemblyAI (universal-2): 148/248 (59.7%) Cohere (transcribe-03-2026): 148/248 (59.7%) Speechmatics (enhanced): 134/248 (54.0%)
P.s. how do I get this to render correctly on here?
It's for code though, not lists or bullet points.
- 1. Soniox (stt-async-v4): +176 new cases, running total 176/248 (71.0%)
- 2. ElevenLabs (scribe_v2): +26 new cases, running total 202/248 (81.5%)
- 3. Speechmatics (enhanced): +12 new cases, running total 214/248 (86.3%)
- 4. NVIDIA Parakeet (TDT 0.6B v2): +6 new cases, running total 220/248 (88.7%)
- 5. Mistral (voxtral-mini): +3 new cases, running total 223/248 (89.9%)
- 6. Gladia: +2 new cases, running total 225/248 (90.7%)
- 7. AssemblyAI (universal-2): +1 new cases, running total 226/248 (91.1%)
- 8. Deepgram (nova-3): +1 new cases, running total 227/248 (91.5%)
- 9. Cohere (transcribe-03-2026): +0 new cases, running total 227/248 (91.5%)
- 10. AssemblyAI (universal-3-pro): +0 new cases, running total 227/248 (91.5%)
And someone has already converted it to onnx format: https://huggingface.co/eschmidbauer/cohere-transcribe-03-202... - so it can be run on CPU instead of GPU.
This kids make sense because "compiling" (training) the model cost inhibitly much, and we can still benefit from the artifacts.
I recently was interviewed for a podcast, and she published it on Apple Podcasts. Apple does a transcript of the podcast. I assume it’s some kind of AI (not sure if it’s the same engine as Siri -which I’m not too thrilled with).
It made quite a few errors (not too bad -but errors, nonetheless), but the thing that annoyed me the most, is that it didn’t differentiate between speakers.
Accurate and fast model, very happy with it so far!
This is a good option. Will check it out.
Seems to not be to difficult in finding or creating training code. So a pretty decent amount of high quality training data should be many hours. And a few hours in high end data enter GPU compute, and many iterations to get it right.