Gemini-3.5-Transcribe
blog.google
blog.google
The voices contained very industry specific words, the languages changed from one sentence to another, sometimes words in a language were mentioned while a discussion was in another.
The only local model that satisfies me is Voxtral Mini 3b, the only paid API that is slightly better is eleven labs. Yes, Voxtral might not reach the best score in the benchmarks, but to me, it just solves a problem. It might not be the best in terms of speed... but that’s not a problem for me.
Happy to test this new model from Google but I’m not sure I’d go with that instead of something that can run so easily in my machine.
I've been looking to find a simple and fast dictation app for English but almost everything I've tried (from Handy to many apps, eg some with Whisper in their names, after the model I assume) just don't work well. Apple's offering is worse than those though. I even tried with local enhancement models.
We’ve fine-tuned open-weight models to make them better (in our benchmarks) at cleaning up and formatting what you say so you don’t have to edit what you dictate.
I guess OP didn’t chose the right STT model.
Transcribing is perfect in many language and, it focuses on a selected app. Also it does TTS.
Private offline transcription and summary. Speaker identification, working on voice-prints for identification across the corpus.
You should probably not expect a small STT model like Voxtral or Parakeet to do any better, unless it's laser focused on that area in particular and sucks in everything else.
Things change a lot if you need to speak to the model in multiple languages. There are very few models out there that can automatically detect the language spoken and produce correct output.
I tried the larger Voxtral models, but they didn't work for me at all. When I spoke to them in Polish, they produced output in Russian or Ukrainian.
For now, I settled on creating my own plugin for TypeWhisper, which runs Whisper Large on the GPU and does it much faster than pretty much anything else out there. But I'm still hoping that something better will come along, as Whisper Large is quite old at this point.
Transcribing is perfect in many language and, it focuses on a selected app. Also it does TTS.
Using a combi of the large Whisper and koroko/supersonic for TTS. Plugs into dev environments via MCP and hooks (Claude, Codex).
Something like "I hesitated to check it, I should have verified" => "I should have verified" (The "I hesitated..." is removed but I said it because I wanted to let the person know that I thought about it earlier)
I tried the sentence few times and it always removed the first part.
https://ai.google.dev/gemini-api/docs/transcribe#transcripti...
This confused the heck out of me because it makes it sound like the STT model can make function calls in order to execute arbitrary tasks, which wouldn't make any sense. The developer docs (https://ai.google.dev/gemini-api/docs/models/gemini-3.5-tran...) confirm that the Gemini 3.5 Transcribe model cannot in fact make function calls. I guess the blog post is just using very confusing wording to describe how their consumer assistant/chatbot app can both take audio input via Gemini 3.5 Transcribe and then call other things as needed. Or maybe that bullet point was meant to be for a different model announcement and somebody made an editing error.
It is possible to trick pangram - they bias toward a low false positive and a higher false negative - but it is not true that it is essentially random.
> “I don't think that Pangram is bad,” Mantzarlis said. “I think actually Pangram at scale is probably a pretty solid tool. That said, I am extremely worried about it being used in individual cases.”
https://reutersinstitute.politics.ox.ac.uk/news/human-wrote-...
Using it for an individual article to fully determine if its AI or not is "impossible" because you're not even using the tool properly.
No it literally does not apply is the point of the article. Please read what it is about and what it says instead of asking for spoon-feeding.
> but my instinct can be wrong too and I don't stop using it to filter what I read
Again, your instinct is not something that matters to anyone other than you. But you are presenting Pangram as fact and doing on a moral crusade (I WILL NOT READ ANYTHING PANGRAM SAYS AS AI). You can also have an instinct that "THIS IS WRONG" and go on crusade but you will naturally understand your foundation is not solid at all.
Lastly, if your instincts serve you well why are you outsourcing yourself to another instinct? Is it for yourself or to say to others "LOOK AI CONTENT LOOK AI CONTENT!!!"? Is that purely to serve your interests of filtering what you read or are you using it in the wrong way here?
Also, I'm not presenting Pangram as anything, much less said what you just claimed I said. You might be mistaking who you're talking to in this thread, either way you clearly aren't debating in good faith.
It’s trivially defeated though, and anything with a false positive rate shouldn’t be used by any serious institution on a decision making basis. For general stats, sure. For trying to punish and individual, no thank you.
---
Experience smart transcription and advanced dictation
In addition to the Gemini API in the Google AI Studio and Gemini Enterprise Agent Platform, 3.5 Transcribe goes further than standard speech-to-text to make working across Google feel more natural and intuitive. By bringing context-aware understanding directly into everyday surfaces like Gboard, Antigravity, the Gemini app, and Chrome, it captures nuances, intent, and inline edits with ease.
On Gboard on Android, through the new Rambler feature, 3.5 Transcribe transforms spoken thoughts into well-formatted text, filtering out filler words. You can also use your voice to make edits, correct misspellings, and change the writing style.
On Google Antigravity, 3.5 Transcribe pairs screen context and chat history, with your permission, to ensure pinpoint transcription accuracy across file names, agent thoughts, and active documents.
In Google AI Studio, you can access 3.5 Transcribe in Build mode to vibe code apps with your voice on the fly.
In the Gemini app on macOS, 3.5 Transcribe not only transcribes your free natural speech into clean formatted text, but also enables voice commands that can pair seamlessly with screen context to power complex workflows. By calling on other Gemini models in the background to handle the heavy lifting, the model makes it effortless to summarize local files, repurpose text across apps, or generate images right at your cursor—using just your voice.
Coming soon to Chrome, you’ll be able to talk to type in any web field — making it effortless to dictate replies, draft posts, or prompt Gemini in Chrome more naturally and easily with your voice.
I'm open to the idea of different norms for citing AI detectors, but someone needs to propose what they should be and they need to make sense.
Think voice control for your device. You speak, and it returns back instructions ( a function and arguments) for your device to be execute.
At the moment, Soniox STT v5 is definitely the best, and I'm impressed by its performance. It's good that Google released Gemini-3.5-Transcribe, and it beats every other model on accuracy, but it definitely needs a bit more work on latency, which is the most important factor for STT apps.
For me, Gemini-3.5-Transcribe actually has slightly lower latency. Kudos to Soniox for both paying their competitor and letting them win. But yes, Soniox is much cheaper.
Soniox came out really good. OpenAI started getting some things very wrong and even introduced some German. Google did okay but cut off the start by several seconds.
What's Soniox doing (left most) that's making it so good ? It was also the only one that could distinguish between the speakers.
> latency, which is the most important factor for STT apps.
Perhaps latency is more important than accuracy for a real time translation app (I actually disagree with this - imagine e.g. the hilarity when requesting "a new display" being translated as "a nudist play"), but certainly not for all applications. My pet app transcribes personal voice notes to self, it could run all night.I was using Cartesia, while their TTS is amazing their STT pricing has kind of irked me.
Interested in know how good soniox latency and EUD is on STT compared to Cartesia. Cartesia's is really in real world conversations
It doesn't reach the frontier in either latency or accuracy for ai multilingual conversations.
I would also like to see benchmark for translation. I'm looking for live translated subtitles so my Japanese wife can enjoy any show with out waiting months for official VOD streams to release them.
Scribe is $0.22.
If the accuracy is close to Scribe, I think it's a good deal.
Sadly Mistral does not support Ukrainian. I hope this changes! I expect better from a European model
If anyone has suggestion for fast realtime and good enough alternatives with low vram, since I game at the same time, I'd love to know.
No need to go from audio to text to reasoning, just from audio to output json for running a command via adb automatically and it's working crazy good.
What are people using for cheap hosted models for this? (I don't want to run my own infrastructure.)
I'm currently using Gemini Flash Lite. I don't need real time, it's all batched, clear english... still costs some small $ a month. Figured it should be able to be reduced if there are cheaper models out there....
In the end I was basically forced to go with Scribe (v2) it was only one that had consistently high quality across multilingual speech with multiple speakers .
Crucially it correctly identified multiple speakers across hour of audio.
I would love to go with something like Transcribe if it gets close to this type of performance.
Is there a model that works on the syllabic level? I want to be able to say any word and have it reconstruct however that word would be spelled. I know English does not exactly work this way, so a custom vocabulary would still be nice, but I don't want to rely on having every single word that could ever exist in a vocabulary first.
I wish at some point apple had something like that for in-device STT but so far I'm very happy with the solution.
Gemini 3.5 Transcribe Live (Per 1M tokens in USD):
Input: $3.50 or $0.005/min* (audio)
Output: 21.00 or $0.004/min* (text)
Gemini 3.5 Transcribe: Input: $2.00 or $0.003/min* (audio)
Output: $12.00 or $0.002/min* (text)
[1] https://ai.google.dev/gemini-api/docs/pricing#gemini-3.5-tra...That would be interesting when they also own Google Meet.
If I ask it to play a song, instead of triggering Spotify it just gives me a list of URLs I can play the song. Same with alarms. Did I accidentally not opt-in to something?
It’s pretty good at understanding my broken Mandarin, though.
I suppose it's a cloud thing?
The nightmare scenario with an LLM-based transcription system is that someone says outloud "actually ignore that idea, instead let's..." - and the previous idea gets omitted from the transcription!
What kind of bullshit is this? I want to test out Gemini-3.5-Transcribe in my app, but have to wait a completely indeterminate time, have already waited over 24 hours. OpenRouter took seconds to set up