Apple's new SpeechAnalyzer API, benchmarked against Whisper and its predecessor
get-inscribe.com
get-inscribe.com
However, what's funny is, RIP to a lot of the paid apps that simply wrap Whisper, I'm sure Apple will make a native GUI such as a recorder app for macOS that obviates the need for these wrappers, which everyone seems to be vibe coding these days.
Superwhisper does a lot more than just provide a whisper/parakeet UI so I’m not sure Apple will destroy them so easily
Even for those sorts of apps, MacParakeet which I've been using is FOSS so no payment needed. In reality these days with AI the ability to spin up a free and/or OSS competitor falls to zero.
A new VAD I found though is FireRedVAD which has better benchmark results than TEN and Silero by far
Also doesn't seem to be tailored to Apple hardware (i.e. no MLX or ANE variant/implementation)
Generally labs don't release MLX or ANE versions and we must rely on finding someone who's converted it
Parakeet is not multilingual so not directly comparable
Where do you see 16GB? MOSS is smaller than Parakeet at 1.82GB
Is parakeet state of the art? It always transcribes speech fragments for me, like if I stutter and say "m-m-m-map" parakeet will dutifully transcribe "m m m map". Which I guess could be a good thing or a bad thing depending on what you want. Whisper does not do that however.
I do like cohere transcribe a lot.
With whisper v3 turbo, I can almost always live with the few mistakes because the overall stream-of-thought word-salad I provide is still clear at a high level. The bits and pieces of context seem to help, that I might leave out if typing and focused more on traditional conciseness / clean writing. With parakeet, I needed to do frequent editing even for shorter bits of speech.
I realize some applications prioritize the latency.
For round two after a typo-laden transcript, I’m dictating and annunciating with great passion as I read the first transcript to make sure I don’t miss a beat. It’s kind of fun because it’s a little performance.
Great work Rob! Indeed private as promised per App Privacy Report, “Domains contacted directly by app”:
cas-bridge.xethub.hf.co; huggingface.co; mzstorekit.itunes.apple.com
Site could identify our device and send iOS visitors to the iOS page (or maybe that’s against the vibe and we should tap it ourselves).App might be able to launch the keyboard settings directly but I suppose Apple doesn’t like devs using those undocumented URIs (uhg but maybe can understand part of it).
Keyboard, given manual app switchbacking + manual pasting, is less convenient for some of my use cases compared to an action button shortcut. (Reference Spokenly w/Local-Only Mode.)
Separating dictation and the history, and having a syncable scratch pad, are some welcome innovations!
>Gemma E4B … takes up 6-7GB RAM.
Google has a spyware-adjacent dictation app (maybe not really, but they demand to connect to servers after you enable their offline toggle). Just like yours, they thought of a cute name too (Google AI Edge Eloquent). Do you know what language model they install on the iPhone? Not a very good one but I’m sure over the next year or two…
I really wanted to have background audio and make it so the keyboard would directly record audio etc, but my first pass didn't make it through app review (and that was just keeping background audio listening AFTER you'd already started a recording). I could maybe have fought it, but figured if I was already butting up against app review there was little point as they'd likely reject in a future release anyway.
Re: analytics, It is quite weird having no idea how many people are using the app. It does leave you a little blind, but I figure people will get in touch if they have big enough problems with it.
For the Google app, I believe Gemma E2B and E4B both have audio input, so I suspect they're using one of those.
you instead bury an opt-in to automatic telemetry, and let us CC ourselves each ~month when the analytics get sent in to you (so we can verify it's all boring data)...
It's hard for me not to opt in to stuff like that, at least periodically. If I'm opted-in by default? Ehhhhh.... I totally get it, weird "flying blind" (quoting a different dev w/similar philosophy), but I guess I'm still weird myself and hence looking for that autonomy or something? Oh and if somebody makes something opt in, and then certainly if they furthermore stick the option somewhere slightly off the beaten path, that seems pretty darn trustworthy.
(I wonder if we'd have enough public data for a decent statistician to calculate a likely number range of how many users you have by extrapolating from the 1% to 10% or 20% of users who'd chose to opt in...)
PS: "bury" concept likely not actually important :)
Apparently MOSS-Transcribe-Diarize is quite good too as it released only a few days ago.
(and if you think Texans have it bad regarding being understood, try being a Scouser [1])
[0] https://huggingface.co/OpenMOSS-Team/MOSS-Transcribe-Diarize
[1] https://huggingface.co/spaces/OpenMOSS-Team/MOSS-transcribe-...
I use TranscriptionSuite, which is focused on offline transcription, and supports Parakeet as well as whisper and others (https://github.com/homelab-00/TranscriptionSuite) The dev was sweet enough to improve the live mode GUI so that I and some friends can use it when playing on Chinese mmo servers.
What's insane to me is that you have all of these low-quality me-too apps, and literally no one could bother to read the damn Human Interface Guidelines or follow iOS design conventions.
Doing so is literally LESS WORK than trying to make your own custom awful iOS UI.
If you use SwiftUI (the native recommendation by Apple), it severely penalizes you, if you want to paint outside the lines (which is a big reason that I don't use SwiftUI for shipping apps). It's insanely easy to write a native app that is 100% in line with HIG.
- parakeet usually runs on Bfloat16. NPU doesn't support that
- CPU is not as fast as the NPU for these ops on A-series, and even on modern CPUs, there's a latency delay
- Parakeet latency is fine but "fine" may not be good enough for Apple's UX team.
- CPU increases power consumption over dedicated float blocks
So I would say that Parakeet was a non-option for Apple to ship, although it should be in the benchmarks anyways!
https://github.com/FluidInference/FluidAudio/pull/507
That means one hour of audio transcribed in 11.25 and 12.75 seconds.
The Inscribe post doesn't give a speed factor for SpeechAnalyzer. However, this Argmax blog post reports 70:
https://www.argmaxinc.com/blog/apple-and-argmax
Based on that, FluidAudio is ~4.6x and ~4.0x faster.
https://github.com/FluidInference/FluidAudio/pull/799
With the updated PR code, ran a test comparing transcribing (using Parakeet V3) a 1 hr stereo 44.1 kHz mp3 vs the same audio in 16 kHz mono wav format. The result was about 21.3% slower with the mp3 vs the wav, i.e. that's the overhead of decoding + resampling.
Currently the decoding + resampling is done up front. If it was done in a pipelined fashion with the inference, that overhead can be eliminated. This is what I did in a recent app I made:
https://apps.apple.com/us/app/drea-podcast-ad-blocker/id6759...
It uses FluidAudio as well, but I forked it and replaced the audio decoding code to (a) use mpg123 instead of the native Apple API and (b) do audio decoding and inference in a pipelined fashion. These two changes effectively eliminated the overhead. mpg123 is quite a bit faster than the native Apple API at mp3 decoding (has some very optimized arm64 assembly routines), and the pipelining ensures that the inference is never starved by the mp3 decoding.
Contributing this pipelined setup to FluidAudio would be good.
Making the ad-finding cheap enough such that I could make it free turned out to be harder than expected. The main issue you run into is dynamic, location-targeted ads. So I came up with a novel technique that uses Shazam-style audio fingerprints for accurate matching, instead of their normal use case, which is identification. This technique is what allows the ad finding to be very cheap, allowing me to make it free.
The SponsorBlock model would actually not work for podcasts, due to dynamic ads. I.e. the location and content of the ads in episodes these days varies by download location. You need the media to be static, like YouTube, for SponsorBlock model to work. Therefore, using an LLM to find the ads + the fingerprints matching in combination is an efficient technique.
Android has def been the most requested thing thus far haha. It'll be a decent undertaking due to me having written the app fully in Swift, i.e. it'll be a complete rewrite. I'll also need to replace FluidAudio with some good, fast Android equivalent.
The goal of making this app was to create something impressive so that I could get a job. Haven't gotten a job yet, but if and when I do, then I'll have time & resources to think about doing an Android version. Currently a bit stressed and occupied from the job search lol.
A native implementation I had thought of was running the audio through a STT LLM that then detects the ad timestamps then returns that to the UI to block.
I explained how the full system works here (someone emailed me and asked): https://pastebin.com/raw/r2YUEkK5
I use fluidaudiocli and it's unfortunate that it doesn't support streaming (e.g. from a named pipe); that would have been an easy workaround to both the pipelining problem and the faster-decoder problem.
I started using a few open source apps for transcription and eventually subscribed to a paid one...
On paper, it's not hard to compete, but for this use case, a few rough edges make it really frustrating to use. Like a keyboard that sometimes doubles the letter "e"
Automatic dictionary, seamless language switch, no issues with accents, etc... Putting the effort in the last mile makes a world of difference.
If anyone has better options, I'm willing to have a look. The best open source solution I found was Handy, and I currently use Wispr Flow
My initial Mac version actually had three additional steps that you could toggle, obviously at the cost of some speed. That is what the website talks about, although nowadays for my own use I've reduced that to just one step and found that it's pretty great. I've got a new version in test to tidy that up, but still lets you define as many steps as you want.
> unfortuntately parakeet-v3 model doesn;t recieve or output language id https://github.com/NVIDIA-NeMo/Speech/issues/14799#issuecomm...
Also this:
> If you are using Parakeet for English only then you should be using V2. V3 is for several languages and is worse at English only. https://news.ycombinator.com/item?id=48897908
Listen and transcribe felt like the easiest thing to do.
Distavo.com
The source is open for anyone to use, and the builds are in github.
I found quite interesting that claude didn't help too much on how to publish to SetApp until Fable.
(Genuine question - I'm a happy Whisper user but am always looking for improvements).
And if someone were broadly comparing all on-device models (instead of just looking at how this new on-device ones compares to what a specific product uses), Nemotron 3.5's WER are actually a bit higher than what they report for SpeechAnalyzer, for both tests.
Anytime I’m talking to an Indian on the other end, I have to have them repeat everything 2 or 3 times.
- transcribed using MacWhisper.
I'd love this, but updated spotlight did not obviate my need for Raycast. I question Apple's ability to make good software at this point.
Splitting the audio in multiple segments and firing it up without hitting the maximum limit of concurrent decoding streams makes it blazing fast. Fair enough you loose the cut, but it’s good enough for just podcast. In one minute it chews through one hour of audio. This on an iPhone 17 Pro.
Edit: all that said, the app is irrelevant. What I want to say is that live transcripts on iOS using Apples frameworks works very well. Only thing I miss is diarization support.
Its so good that I'm not sure that it's possible to get any better. Speech to text seems like basically a solved problem, if not now then definitely in 5 years. I don't know if any of these speech to text businesses will work in the long run, but for consumers they are great. My guess is the 2030 version of Apple's SpeechAnalyzer will be so good that nobody will need to use 3rd party software.
If I say 'useSuspenseQuery' I want it to come out as useSuspenseQuery not 'use suspense query'. Even if I had to say 'symbol useSuspenseQuery' to give a hint that i'm referencing a symbol, that would be fine.
Looks like Voxtral and Nvidia's Nemotron are best.
[0] https://artificialanalysis.ai/speech-to-text/non-streaming
So I ended up organically testing and ending up with SpeechAnalyzer because it was not only fast and accurate enough, but you also see live results as you talk. It also has speaker identification and people can register their voices. And it does all processing on device.
It also had the best model for Indian accented English, given she lives in India.
So I was quite impressed, but the holy grail to me is transcription that does speaker identification but also works in a standard family conversation, where multiple people interrupt each other all the time.
I will say though, I'm really curious as to what Claude Code Desktop uses for their voice mode, because it seems even better than Apple's, and it provides realtime feedback. Maybe they're using apple's model?
If I start typing and the existing text is in Spanish, then a sensible default is to select the Spanish keyboard I have installed and let me adjust otherwise.
App developers should also be allowed to supply mini-dictionaries within a context to allow autocorrect to work correctly in that context, so for example in this thread [SpeechAnalyzer, API, Whisper, Parakeet, Nemotron] should be supplied so that these terms are autocorrected.
It can struggle with proper nouns but will return something phonetically similar.
My main gripe is that it requires a separate model download per language. I understand the why they did this (to save disk space). But it makes multi-lingual audio hard to transcribe unless you know ahead of time the languages in the audio.
As an app developer the biggest win from using Apple's model is I don't have to bundle it in my app so my app looks much smaller. If a user has many transcription apps each one could have their own model. If Apple's model is used only one copy is needed.
And yea, Nvidia's Parakeet v3 is good enough for my own just local transcription most of the time.
When I need local transcription to be more reliable and I don't have the energy to proof read a long ramble, I still often just pop open chatGPT, dictate, cut, paste.
But we're pretty much already to the point where local transcription models can replace cloud ones for personal use. They're still a bit rough around the edges in terms of polish and latency, but plenty of people are fine with that to avoid yet another app subscription and not having to worry about wondering what's potentially happening with their data.
It's very welcome, to start with this.
However, just from this heading, I knew with 90% certainty that the post was AI generated.
I wonder if I would have less of an issue with this if such blog posts would just start with: “I asked $model $model-version the following: $prompt. Here is what I got.”
> The new API cuts word error rate by 3.5 to 4x on the same audio: from 9.02% to 2.12% on clean speech
Shouldn't they have said "cuts error rate by 78%" or something?
- it implies that error could be increased n-times, but a 15x _increase_ in 9% error would be an error rate of 135%, which is nonsensical.
- a reduction from 90% error to 20% error is clearly a bigger improvement in rightness to a reduction from 9% to 2%. One is “almost all wrong to almost all right”, the other is “more right”, but they are both a 4.5x reduction in error which means that the 4.5 quantity doesn’t have a constant meaning.
The answer is something like log odds ratios, but that introduces the additional need for a reader to know what that is, and that would be unusual.
Im looking for the same experience I have when talking to chatGPT. As for past two years or more talking to GPT within it's app and on my iPhone Pro Max 15 it runs smooth as butter :-). This is the experience I was and still am hoping with Apple, but Im thinking all the extra layers of privacy and security might be slowing them down?
Overall, Apple who is suing Open AI should just buy them and let me have the best conversational AI out there baked into my old ass iPhone. Because as so far the new Siri on my old phone (tho again GPT works great talking to it and for years) doesnt come close. It's the same old "Could you try that again," Siri. BOO!!!
1. In Shortcuts app, make shortcut named "AI Voice Mode" (or whatever you want, YMMV)
2. Set it to run the ChatGPT action "Voice Mode" (requires at least the minimum paid tier, I think)
3. To trigger, say "Hey Siri, AI Voice Mode" (or whatever you called the shortcut)
This is a pretty slick integration, but yeah, if it were baked in that would be all the better.Thanks for the tip and if Im not mistaken it's similar to asking Siri to ask chatGPT to ask XYZ?
Effectively, it sort of does that, but really it just listens to the wakeword and opens/switches to the requested app & modality.
FWIW, I get a very different functional result using the Shortcut method vs. asking Siri to delegate natively. To compare, I asked Siri (non-beta here) now to "ask ChatGPT <x>" and I got a top-card with some fairly low quality SEO-ranked weblinks.
New Siri is impressive in that it answers satisfactorily now 80% of the time vs 10% with old Siri.
But it’s slow as shit. GPT, Claude, and Gemini can answer me in 5-10 seconds. Google AI Mode can answer in 2 seconds.
New Siri usually takes 25 seconds to respond to me. This morning it timed out (with strong network connection) when asked a simple multiplication question.
Apple would never do that, if anything they did not offer their Siri with the most advanced AI on iPhone 16 Pro Max, which is one year-old only.
Supports SRT/TXT/VTT or JSON-with-optional-word-level-timestamps output and progress meter.
Also it can transcribe live system audio.
In my own tests a few months ago, it was faster than both small/large Whisper models I compared it to with accuracy competitive with both of them (each model had different quirks).
Still sucks
This is useless test and benchmark when you have these day Whisper-V3-Large and Whisper V3-Turbo that you can faster than realtime on 5 years old macbook on apple sillicon (ANE). They didn't even compared to parakeet v2 or parakeet v3. And only english language...
Kind of a bait and switch. How can we test the product with such short time limits and what, exactly are you offering if all the processing is done on device by Apple?
https://github.com/cjpais/Handy
It's great for me when configured to use Parakeet v3.
Cloud models are usually protected by trade secret laws, leaking them would get you in trouble. However if the model is made available publicly, as long as you don't break the law to get them, anything after that would be fair game unless Apple can prove that humans have significant authorship over the weights, which hasn't been tested and is a significant burden to prove/disprove.
The Jedi Hand Wave-y nature of the way people talk about AI these days is going to make reigning in the AI superpowers nearly impossible. Because there are people out here who believe models of this quality are easily replicated or reverse engineered. Neither is really doable on any reasonable timeline by people who are not AI experts. Real AI experts. Not TF/PyTorch monkeys or Agent Slop Slingers.
And those people are already highly incentivized to not make anything performing better than SOTA models open source.
Edit: Getting downvoted by Apple fanboys for telling the truth is a badge of honor.
Apple published no accuracy numbers for SpeechAnalyzer (or for SFSpeechRecognizer, ever, as far as I can tell), so the migration question has been guesswork. Short version: the new API cuts WER 3.5-4x vs the old one (2.12% vs 9.02% on test-clean), and it also beat Whisper Small on both splits at about 3x the speed. The old API came in last on clean speech, behind even Whisper Tiny.
On "why should I trust a vendor benchmark": the Whisper column reproduces OpenAI's published LibriSpeech WERs within +0.11 to +0.42 on all six measurements (same corpus, same normalizer, same scorer for every engine), and the raw per-utterance transcripts are downloadable from the article if anyone wants to rescore with their own normalizer.
Limitations worth stating up front: English only, read speech rather than meeting audio, one machine. Precise per-engine timing isn't in the article yet because the accuracy runs shared the machine with a dev workload; WER is load-independent, timing isn't.
Two things that might interest people migrating: SFSpeechRecognizer sends audio to Apple's servers unless you set requiresOnDeviceRecognition, and with SpeechAnalyzer, finishing your input stream is not enough to end a session. If you never call finalizeAndFinishThroughEndOfInput(), the results sequence never terminates and your await hangs forever. I found that one because it was shipping in my own app.
Happy to answer questions about the harness or the normalizer.
On the more cutting edge front, Granite Speech 4.1 has proven to be a reliable workhorse for me, but it is larger than Parakeet. Cohere Transcribe is interesting, but how strong it is seems to vary more from task to task.
Parakeet Unified 0.6B came out a few months ago, combining both online streaming and offline transcription into one model, and that is one that I need to test more, but it seems promising.
As others have mentioned macOS 27/iOS 27 is supposed to have a new model, particularly on devices with 12GB of RAM or more. I have not actually seen the option to enable that new model yet, though, despite being on the beta on a device that meets the requirements. Maybe a benchmark would reveal that it is already active?
Also, just out of curiosity, seems like everyone and their mother is making Whisper wrappers, how is your app different?
MOSS-Transcribe-Diarize
> What this means if you just want good transcription
> If you are on a current iPhone or Mac, the best on-device transcription engine for English is already in the operating system, and the private option is no longer the compromise option
How can you be sure this isn't leaking data or metadata to Apple? Can Apple really be trusted?
https://alternativeto.net/software/little-snitch/
https://www.g2.com/products/little-snitch/competitors/altern...
There are many alternatives for trying to find out what’s going on. If you don’t want to bother, and most people don’t, well, what else is there to say?
It is generally a good idea to know what software is phoning home, if you can pinpoint it.
If you have any software recommendations, I’d be happy to know.
> If you are on a current iPhone or Mac
Presumably if you don't trust apple you wouldn't purchase their products and even if you were for example forced to use it via work or something you wouldn't use this feature ... so it doesn't really change the calculus as presented by this article - IF you ALREADY HAVE a MODERN Mac (and trust apple) this is your best option
If trust Apple, then no need for privacy from Apple
The appeal is that users only have to download it once across all apps that use it. Instead of convincing a user to give a couple gigs for just your one app