However, what's funny is, RIP to a lot of the paid apps that simply wrap Whisper, I'm sure Apple will make a native GUI such as a recorder app for macOS that obviates the need for these wrappers, which everyone seems to be vibe coding these days.
However, what's funny is, RIP to a lot of the paid apps that simply wrap Whisper, I'm sure Apple will make a native GUI such as a recorder app for macOS that obviates the need for these wrappers, which everyone seems to be vibe coding these days.
(and if you think Texans have it bad regarding being understood, try being a Scouser [1])
[0] https://huggingface.co/OpenMOSS-Team/MOSS-Transcribe-Diarize
[1] https://huggingface.co/spaces/OpenMOSS-Team/MOSS-transcribe-...
I use TranscriptionSuite, which is focused on offline transcription, and supports Parakeet as well as whisper and others (https://github.com/homelab-00/TranscriptionSuite) The dev was sweet enough to improve the live mode GUI so that I and some friends can use it when playing on Chinese mmo servers.
Is parakeet state of the art? It always transcribes speech fragments for me, like if I stutter and say "m-m-m-map" parakeet will dutifully transcribe "m m m map". Which I guess could be a good thing or a bad thing depending on what you want. Whisper does not do that however.
I do like cohere transcribe a lot.
With whisper v3 turbo, I can almost always live with the few mistakes because the overall stream-of-thought word-salad I provide is still clear at a high level. The bits and pieces of context seem to help, that I might leave out if typing and focused more on traditional conciseness / clean writing. With parakeet, I needed to do frequent editing even for shorter bits of speech.
I realize some applications prioritize the latency.
For round two after a typo-laden transcript, I’m dictating and annunciating with great passion as I read the first transcript to make sure I don’t miss a beat. It’s kind of fun because it’s a little performance.
Great work Rob! Indeed private as promised per App Privacy Report, “Domains contacted directly by app”:
cas-bridge.xethub.hf.co; huggingface.co; mzstorekit.itunes.apple.com
Site could identify our device and send iOS visitors to the iOS page (or maybe that’s against the vibe and we should tap it ourselves).App might be able to launch the keyboard settings directly but I suppose Apple doesn’t like devs using those undocumented URIs (uhg but maybe can understand part of it).
Keyboard, given manual app switchbacking + manual pasting, is less convenient for some of my use cases compared to an action button shortcut. (Reference Spokenly w/Local-Only Mode.)
Separating dictation and the history, and having a syncable scratch pad, are some welcome innovations!
>Gemma E4B … takes up 6-7GB RAM.
Google has a spyware-adjacent dictation app (maybe not really, but they demand to connect to servers after you enable their offline toggle). Just like yours, they thought of a cute name too (Google AI Edge Eloquent). Do you know what language model they install on the iPhone? Not a very good one but I’m sure over the next year or two…
I really wanted to have background audio and make it so the keyboard would directly record audio etc, but my first pass didn't make it through app review (and that was just keeping background audio listening AFTER you'd already started a recording). I could maybe have fought it, but figured if I was already butting up against app review there was little point as they'd likely reject in a future release anyway.
Re: analytics, It is quite weird having no idea how many people are using the app. It does leave you a little blind, but I figure people will get in touch if they have big enough problems with it.
For the Google app, I believe Gemma E2B and E4B both have audio input, so I suspect they're using one of those.
you instead bury an opt-in to automatic telemetry, and let us CC ourselves each ~month when the analytics get sent in to you (so we can verify it's all boring data)...
It's hard for me not to opt in to stuff like that, at least periodically. If I'm opted-in by default? Ehhhhh.... I totally get it, weird "flying blind" (quoting a different dev w/similar philosophy), but I guess I'm still weird myself and hence looking for that autonomy or something? Oh and if somebody makes something opt in, and then certainly if they furthermore stick the option somewhere slightly off the beaten path, that seems pretty darn trustworthy.
(I wonder if we'd have enough public data for a decent statistician to calculate a likely number range of how many users you have by extrapolating from the 1% to 10% or 20% of users who'd chose to opt in...)
PS: "bury" concept likely not actually important :)
Apparently MOSS-Transcribe-Diarize is quite good too as it released only a few days ago.
Superwhisper does a lot more than just provide a whisper/parakeet UI so I’m not sure Apple will destroy them so easily
Even for those sorts of apps, MacParakeet which I've been using is FOSS so no payment needed. In reality these days with AI the ability to spin up a free and/or OSS competitor falls to zero.
A new VAD I found though is FireRedVAD which has better benchmark results than TEN and Silero by far
Also doesn't seem to be tailored to Apple hardware (i.e. no MLX or ANE variant/implementation)
Generally labs don't release MLX or ANE versions and we must rely on finding someone who's converted it
Parakeet is not multilingual so not directly comparable
Where do you see 16GB? MOSS is smaller than Parakeet at 1.82GB
What's insane to me is that you have all of these low-quality me-too apps, and literally no one could bother to read the damn Human Interface Guidelines or follow iOS design conventions.
Doing so is literally LESS WORK than trying to make your own custom awful iOS UI.
If you use SwiftUI (the native recommendation by Apple), it severely penalizes you, if you want to paint outside the lines (which is a big reason that I don't use SwiftUI for shipping apps). It's insanely easy to write a native app that is 100% in line with HIG.
I started using a few open source apps for transcription and eventually subscribed to a paid one...
On paper, it's not hard to compete, but for this use case, a few rough edges make it really frustrating to use. Like a keyboard that sometimes doubles the letter "e"
Automatic dictionary, seamless language switch, no issues with accents, etc... Putting the effort in the last mile makes a world of difference.
If anyone has better options, I'm willing to have a look. The best open source solution I found was Handy, and I currently use Wispr Flow
My initial Mac version actually had three additional steps that you could toggle, obviously at the cost of some speed. That is what the website talks about, although nowadays for my own use I've reduced that to just one step and found that it's pretty great. I've got a new version in test to tidy that up, but still lets you define as many steps as you want.
> unfortuntately parakeet-v3 model doesn;t recieve or output language id https://github.com/NVIDIA-NeMo/Speech/issues/14799#issuecomm...
Also this:
> If you are using Parakeet for English only then you should be using V2. V3 is for several languages and is worse at English only. https://news.ycombinator.com/item?id=48897908
- parakeet usually runs on Bfloat16. NPU doesn't support that
- CPU is not as fast as the NPU for these ops on A-series, and even on modern CPUs, there's a latency delay
- Parakeet latency is fine but "fine" may not be good enough for Apple's UX team.
- CPU increases power consumption over dedicated float blocks
So I would say that Parakeet was a non-option for Apple to ship, although it should be in the benchmarks anyways!
https://github.com/FluidInference/FluidAudio/pull/507
That means one hour of audio transcribed in 11.25 and 12.75 seconds.
The Inscribe post doesn't give a speed factor for SpeechAnalyzer. However, this Argmax blog post reports 70:
https://www.argmaxinc.com/blog/apple-and-argmax
Based on that, FluidAudio is ~4.6x and ~4.0x faster.
https://github.com/FluidInference/FluidAudio/pull/799
With the updated PR code, ran a test comparing transcribing (using Parakeet V3) a 1 hr stereo 44.1 kHz mp3 vs the same audio in 16 kHz mono wav format. The result was about 21.3% slower with the mp3 vs the wav, i.e. that's the overhead of decoding + resampling.
Currently the decoding + resampling is done up front. If it was done in a pipelined fashion with the inference, that overhead can be eliminated. This is what I did in a recent app I made:
https://apps.apple.com/us/app/drea-podcast-ad-blocker/id6759...
It uses FluidAudio as well, but I forked it and replaced the audio decoding code to (a) use mpg123 instead of the native Apple API and (b) do audio decoding and inference in a pipelined fashion. These two changes effectively eliminated the overhead. mpg123 is quite a bit faster than the native Apple API at mp3 decoding (has some very optimized arm64 assembly routines), and the pipelining ensures that the inference is never starved by the mp3 decoding.
Contributing this pipelined setup to FluidAudio would be good.
Making the ad-finding cheap enough such that I could make it free turned out to be harder than expected. The main issue you run into is dynamic, location-targeted ads. So I came up with a novel technique that uses Shazam-style audio fingerprints for accurate matching, instead of their normal use case, which is identification. This technique is what allows the ad finding to be very cheap, allowing me to make it free.
The SponsorBlock model would actually not work for podcasts, due to dynamic ads. I.e. the location and content of the ads in episodes these days varies by download location. You need the media to be static, like YouTube, for SponsorBlock model to work. Therefore, using an LLM to find the ads + the fingerprints matching in combination is an efficient technique.
Android has def been the most requested thing thus far haha. It'll be a decent undertaking due to me having written the app fully in Swift, i.e. it'll be a complete rewrite. I'll also need to replace FluidAudio with some good, fast Android equivalent.
The goal of making this app was to create something impressive so that I could get a job. Haven't gotten a job yet, but if and when I do, then I'll have time & resources to think about doing an Android version. Currently a bit stressed and occupied from the job search lol.
A native implementation I had thought of was running the audio through a STT LLM that then detects the ad timestamps then returns that to the UI to block.
I explained how the full system works here (someone emailed me and asked): https://pastebin.com/raw/r2YUEkK5
I use fluidaudiocli and it's unfortunate that it doesn't support streaming (e.g. from a named pipe); that would have been an easy workaround to both the pipelining problem and the faster-decoder problem.
And if someone were broadly comparing all on-device models (instead of just looking at how this new on-device ones compares to what a specific product uses), Nemotron 3.5's WER are actually a bit higher than what they report for SpeechAnalyzer, for both tests.
I'd love this, but updated spotlight did not obviate my need for Raycast. I question Apple's ability to make good software at this point.
Anytime I’m talking to an Indian on the other end, I have to have them repeat everything 2 or 3 times.
- transcribed using MacWhisper.
Listen and transcribe felt like the easiest thing to do.
Distavo.com
The source is open for anyone to use, and the builds are in github.
I found quite interesting that claude didn't help too much on how to publish to SetApp until Fable.
(Genuine question - I'm a happy Whisper user but am always looking for improvements).