MacWhisper: Transcribe audio files on your Mac
goodsnooze.gumroad.com
goodsnooze.gumroad.com
Sometimes I'll send a mp3 or mp4 video through it and use the resulting transcript directly.
Other times I'll run a second step through https://claude.ai/ (because of its 100,000 token context) to clean it up. My prompt for that at the moment is:
> Reformat this transcript into paragraphs and sentences, fix the capitalization and make very light edits such as removing ums
That's often not necessary with Whisper output. It's great for if you extract captions directly from YouTube though - I wrote more about that here: https://simonwillison.net/2023/Aug/6/annotated-presentations...
I was working on a native version in the form of a taskbar app with customizable prompt and all. However I quickly realized that the behaviors I want the app to do require a bunch of accessibility permissions that would block it from the app store and require more setup steps.
Would anybody still find that useful?
Edit: looks like Mac native runs locally only with dapple Silicon, and maybe has a more limited geographic/linguistic reach?
My old Sun Ultra 40 M2 had a ton of electrical noise on my headphone jack, and I could def. tell when the CPU was busy from what I was hearing.
- The accuracy is decent but the vocabulary is very limited. With Whisper, you can customize the prompt to include industry-specific terms and acronyms.
See my example from the repo. Apple recognizes:
> Popular Linux distributions include Debby and Fed or Linux, and do Bantu. You can use windowing systems such as X eleven or Weiland with a desktop environment like KD plasma.
Whisper recognizes:
> Popular Linux distributions include Debian, Fedora, Linux, and Ubuntu. You can use windowing systems such as X11 or Wayland with a desktop environment like KDE Plasma.
Popular Lennox distributions include Debbie and Fedora, Lennox, and Beau you can use windowing system such as X11 or Whelan with a desktop environment like Katy plasma
Popular Linux distributions include Devion fedora Linux and a bunch to you can we use when doing system such as excellent Wayland with a desktop environment like Katie plasma.
Which behaviours specifically?
Personally, I wouldn't worry too much about the App Store. I'm distributing Enso (http://enso.sonnet.io) via gumroad.com, and people download/pay for it. I think it's easier than using the App Store Connect route anyway.
Here's a good intro: https://rambo.codes/posts/2021-01-08-distributing-mac-apps-o...
Thanks for the info about your app. It looks great!
You can do that using something like:
var reactOnOptionKeyHeld: DispatchWorkItem? { didSet { oldValue?.cancel() } }
NSEvent.addGlobalMonitorForEvents(matching: .flagsChanged) { (event) in
guard event.modifierFlags == [.option] else {
reactOnOptionKeyHeld = nil
return
}
reactOnOptionKeyHeld = DispatchWorkItem {
// start recording
}
// schedule to run if held for at least 1 second
DispatchQueue.main.asyncAfter(deadline: .now() + 1, execute: reactOnOptionKeyHeld!)
}
I see you're using Python with pynput though, which is creating a full key listener so I guess that is why you need the permissions.- The store itself is convenient for browsing and discovery.
- I don't need to do any kind of background checking on the developer prior to running their app
- Similarly they don't get access to my credit card details, so I don't have to be concerned about them storing it incorrectly or abusing it later
- It's also easier for me to pay them, as often foreign transactions are blocked despite me specifically approving them.
- I don't need to hand over my email or other personal details, small developers seem really bad at storing information and CC details properly. I use custom emails for everyone and numerous times I've seen my data on-sold, stolen by ex-employees, or simply thefted from the company by hacking groups.
- I don't have to do any reviews when there is an update, I can just accept it knowing that the developer is still trusted and that the project hasn't been hijacked such as the numerous painful times that popular open source projects have had malware snuck into them.
- If the app doesn't properly do what claims (or at least what I thought it would), it's a few clicks to get a refund from Apple.
- Apple carrot and stick developers to keep their apps up to date with the system. First carrot, and eventually the stick (delisting).
Some may argue that some of these things can still happen with an app store, but it's demonstrably less and there are processes in place to deal with that.
It's not a popular opinion, especially on HN, but there are plenty of developers who, whether through frustration or dealing with pedantic/rude* customer requests, treat their customers like shit, the store prevents that.
The app store ain't perfect, and there are plenty of functions which apps can't have if they're sold through the app store, but for its flaws it's trustworthy and helps me utilise a far larger number of utilities than I'd normally be comfortable with personally maintaining.
* I sit on enough discords to see how breathtakingly rude and demanding some users are without noticing it, even for totally free software.
SRC_DIR="$( cd "$( dirname "${BASH_SOURCE[0]}" )" >/dev/null 2>&1 && pwd )"
python3.10 "$SRC_DIR/whisperer.py" $@
Also did you solve the python app icon from bouncing all the time?
I have not solved the bouncing icon. That's one of the reasons it needs to be rewritten in Swift!
On top of that, constantly sending data to google would have chewed a ton of battery compared to the "activation word" style solutions ("ok google/siri") that can be done on-device. The power for on-device processing was obviously going to come down over time, while wireless is much more governed by the laws of physics, and connectivity power budgets haven't gone down nearly as much over time. I am pretty sure there is a fundamental asymptotic limit for this, governed by Shannon entropy limit/channel width and power output. In the presence of a noise floor of X, for a bandwidth of Y, you simply cannot use less than Z total power for moving a given amount of data.
BTLE is really the first game-changer (especially if you are hooking into a broad network of receivers like apple does with airtags) but even then you are not really breaking this rule - you are just transmitting less often, and sending less data. It's just a different spot on the curve that happens to be useful for IOT. If you are, say, doing a keyboard over BTLE where the duty cycle is higher, the power will be too. Applications that need "100% duty cycle"/"interactive" (reachable at any time with minimal latency") still have not improved very much.
In hindsight I guess the answer would have been writing a mobile app that ties into google/siri keywords and actions, and letting the phone be the UI and only transmit BT/BTLE to the device. But BTLE hadn't hit the scene back then (or at least not nearly to the extent it has now) and I was less experienced/less aware of that solution sapce.
[0] https://github.com/guillaumekln/faster-whisper/discussions/3...
Not a deal breaker, but it was last updated 3 months ago and lacks a few QoL features of MacWhisper. Jordi is frequently pushing updates to MacWhisper: https://nitter.net/jordibruin/status/1692133387299864638
This project has been alright for transcribing audio with speaker diarization. A big finicky. The OpenAI model is better than other paid products(Descript, Riverside) so I’m looking forward to trying MacWhisper.
I ask because I asked a friend to record a (for fun) lecture I couldn't attend, and unfortunately the speech audio levels are quite low, and I'm trying to figure out how to extract as much info as possible so I can hear it. If I could add context to the transcriber like "This is about the Bronze Age collapse and uses terminology commonly used in discussions on that topic", it might be even more useful.
Has anyone else tried to do something similar? How did you achieve it?
Any chance there's an iOS version of this coming down the pike? It would be great to have a voice-based note taking app that you can use when you are driving or walking and you don't want to type into your phone, but you just want to save that thought you just had somewhere by quickly dictating it, and having it accessible as text later.
Granted my use cases are not high volume or frequent but being able to take output from Whisper and pipe it to other models has been very powerful for me. It is also amazing how good the quality of Whisper is when handling non English audio.
We added LocalAI (https://localai.io) support to LLMStack in the last release. Will try to use whisper.cpp and see how that compares for my use cases.
The developer Jordi has a great speech online about product development.
Supports Tiny (English Only), Tiny, Base, Small, Medium and Large models
Translate audio file into another language through Whisper (use the Medium or Large models, the results will not be perfect and I'm working on more advanced ways to do this)
Whisper itself is open source, and so is that implementation, the OpenAI endpoint is merely a convenience to those who don't wish to host a Whisper server themselves, deal with batching, renting GPUs etc. If you're making a commercial service based on Whisper, the API might be worth it for the convenience, but if you're running it personally and have a good enough machine (an M1 MacBook Air will do), running it locally is usually better.
Web browsers are mostly free and don't try to upsell you to a Pro paid version. The MacWhisper author deserves to be compensated for their work, so I'm not objecting to the existence of a paid version. This feels like yet another relatively low value freemium/upsell wrapper in the Mac shareware ecosystem to me.
I'm probably wrong and there's a real population that benefits from this work, clearly some folks perceive it as useful enough to pay for it and I'm just not in that audience to see it.
I think part of what rubs me the wrong way about this is that it feels to me like commercial freeloading due to the thinness of the commercialized wrapper around a free/open core in this case (whisper model + code); it feels ethically questionable unless the author contributes back some portion of the proceeds to research in some way -- I didn't see evidence of that. I'm probably being naive here, happy to have a less snarky discussion about it though.