Show HN: I made a free transcription service powered by Whisper AI
freesubtitles.ai
freesubtitles.ai
I feel compelled to point out whisper.cpp. It may be cheaper for the author but is relevant for others.
I was running whisper on a gtx 1070 to get decent performance; it was terribly slow on M1 Mac. Whisper.cpp has comparable performance to the 1070 while running on M1 CPU. It is easy to build and run and well documented.
https://github.com/ggerganov/whisper.cpp
I hope this doesn’t come off the wrong way, I love this project and I’m glad to see the technology democratized. Easily accessible high-quality transcription will be a game changer for many people and organizations.
https://scotts-podcasts.s3.amazonaws.com/temp/whisper/Intern...
Not sure if there's a use for this that's not me, but I like the idea of having subtitles for a podcast I'm listening to.
I can usually get an (unverified) 1x RTX 3090 instance for about $0.10/hr, and that processes audio at something like 1.5X speeds. Unverified instances do crash once in a while, but as long as you back up the output every few hours, it's fine, you just set up a new one in case something happens. I wouldn't use this for confidential company meetings, but it's good enough for podcasts, Youtube Videos and other public or semi-public stuff.
I’m in love with the idea of pressing button on my Lock Screen and getting a perfect transcription in my inbox.
Also, just added emoji summarization in email subject, a small visual reminder of what your memo was about.
I hope this is useful to someone!
Anyway, I'd love to get it running well on g5g, but they seem extremely temperamental. If anybody has, please let me know your secret. :)
I imagine future models will allow the user to input some context to disambiguate. Like if I could give it the audio along with the context "Seattle Fire Department and EMS radio traffic", it would bias towards the type of things you'd likely hear on such a channel.
But, for me, the English model works really well. Using the 'large' model works about perfectly for me, I can't think of anything I thought the large model got too badly wrong, is that the model you tried?
I'd need to:
- Wait for a MED6 or AID response code to come across the live event stream
- Listen to the radio chatter to see if it was a pedestrian getting hit (use GPT3 on the transcription to determine if the text was about a ped/cyclist getting hit)
- Maybe also correlate to SPD logging a 'mvc with injuries' at the same location
Did you try a telephony oriented model like aspire or similar? They’re trained on sort-of 8 kHz audio and might work better.
I tried something similar for my SDR feeds and gave up because it’s just too challenging and niche - the sampling, the jargon, the 10 codes, the background noise, static on analog systems/drop outs on digital systems, rate of speech, etc all contribute to very challenges issues for an ML model.
Is the reduced bandwidth really the most significant problem? Naively I'd think everything else you mentioned would matter a lot more, I'm curious how much you experimented with that specifically.
I'm not sure which contributes most but I know from my prior experiences with ASR for telephony even clean speech on pristine connections does much worse with models trained on 16 kHz being fed native 8 kHz audio that gets resampled.
I've done some early work with Whisper in the telephony domain (transcribing voicemails on Asterisk and Freeswitch) and the accuracy already seems to be quite a bit worse.
Hah, in all seriousness I'm more of a practitioner in this space. If this was something I absolutely needed to get done who knows where it would have went. For a little side hacking project once I encountered these issues I moved on - back in the day expectations were lower for telephony and the 8 kHz aspire models and kaldi were adequate to get that "real work" done.
I'm still getting spun up on this but base delivers a pretty impressive 5-20x realtime on my RTX 3090. I haven't gotten around to trying the larger models and with only 24GB of VRAM I'm not sure what kind of success I'll have anyway...
In my case the goal was to actually generate tweets based on XYZ. As I've already said there were serious technical challenges so I abandoned the project but I was also a little concerned about the privacy, safety, etc issues of realtime or near-realtime reporting on public safety activity. I also streamed to broadcastify and it really seems like they insert artificial delay because of these concerns.
For a 1:17 file it takes:
6s for base.en, I think 2s to load the model based on the sound of my power supply.
33s for large, I think 11s of which is loading the model.
Varies a lot with how dense the audio file is, this was me giving a talk so not the fastest and quite clean audio.
While I saw near perfect or perfect performance on many things with smaller models, the largest really are better . I'll upload a gist in a but with Rap God passed through base.en and large.
edit -
Timings (explicitly marked as language en and task transcribe):
base.en => 23s
large => 2m50
Audio length 6m10
Results (nsfw, it's Rap God by Eminem): https://gist.github.com/IanCal/c3f9bcf91a79c43223ec59a56569c...
Base model does well, given that it's a rap. Large model just does incredibly, imo. Audio is very clear, but it does have music too.
Hah, I love that - "benchmark by fan speed".
Good to know - I've tried large and it works but in my case I'm using whisper-asr-webservice[0] which loads the configured model for each of the workers on startup. I have some prior experience with Gunicorn and other WSGI implementations so there's some playing around and benchmarking to be done on the configured number of workers as the GPU utilization of Whisper is a little spiky and whisper-asr-webservice does file format conversion on CPU via ffmpeg. Default was two workers, is now one but I've found as many as four with base can really improve overall utilization, response time, and scale (which certainly won't be possible with large).
OPs node+express implementation shells out to Whisper which gives more control (like runtime specification of model) but almost certainly has to end up slower and less efficient in the long run as the model is obviously loaded from scratch on each invocation. I'm front-ending whisper-asr-webservice with traefik so I could certainly do something like having two separate instances (one for base, another for large) at different URL paths but like I said I need to do some playing around with it. The other issue is if this is being made available to the public I doubt I'd be comfortable without front-ending the entire thing with Cloudflare (or similar) and Cloudflare (and others) have things like 100s timeouts for final HTTP response (Websockets could get around this).
Thanks for providing the Slim Shady examples, as a life-long hip hop enthusiast I'm not offended by the content in the slightest.
I was initially going to use Azure Cognitive Services and train it on a small amount of test data, after Whisper released for free I use Whisper + openai GPT-3 trained to fix the transcription errors by 1) taking a sample of transcripts by Whisper 2) fixing the errors and 3) fine-tuning GPT-3 by using the unfixed transcriptions as the prompt and the corrected transcripts as the result text.
Whisper with the --initial_prompt containing industry jargon plus training GPT-3 to fix the transcription errors should be nearly as accurate as using a custom-trained model in Azure Cognitive Services but at 5-10% of the cost. Biggest downside is the amount of labor to set that up, and the snail's pace of Whisper transcriptions.
https://github.com/ggerganov/whisper.cpp
This is a C/C++ version of Whisper which uses the CPU. It's astoundingly fast. Maybe it won't work in your use case, but you should try!
I am no-longer involved in the project but you're welcome to contact the CTO if you're curious how it worked:
Read.ai - https://www.read.ai/transcription
Provides transcription & diarization and the bot integrates into your calendar. It joins all your meetings for zoom, teams, meet, webex, tracks talk time, gives recommendations, etc.
It’s amazing how quickly this space is moving. Particularly, with the increase in remote work. Soon you’ll be able to search all your meetings and find exactly when a particular topic was discussed! It’s exciting.
Dictation is also great when writing in a foreign language: I speak German ok-ish, writing is harder. Dictation helps writing more correct German.
Also, labels’ [for] attributes are all "file" instead of "language" and "model", so all labels trigger the file selection dialog on click :-)
Just given your site a try, nicely done. One feedback - would be great to have a progress indicator on the processing page, I have no idea what stage it's at or how much longer I need to wait.
It should show the data via processing.. it's setup to just take whatever stdout/stderr comes back from Whisper and send it directly to the frontend via websockets, I'm surprised you got stuck there :thinking:
I'm going to hope/assume you're doing some sort of sanitisation on those inputs.
Additionally, wouldn't you lose the language detection that's done for no language input? (IIRC, it uses the first 30 seconds to detect language if you don't specify one)
Well those inputs should all error unless they are a valid value.
Yes if nothing is input it will automatically detect the language based on the first 30s of input
Wait what. Not being able to safely restart the server sounds like a disaster waiting to happen.
> I'm just running this off of a 2x RTX A6000 server on Vast.ai at the moment, about $1.30/h
whether that's a lot is a matter of perspective
users would lose the session and have to start over, not the end of the world
from https://github.com/openai/whisper#available-models-and-langu...
It should be all local until it needs information from the internet...
Curious if there was a benefit to using whisper over something like vosk which can transcribe on mobile device pretty decently.
Whisper has other interesting functionality but for straight transcription it seems a bit heavy. Still learning about it and putting it through its paces.
https://alphacephei.com/nsh/2022/10/22/whisper.html
In general, Whisper is more accurate but much more resource heavy. Vosk runs on single core while Whisper needs all CPU cores.
Accuracy difference for clean speech between Vosk-small and Whisper tiny is 2-3% absolute, 20% relative. Not sure how important is it, I would claim it is not that critical.
Numbers there are for original Whisper. Whisper.cpp recommended here is actually 10% worse than stock Whisper for speed considerations. Not that simple.
Vosk is streaming design, you get results with minimum latency of 200ms. Whisper requires you to wait for significant amount of time. If you refactor Whisper for lower latency you will loose a lot of accuracy advantage. Latency is very important for interactive applications like assistants.
Whisper is multilingual and has punctuation, that is a clearly a good advantage. It also can use context properly improving for long recordings.
So on mobile Vosk is still a viable option actually as many others mobile-focused engines.
For server based transcription Whisper is certainly better. But not much better than Nvidia Nemo for example. Not that much publicity for the former though.
I setup a whisper-asr-api backend this week with gobs of CPU and RAM and an RTX 3090. I’d be interested in making the API endpoint available to you and working on the overall architecture to spread the load, improve scalability, etc.
Let me know!
Open an issue on the Github repo and we can collab for sure!: https://github.com/mayeaux/generate-subtitles/issues
Through a series of events I'm in the beneficial position of my hosting costs (real datacenter, gig port, etc) being zero and the hardware has long since paid for itself. I'm almost just looking for ways to make it more productive at this point.
Anyway, I'll be making an issue soon!
I also implemented some anti-abuse-ish features between traefik and Cloudflare that should help it stand up better in the face of bad actors abusing it.
Certainly not something to necessarily depend on but I thought I'd mention it.
They are donating some spare capacity.
On Firefox 102.4.0esr, also uBlock Origin.
edit: don't know if you'll see this any time soon, but I've had it fail/hang again. You might want to take a hash of uploads, so if the lost connections still end up getting transcribed, if they're reuploaded they won't get transcribed again.
Also I haven't had success in Firefox, only Chromium.
Make a JSON API and I’ll be your first customer.
I tried out this notebook about a month ago, and it was rough. After spending an evening improving it, I got everything "working", but pyannote was not reliable. I tried it against an hour-ish audio sample, and I found no way to tune pyannote to keep track of ~10 speakers over the course of that audio. It would identify some of the earlier speakers, but then it felt like it lost attention and would just start labeling every new speaker as the same speaker. There is an option to force the minimum number of speakers higher, and that just caused it to split some of the earlier speakers into multiple labels. It did nothing to address the latter half of the audio.
So, sure, someone should continue working on putting the pieces together, and I'm sure the notebook in the discussion I linked has probably improved since then, but I think pyannote itself needs some improvement first.
Sadly, I think using separate models for transcription and diarization ends up being clunky to the point that it won't ever be polished, no matter how good pyannote might get. If you have a podcast-like environment where people get excited and start talking over each other, then even if pyannote correctly identifies all of the speakers during the overlapping segments and when they spoke... Whisper cannot be used to separate speakers. You end up with either duplicate transcripts attributed to everyone involved, or something worse. Impressively, I have seen pyannote do exactly that, when it's working.
At the end of the day, I think someone is going to need to either train Whisper to also perform diarization, or we're going to need to wait until someone else open sources a model that does both transcription and diarization simultaneously. Unfortunately, it seems like most of these really big advances in ML only happen when a corporate benefactor is willing to dump money into the problem and then release the result, so we might be waiting awhile. I'm trying to learn more about machine learning, but I'm not at the point where I have any realistic chance of making such an improvement to Whisper. Maybe someone else around here can proven me wrong by just making it happen.
Whisper is great but at the point we get to kludging various things together it might start to make more sense to use something like Nvidia NeMo[1] which was built with all of this in mind and more.