Show HN: Gemini LLM corrects ASR YouTube transcripts
ldenoue.github.io
ldenoue.github.io
https://news.berkeley.edu/2017/02/24/faq-on-legacy-public-co...
Discussed at the time (2017) https://news.ycombinator.com/item?id=13768856
IME youtube transcripts are completely devoid of meaningful information, especially when domain-specific vocabulary is used.
Anything that's even remotely domain specific becomes a garbled mess. Even watching documentaries about light engineering/archeology/history subjects are hilariously bad. Names of historical places and people are randomly correct and almost always never consistent.
The second anyone has a bit of an accent then it's completely useless.
I keep them on partially because I'm of the "everything needs to have subtitles else I can't hear the words they're saying" cohort. So I can figure out what they really mean, but if you couldn't hear anything I can see it being hugely distracting/distressing/confusing/frustrating.
Soundex [0] is a prevailing method of codifying phonetic similarity, but unfortunately it's focused on names exclusively. Any correction-by-LLM really ought to generate substitution probabilities weighted heavily on something like that, I would think.
There may be mistakes like the ones you mentioned (getting names wrong/inconsistent), but if I know what was intended, it's pretty easy to ignore that. I think expecting "textual" correctness is unreasonable. Usually when there are mistakes, they are "phonetic", i.e. if you spoke the caption out loud, it would sound pretty similar to what was spoken in the video.
Of course you think that, you don't have to rely solely on closed captions! It's usually not even posed as an expectation, but as a request to correct captions that don't make sense. Especially now that we have auto-captioning and tools that auto-correct the captions, running through and tweaking them to near-perfect accuracy is not an undue burden.
> if you spoke the caption out loud, it would sound pretty similar to what was spoken in the video.
Yes, but most deaf people can't do that. Even if they can, they shouldn't have to.
Deleting thousands of hours of course material because you're worried they're not able to understand autogenerated captions just ensures everyone loses. Don't be so ridiculous.
Even mentally sounding them out (which is fine for me since I have no relevant disabilities, I just despise trying to take in any meaningful quantity of information from a video) when they look weird doesn't make them tolerable *for me*.
It's still a good thing overall that they're tolerable for you, though, and I hope other people are on average finding the experience closer to how you find it than how I find it ... but I definitely don't, yet.
Hopefully in a year or so I'll be in the same camp as you are, though, overall progress in the relevant class of tech seems to've hit a pretty decent velocity these days.
I think that the biggest problem is that the subtitles do not distinguish between the speakers.
If the automatically-generated captions are now of a similar quality as human-generated ones, then that changes things.
[1] https://news.berkeley.edu/wp-content/uploads/2016/09/2016-08...
Is building a ramp to meet ADA requirements not using technology to solve a legal issue?
Building a ramp solves a problem. Pointing at a ramp 5 blocks away 7 years later and asking "doesn't this solve this issue" doesn't.
I'm somewhat tempted to think that whoever sued berkeley and had the whole thing taken down in this specific case was just being a knob, but OTOH there's issues even with that POV in terms of letting precedents be set that will de facto still become "screw over deaf people entirely" even when everybody involved is doing their best to act in good faith.
Hopefully speech-to-text and text-to-speech will make the question moot in the medium term.
I really think this and other tech advances are going to be our saviors. It's still early days and it sometimes gets things wrong, but it's going to get good and it will basically allow us to have our cake and eat it too (as long as we can prevent having automated solutions banned).
But the print form was even less accessible, and they kept publishing that...
Clearly the result of a regulation that meant well. But the road to hell is paved with good intentions.
It's a bit reminiscent of a law that prevents institutions from continually offering employees non-permanent work contracts. As in, after two fixed-term contracts, the third one must be permanent. The idea is to guarantee workers more stable and long-term perspectives. The result, however, is that the employee's contract won't get renewed at all after the second one, and instead someone else will be hired on a non-permanent contract.
The longer I live the more the truth of this gets reinforced. We humans really are kind of bad at designing systems and/or solving problems (especially problems of our own making). Most of us are like Ralph Wiggum with a crayon sticking out of our noises saying, "I'm helping!"
If it's, say, 5000 hours then through the best model at assembly.ai with no discounts it's cost less than $2000. I know someone could do whisper for cheaper, and there likely would be discounts at this rate but worst case it seems very doable even for an individual.
It is not perfect, it'd sometimes replace words with a synonym, but it is much faster and cheaper.
The low cost of Gemini 1.5 Flash-8B costs $1 per 500 hours of transcript.
Grabbing the audio from thousands of hours of video, or even just managing getting the content from wherever it's stored, is probably more of an issue than actually creating the transcripts.
If anyone reading this has access to the original recordings, this is a pretty great time to get transcriptions.
It would be great if they were annotated and served in a more user-friendly fashion.
As a bonus link, one of my favorite courses from the time: https://archive.org/details/ucberkeley_webcast_itunesu_35482...
And to run it on a "free" product they probably use a very tiny, heavily quantized version of their already weak ASR.
There's lots and lots of better meeting bots if you don't mind paying or have low usage that works for a free tier. At Rev we give away something like 300 minutes a month.
Since you have experience in this, I’d like to hear your thoughts on a common assumption.
It goes like this: don’t build anything that would be feature for a Hyperscalar because ultimately they win.
I guess a lot of it is a question of timing?
It is hard to compete with these hyperscalers because they use pseudo anti-competitive tactics that honestly should be illegal.
For example, I know some ASR providers have lost deals to GCP or AWS because those providers will basically throw in ASR for free if you sign up for X amount of EC2 or Y amount of S3, services that have absurd margins for the cloud providers.
Still, stuff like Supabase, Twilio, etc show there is a market. But it's likely shrinking as consolidation continues, exits slow, and the DOJ turns a blind eye to all of this.
But you do have to be next to amazing at execution
Which is not to disagree with you, only to "yes, and" to emphasise that it's a fairly narrow path and 'amazing at execution' is necessary but not sufficient.
We also compared Amazon, Google, Microsoft Azure as well as a bunch of smaller players (from Edinburgh and Cambridge) and - consistent with what you reported - we also found Google ranked worst - but that was a one-off study from 2019 (unpublished) on financial news.
Word Error Rate (WER), the standard metric for the tast, is not everything. For some applications, the ability to upload custom lexicons is paramount (ASR systems that are word-based (almost all) as opposted to phoneme based require each word to be defined ahead of being able to recognize said word).
>the only hyperscaler with a good ASR is Azure
How would you say the non-hyperscalers compare? Speechmatics for example?
I don't think it's quite out of the realm of the possibility to have interpreted as "Gemini LLM corrects ASMR YouTube transcripts". Because you know..they're whispering so might be hard to understand or transcribe.
1. It brings everything back to the "average." Any outliers get discarded. For example, someone who is a circus performer plays fetch with their frog. An LLM would think this is an obvious error and correct it to "dog."
2. LLMs want to format everything as internet text which does not align well to natural human speech.
3. Hallucinations still happen at scale, regardless of model quality.
We've done a lot of experiments on this at Rev and it's still useful for the right scenario, but not as reliable as you may think.
The extra dimensions of analysis cause increased hallucination at times. So maybe it solves the frog problem, but now it's hallucinating in another section because it got confused by another frame's tokens.
One thing we've wanted to explore lately has been video based diarization. If I have a video to accompany some audio, can I help with cross talk and sound separation by matching lips with audio and assign the correct speaker more accurately? There's likely something there.
https://research.google/blog/looking-to-listen-audio-visual-...
This is just like playing a game of markov telephone where the step in OP's solution is likely higher compute cost than the step YT uses, because YT is interested in minimizing costs.
One of the better reasons to use Cerebras/Groq that I've found so you can return huge amounts of clean text back fast for processing in other ways.
Are you saying it works with 70B models on Groq? Mixtral, Llama? Other?
I have had little success with Gemini and long videos. My pipeline is video -> ffmpeg strip audio -> whisperX ASR -> groq (L3-70b-specdec) -> gpt-4o/sonnet-3.5 for summarization. Works great.
https://github.com/google-gemini/generative-ai-js/issues/269...
Once you accept it okay for the LLM to just replace words in a transcript, you might as well just let it make up a story based on character names you've provided.
That's a wild exaggeration. Professional transcripts often have small (and not so small) mistakes, caused by typos, mishearing or lack of familiarity with the subject matter. Depending on the case, these are then manually proofread, but even after proofreading, some mistakes often remain, and occasionally even introduced.
Talking has a lot of nuances to it. Just try to read a Donald Trump transcript. A professional author would never write a book's dialogue like that.
Using a generic LLM on transcripts almost always reduces accuracy as a whole. We have endless benchmark data to demonstrate this at RevAI. It does, however, help with custom vocabulary, rare words, proper nouns, and some people prefer the "readability" of an LLM-formatted transcript. It will read more like a wikipedia page or a book as opposed to the true nature of a transcript, which can be ugly, messy, and hard to parse at times.
That's a bit too far. Ever read Huck Finn?
But I understand it can be difficult to trust: that’s why the project is on GitHub so you can run it on your own machine and look at how the key is used.
I will try to offer a version that doesn’t require any key.
Whether you want to deal with it being annoying is your call.