AI-powered transcription tool used in hospitals invents things no one ever said
abcnews.go.com
abcnews.go.com
In others words, LLMs are only clearly useful if the results don't really matter or they can and will be externally verified.
LLMs negate a fundamental argument for computing --- instead of accurate results at low cost, we now have inaccurate results at high cost.
There is undoubtedly some utility to be had here but it is not at all clear or obvious that this will be widely transformative.
Main thing is to measure the accuracy of the approach (regardless of whether it's traditional, statistical, or human) to determine if it's fit for purpose. In this case it sounds like the transcription shouldn't be solely relied on for high-risk decisions in its current state, but could be useful for something like searching through the reference audio if it were available.
That the issue tends to be from "pauses, background sounds or music playing" also makes me suspect a lot of the cases could be relatively low hanging fruit - check the noise gate and normalization on the microphones, or potentially have the model output a quality score for each word so that low-confidence background noise can be displayed to the end user as smaller fainter text for instance, instead of part of the conversation.
It’s not quite a solved problem but it’s close.
As long as the results don't really matter and no one is auditing, it appears more "solved" than it actually is.
Another example in medicine, radiologists will start handling orders of magnitude more cases. But the number of scans done might also increase exponentially as costs likewise drop.
In the real world "better" typically translates to lower cost.
Which costs less? 1) Pay someone to transcribe a recording or 2) pay for a LLM transcription + pay someone to verify the transcription from a recording.
It is far from certain or obvious that #2 is actually "better".
A machine learning engineer said he initially discovered hallucinations in about half of the over 100 hours of Whisper transcriptions he analyzed.
The '100 hours' is almost useless information. 'About half' is meaningless without knowing the sample size. Perhaps he had 5 transcripts averaging 20 hours each, and 2 of the 5 had issues. Or perhaps there were hundreds of short transripts, where the 'almost half' would imply significance.> But the transcription software added: “He took a big piece of a cross, a teeny, small piece ... I’m sure he didn’t have a terror knife so he killed a number of people.”
> A speaker in another recording described “two other girls and one lady.” Whisper invented extra commentary on race, adding "two other girls and one lady, um, which were Black.”
> In a third transcription, Whisper invented a non-existent medication called “hyperactivated antibiotics.”
I didn't expect it to be this bad.
I use a digital recorder app to record audio from my clinical consultations. It's important for me, as a patient, to have a record, because I'm alone in there, and I frequently misremember or misunderstand things that were said.
My current recorder app has a transcription feature. It's fairly good at picking out words. It's supposed to recognize and label speakers as well, but that requires a lot of manual editing after the fact.
Still, it's fantastic having my own durable record of what was said to me, and by me. There are usually a few surprises in there!
Now, I've stopped asking for permission to record, because usually they become hostile to it. Nevertheless, it's legal, and it's my right to have.
(Literally)