Automated transcription for anything official is scary to begin with, because some noise in the background is all it takes to turn "I've never taken mushrooms" to "I take mushrooms," or whatever. And then the LLM will simply report "Patient reported using mushrooms."
One of the main issues is around homophones in an accent (Adam/Atom in American English, Bath/Barf in London English, etc.). Not to mention pronunciation variations due to fast speech, speech impedements, or parts of words side-by-side that sound like a different word.
Another big issue is around misaligned training data. For example, Whisper is known to hallucinate on silence [1].
[1] Investigation of Whisper ASR Hallucinations Induced by Non-Speech Audio (https://arxiv.org/html/2501.11378v1)
My problem is who is accountable when the AI is given autonomy and messes up
It seems like AI is being deployed so it can take the blame for some individuals decisions that will have negative impacts. Then they can shrug and say "wasn't me, it was the AI"
Or even worse they hold a fall person accountable. For example, a company pushing its employees to give more autonomy to LLMs for automating tasks and then blaming “human error” when the next token predictor inevitably fucks up something important.
It should be completely and utterly intolerable that a computer produces a different output given the same input. We shouldn’t couch that behavior in soft terms like “hallucination”. A computer system that non-deterministically makes mistakes is a defective computer system.
1. accents -- Especially around mergers (cot-caught [AmE], trap-bath [BrE] vs palm-bath [LondonE], pin-pen [Some AmE]). These can even be hard for native speakers -- try transcribing a broad Scottish, London, Brooklyn, or Indian accent and see how well you do.
2. sound/phoneme variation based on surrounding phonemes -- It is common for the 'n' sound to be realised as an 'ng' sound before a 'k' or 'g' sound due to velarization ('ng' is the velar variant of 'n' and 'k' and 'g' are velar sounds). It is common for vowels to be nasalized before nasal sounds ('n', 'm', 'ng'). It is also common in non-rhotic (don't pronounce the 'r's next to vowels like in 'start' and 'north') to pronounce an 'r' between two adjacent vowels in words ending/beginning with vowels (the "intrusive r", e.g. in "there and back").
3. sound changes due to fast speech ("I'm gonna see 'bout it t'day.", etc.)
4. ambiguity about where words start/end (e.g. "to Damon" vs "today mon" where the "mon" is the variant of "man" in Caribbean English).
5. word play, puns, etc. due to accent and other speech.
6. technical words in a given domain, specific place names, etc.
7. other things that can affect speech such as mumbling, stuttering, or slurred speech.
https://en.wikipedia.org/wiki/British_Post_Office_scandal https://en.wikipedia.org/wiki/Robodebt_scheme
It seems to me that most people regard computers as some kind of infallible truth machine. If told its spewing garbage they're more likely to double down and shoot the messenger than try and get it sorted out.
"AI".
If concepts can be that sloppy, then the party that believes it an argument that NNs surpass humans get a point.
Edit: in fact, there is a point: we compare AI (proper AI) to optimal professionals, but that is not the real scene. And this is why in computing we bet on deterministic algorithms: they do not guess a solution, they compute it. There is no comparison with the possibility of failure from a biology based system - in deterministic computing the failure is restricted to exceptions.