If it turned out an LLM embezzled funds and spent them at an internet casino, we would have folks in the comments explaining that what really matters is the embezzlement rate compared to humans doing the same job.
In addition to not enjoying my work as much as what I used to because it's become babysitting an superpowered AI toddler, I now have to deal with this kind of opinion online.
That's hopefully temporary: you are dealing with provisional architectures - as is obvious by their lackings (transparency; reflection; evolution; one shot learning...).
The very fact that you use the term 'AI' for LLMs when some of us would not ("NNs are used in AI" does not mean that all NNs would be AI), or would be wary of that use signifies a problem that is being tackled and will be worked on until the next stage.
Because thats how probabilistic machines work. You can’t change that.
Clearly AI has been extremely useful for doctors despite its “flaws”.
I dunno, for every new hype automation, the machine is given a lot more leeway then people. At least by some on HN. I remember FSD discussion years ago where FSD was already supposedly better then people and all its problems explained away.
I think this sums up what I find most frustrating about all of this from the executive level. It feels like everything that was important just a few years ago is now considered unnecessary baggage, that is merely there to slow everything down. What would've got you sacked is now applauded at times, and this is only really a year or two into it proper.
Automated transcription for anything official is scary to begin with, because some noise in the background is all it takes to turn "I've never taken mushrooms" to "I take mushrooms," or whatever. And then the LLM will simply report "Patient reported using mushrooms."
One of the main issues is around homophones in an accent (Adam/Atom in American English, Bath/Barf in London English, etc.). Not to mention pronunciation variations due to fast speech, speech impedements, or parts of words side-by-side that sound like a different word.
Another big issue is around misaligned training data. For example, Whisper is known to hallucinate on silence [1].
[1] Investigation of Whisper ASR Hallucinations Induced by Non-Speech Audio (https://arxiv.org/html/2501.11378v1)
It should be completely and utterly intolerable that a computer produces a different output given the same input. We shouldn’t couch that behavior in soft terms like “hallucination”. A computer system that non-deterministically makes mistakes is a defective computer system.
1. accents -- Especially around mergers (cot-caught [AmE], trap-bath [BrE] vs palm-bath [LondonE], pin-pen [Some AmE]). These can even be hard for native speakers -- try transcribing a broad Scottish, London, Brooklyn, or Indian accent and see how well you do.
2. sound/phoneme variation based on surrounding phonemes -- It is common for the 'n' sound to be realised as an 'ng' sound before a 'k' or 'g' sound due to velarization ('ng' is the velar variant of 'n' and 'k' and 'g' are velar sounds). It is common for vowels to be nasalized before nasal sounds ('n', 'm', 'ng'). It is also common in non-rhotic (don't pronounce the 'r's next to vowels like in 'start' and 'north') to pronounce an 'r' between two adjacent vowels in words ending/beginning with vowels (the "intrusive r", e.g. in "there and back").
3. sound changes due to fast speech ("I'm gonna see 'bout it t'day.", etc.)
4. ambiguity about where words start/end (e.g. "to Damon" vs "today mon" where the "mon" is the variant of "man" in Caribbean English).
5. word play, puns, etc. due to accent and other speech.
6. technical words in a given domain, specific place names, etc.
7. other things that can affect speech such as mumbling, stuttering, or slurred speech.
"AI".
If concepts can be that sloppy, then the party that believes it an argument that NNs surpass humans get a point.
Edit: in fact, there is a point: we compare AI (proper AI) to optimal professionals, but that is not the real scene. And this is why in computing we bet on deterministic algorithms: they do not guess a solution, they compute it. There is no comparison with the possibility of failure from a biology based system - in deterministic computing the failure is restricted to exceptions.
https://en.wikipedia.org/wiki/British_Post_Office_scandal https://en.wikipedia.org/wiki/Robodebt_scheme
It seems to me that most people regard computers as some kind of infallible truth machine. If told its spewing garbage they're more likely to double down and shoot the messenger than try and get it sorted out.
My problem is who is accountable when the AI is given autonomy and messes up
It seems like AI is being deployed so it can take the blame for some individuals decisions that will have negative impacts. Then they can shrug and say "wasn't me, it was the AI"
Or even worse they hold a fall person accountable. For example, a company pushing its employees to give more autonomy to LLMs for automating tasks and then blaming “human error” when the next token predictor inevitably fucks up something important.
And I am sure people will still defend that dystopia with "companies send canned response all the time".
the problem is that introduction of any new technology usually happens with so much emotional baggage, that when there's an error (human or otherwise) some humans will understandably see their biases confirmed in them, and will signal boost everything to the Moon.
LLM's will never be reliable enough to let loose on tasks that require 100% accuracy, therefore a human will have to review their work. So will any time actually be saved, or at least enough time to justify the cost and extra complexity of the new system?
Also, the types of mistakes are completely different. A person may mishear something and ask to verify; LLM is always certain that what it transcribes is a fact. A person might omit something but won't make up the facts like that.
So, a mistake is not the same thing as a hallucination.
Based on watching the medical software field as a consumer (patient) and friends who are doctors, this is a fantasy. The quality of software in this field is abysmal and there seems to be almost no repercussions to those who develop or sell it.
Which is precisely why this sort of thing can be rolled out without much fear by those pushing it.
I get why it's problematic, obviously, but if it produces statistically better results (which I have no idea of), I don't think it's right to just write it off because of this.
I do not oppose AI integration; I'm not a Luddite. But having a "move fast, who cares if a couple die" isn't the way to go with sensitive fields, like the medical field.
I suppose we will come up with proper responsibility-hierarchies and guardrails around AI, but until then, people have a right to complain about the lack of them.
I can say for a fact that reliable medical transcrption and dictation is worth handling hallicinations..
its an order of magnitude worse in real life.. or else its just ommitted info since most docs and nurses dont have time for details..
By all means lets be accurate but we must remember all these complex workflows are filled with human error..