But the real danger with these IMO is that they're turning casual conversations into a permanent record, and one that will be completely discoverable in court, should the company get into trouble later.
But the real danger with these IMO is that they're turning casual conversations into a permanent record, and one that will be completely discoverable in court, should the company get into trouble later.
The problems start when using conference room audio or someone is on their laptop mic. If they miss a word they never do unintelligible, they just start playing madlibs based on the rest of the sentence.
We just went through a round of 100+ (non-sensitive) VoC interviews and they really cut down the workload of compiling all of the feedback. If the audio was a little shaky though, we pretty much had to throw away the transcripts and do them from scratch like we used to.
Imo this is the single biggest flaw of LLMs. They're great at a lot of things, but knowing when they're wrong (or don't have enough information to actually work on) is a critical flaw.
IMO there's nothing structural about why they shouldn't be able to spot this and correct themselves - I suspect it's a training issue. But presumably bots that infer context/fill in the dots rank better on what people like... at the cost of accuracy.
Lots of tools in our toolbelts to do better uncertainty calibration but it trades off against other capabilities and actually can be rather frustrating to interact with in agentic contexts since it will constantly need input from you or otherwise be indecisive and overly cautious. It’s not technically a limitation of transformer architecture but it is more challenging to deal with than other architectures/statistical paradigms.
Like you can maintain a belief state and generate conditional on this and train to ensure belief state is stable and performant. But evals reward guessing at this point, and it’s very very hard to evaluate the calibration in these open ended contexts. But we’re slowly getting there, just not nearly as fast as other capabilities.
The confidence level can be any, as long as it's reported accurately often enough. "This is my conjecture, but", "I'm not completely sure, but", and "most historians agree that" are all perfectly valid ways to start a sentence, which LLMs never use. They state mathematical truth, general consensus, hotly debated stances, and total fabrication, with the exact same assertiveness.
> ways to start a sentence, which LLMs never use
A huge part of the problem is we've invented a document-generator setup which exploits human cognitive illusions, and even the smartest person can't constantly override the instinctive brain-bits that "sees" fictional entities and infers the intent of a mind. That makes it weirdly-hard to discuss the setup's shortfalls or how to improve it.
To wit: The machine does not possess any kind of confidence about how Rome fell. Or even whether Rome fell. It has "confidence" about which word/token will next in a "typical" document given the document-so-far has text like "How did Rome fall?" It may be straightforward to burn money training the system so that its "typical" story never has a computer-character with confident words about Roman history, but that's just papering over the underlying problem.
TLDR: We can't fix the thinking-habits or beliefs inside the mind of an entity that doesn't actually exist. Changing the story-generator to contain a tee-totaling Dracula dispensing life-advice doesn't mean we "cured the disease of vampirism."
That's why I'm still cautiously optimistic about LLMs somewhere being good enough. I don't know if or when someone will manage to do it, but I'm hopeful.
Do stochastic parrots dream of the number of 'e's in "electric sheep"?
Not always though. Let’s say that someone is saying ”1 2 3 4 <unintelligible> 6 7 8” then it will happily write 5 in the middle and give it good confidence as based on the context, it is the only likely word. Varies between TTS providers though.
Basically, why they are so good in average is that they estimate what is said most often based on the context. The context being then not only the audio but what was transcribed previously.
And if you don’t want it to be based on what is most likely to be said in context and only based on the audio around 1 word it is going to be awfully wrong most of the time.
Add accents, and half the words would be indistinguishable from each other (note that word "indistinguishable", ironically, would be quite distinguishable).
People parse things like that in so much context, based in their own understanding of a situation, their grasp on speakers accent or speech impairments, etc.
Add to that that most native english speakers blur words together. The pause that in some languages is used to separate words, is used in english to separate sentences. English language as spoken doesn't separate words natively.
The text-to-speech before LLMs was meh. I think it's the ability to generate filler for uncertain words that makes it feel magic compared to before.
If you ask a good model something that makes no sense, it will tell you it makes no sense and it can't answer the question; so I know it's possible.
The reason AI companies won’t do this of course is it would completely ruin the illusion of confident confidence these machines project.
(Which is intensely depressing to a human that doesn’t.)
Of course there's a secondary problem that the model may then overuse the unintelligible option, but that's something that's a matter of training them properly against that eval.
You could also try thresholding the output based on perplexity to remove the parts that the model is less sure about, but that's not going to be super accurate I think.
Which reminds me that that's another big issue with LLMs - they'll blindly do whatever you ask them to, without pushback. (Again, I miss 3.5/3.6 era Sonnet which actually had half a spine. Fuck anthropic for blindly chasing coding benchmarks at the cost of everything else.)
I've engaged in several "CMVs" (or "tell me why X is bad") with LLMs, and very often it's clear it's just saying stuff to say it, giving very terrible points on unjustifiable positions that collapse the moment I counter argue even slightly rationally.
AA-Omniscience is a knowledge and hallucination benchmark that rewards accuracy, punishes bad guesses and provides a comprehensive view of which models produce factually reliable outputs across different domains. The benchmark contains 6,000 questions across 6 major domains, derived from authoritative academic and industry sources and generated automatically using an LLM-based question generation agent to ensure unambiguity, scalability and factual precision
This is a solvable issue, the current model and harnesses just aren't made with that assumption - hence they're doing "best effort while guessing if unsure".
Give it a few more months to years and things will likely settle how he pitched - at least in the context of note taking: only let it become "lore" if it didn't have to guess a word.
Currently there is basically only one mode - and it's optimized for conversation. The note taking is just glued on with that functionality as the backbone, and that's probably not going to stay.
I'm hesitant to admit even that. Like any computational linguistics problem, accuracy relies on coverages of all levels: form morphology, through syntax and semantics to speech act and world knowledge.
I worked with state of art speech recognition in healthcare setting. The model was specifically trained on small set of languages with emphasis on covering medical terminology.
It worked great for conversations most of the time, but sometimes messed up very badly. For instance when patient would mention the name of a relative, a street address or phone number. Spelling out an email address would mess it up completely.
It's just like when you're a horrible typist and rely on spell checking: The red squibles are gone, but the story no longer makes sense. Or when you "autofix" a syntax error, but the meaning diverges from your intention.
As the technology improved the number of words decreases, but the mistakes get more severe.
The point isn't that it's unexpected. It's that prior text-to-speech systems were much better about this particular failure mode, prone to spitting out entirely incorrect words but not rephrasing entire sentences.
This is a particularly bad failure mode because people don't notice it.
> What we need are tools that embrace that and ping the agent to validate what it just said or double check.
This is not a problem that can be fixed by throwing more AI at it. It's a shared problem to all such systems, whether they're audio-text transformers or LLMs. Agentic review would just further push the system towards creating output that looks correct, but is not.
LLM translation does the same, yielding more natural text, but generally not better translation. In several cases, especially the "easy" translation between similar languages (e.g. within a language group like Germanic or Nordic) LLM-powered translation is notably worse than more primitive "word & phrase book" systems, tending to change the meaning of the text in order to have good grammar whereas these older systems would give crude or grammatically incorrect translations that still retained the core meaning.
Maybe it depends on topics or length, for me it's usually 1-2 paragraphs of a German article to share online.
Same languages, same use case. My experience is different. On both google translate and others. ¯\_(ツ)_/¯
Are you native in both languages? If you are only native in one of them, it would be insightful to find if people with your skillset but native in the language you are not have the same opinion as you.
Sadly there are no examples here to compare.
If the prediction strength is below X, put an indicator that it couldn't make a valid prediction?
Someone tell Altman
RTO problems
Got a team with Indian, Chinese, Texan, British, and Australian? Your A.I.-powered translation tool is going to get 80% of your conversation wrong.
But key in my prompt is asking 1) for it to flag any low confidence or context-nonsensical statements in the transcript, with the timestamp, so then I can listen to the original audio and either clarify, correct, or say "I couldn't understand that either, here's my best guess and mark it low confidence", then 2) which I see as critical: Claude also is told to create a "context" document that it maintains based on my answers, so it starts to gather ASR things like "transcript commonly hears A B and C as variants of name X", who is who, internal product and project names and context info on them. 3) Claude is told specifically to read this prior to summarizing the transcript, and to consult it as it is doing so, and to ask me on anything it's not confident on.
What is then starting to get quite powerful for me is moving from full text search of my meeting notes in Obsidian (I'm a PM in a lot of meetings), but I can point Cowork to the Obsidian notes folder (because they're all Markdown) and start doing rich "querying" of it. "When did [stakeholder] first mention [feature] as a release blocker?" and it can point to the meeting.
My system works well, and I've done a bit to fine tune the automation and friction reduction, and it's a bit easier to manage because I'm not generally creating summaries for broader consumption but as my second brain (I have a separate prompt that utilizes some of that "knowledge" to build those).
One thing I've found helpful with this is moving the summarization itself into something with "context/memory". Krisp is capable of generating summaries but can't/doesn't review prior transcripts. Its role is just "give me the transcript as you heard it".
Or mostly just confirm what you half-remembered?
Trying to figure out whether the value of the loop is rediscovery or just precise lookup.
- the person said 8 to 10
- LLM transcribed as 18
Granted, the person had a foreign accent and didn't enunciate very clearly. But I knew they meant 8-10 if for no other reason than 18 didn't make sense given the context. But the AI isn't smart enough, and then 18 goes into the record.
Half- vs. full duplex. Headphones is all you really need, though of course a directional mic and/or one closer to your mouth will yield a clearer audio recording as well.
Isn't that what people do?
Nixon tapes for example: https://kagi.com/search?q=site%3Anixonlibrary.gov+%22unintel...
But the summarization feature is where the most ridiculous errors and omissions happens.
I sincerely hope these aren't used in court.
Potentially sinister due to the biases of the model, as the model may have been trained using internet content that has a lot more fictional titillating evil overlord board meetings than the actual mind-numbing real thing. Training that included extremist anti-corporate dogma might even bias the language models towards hallucinating the worst possible misinterpretation.
I've seen whisper hallucinate whole legal arguments whole cloth when the AGC was broken in it and the audio went quiet-- so I think the language models in it are more than powerful enough to politically load a transcript.
Good practice should be to minimize any unnecessary stored records because ANY record just means more processing costs in discovery and god knows how much extra cost in litigation should it happen to have an unfavorable interpretation in light of some impossible to anticipate future litigation.
But if AI transcription must be used it would be might be prudent to save a copy of the original audio along with it.
Ironic use of “sinister” when you probably mean “nefarious” and don't mean to perpetuate silly old superstitions about “left-handed” people being evil :p
I’ve been saying it since the mid-10s, but it’s worth repeating: data isn’t gold, it’s more like oxygen in a room in that the higher the concentration, the more likely it is to poison the inhabitants or explode with an errant spark (lawsuit).
Collect only what’s needed to perform the function, and store it only as long as necessary for compliance. Anything else is going to spool counsel.
Limiting data retention doesn't mean hiding bad things, it means limiting exposure in general. The more of a thing - anything - that you have, the bigger a target you are to bad actors. By extension, companies holding vast sums of data beyond what's needed to process a given transaction or remain compliant with the law end up placing themselves at risk of being targeted and said data used as leverage against them.
You don't limit data to hide bad shit you're doing, you limit it to avoid others using it to do bad shit against you or your customers. If someone or something is engaged in bad shit, there will always be evidence somewhere regardless of data retention policies.
I would add that their is no guarantee their are correct as well.
“At timestamp X, person Y said Z” says the robot, and then you dutifully scrub the audio to timestamp X to verify.
I’m overall an AI optimist but this is going to blow up in people’s faces very quickly. (I would explain this to my manager but he has AI note taking turned on in all his meetings!)
And that’s not even getting into the use of it for sensitive clinical notes in eg. mental health…
Also social settings will change, when everything you say stays on record forever in every meeting...
The parts that aren’t privileged. On the other hand, perhaps the truth-seeking function of the justice system will be better equipped than before when we had to rely on (more) faulty human recollection.
I have a friend who works at a large-ish company that imports and manufactures things (in one of the clerical/quantitative professions). A few years back, they had the IT department go on a kind of "inquisition", wherein they forced employees to disable the summarization function that came with MS Teams, and threatened to fire them if they did not. The resistance to this demand was surprising -- most people are clueless about the cost of their own convenience. Worst of all, people would zone out of meetings, because the AI was producing summaries, which they would then never read.
The effect of the technology was that it made meetings infinitely more expensive, because the supposed benefit of meetings was nullified by complacency, _and_ it made the meetings a liability (incorrectly summarized meetings, that could be used in the discovery process, sure, but could also be sold by MSFT as a kind of market-research-data to competitors in the space).
Nothing illegal has to happen in these meetings at all, for this tech to cause an infinity of problems for the corporation. Every employee that uses these is effectively an unwitting spy. And if that is the case, then the meetings might as well be recorded and uploaded to YouTube (or whatever people watch these days)[1].
[1]: Maybe this is the future. Which I am okay with, but only if the entire planet has to do it, and the penalties for not doing it are irrecoverably severe.
Modernized. Industrial AI scale.
I’ll be honest, this is something that I hope AI note taking tools capture and incorporate into summaries of the company’s status. Especially if they act as an intermediary without revealing the specific person who said it. There’s a lot of information latent within organizations that doesn’t get properly shared due to concerns of retaliation or simply embarrassment that would benefit everyone by being communicated sooner.
Maybe some smaller shops are not like this, but the bigger your company is, the more you'll find this type of thinking to persist.
In theory, I do like your idea - anonymously cascading feedback upstream. I just see no avenue for this to succeed in practice.
"It seems that starting in 2025 one of your employees began spreading many bigly unfair and hateful lies about Dear Leader in team meetings. We at the Department of Truth would hate to see your operating license revoked for encouraging such unpatriotic behavior..."
The only question is whether everyone gets a slice, or it ends up locked down so only governments and corporations have access to it. Obivously I come down on the sousveillence side of the fence - it's the lesser of two evils. If it exists I want everyone to have it, and you can't stop it existing.
Be careful what you wish for. Particularly when it involves tech that often gets it very, very wrong.
It is true accusation and potential for success of it.