Experts starting to doubt AI ‘hallucinations’ will go away: ‘This isn’t fixable’
fortune.com
fortune.com
When one confabulates, one shamelessly inserts words which sound smooth and grammatical and potentially acceptable to a listener (and the speaker), but shamelessly with no hesitation or corrections to achieve actual truth.
I suspect there needs to be some layering on of reflection and inhibition to avoid this.
The generated text is wrong, false, inaccurate. I don't think we really need a jargon-y term for it.
At least the day wasn't all lost. Learnt a new word.
1. How to do it? My thought is the answer involves building AI's that are much more than just LLM's. You need a system that has knowledge and a world model and common-sense reasoning, etc. When you're doing nothing but "predict the next token based on statistical patterns" then yeah, you're going to get hallucinations.*
2. How much will it cost?
3. How long will it take?
And so on.
* Let me add that that's a very "off the cuff" view of my current thinking on building AI's that are less susceptible to "hallucinations" (or "confabulations" if you prefer). But I'm not saying that is the only way, or that it will ever be possible to construct AI's that never make mistakes at all. After all, our current "gold standard" for intelligence to compare with (eg, human intelligence) is quite prone to making mistakes. So I'm only talking about getting to a level that eliminates the really egregious hallucofabulations (to coin a word).
Those constraints, as you eluded to, are most often things like “within a certain timeframe” or “under a certain cost”.
Saying that something cannot be done without specifying those unspoken constraints - or being explicit that you do actually mean it can’t be done, even on an infinite timescale or with infinite money - is kind of a useless statement to make IMHO.
I’ve never really understood that tendency but maybe working in tech kind of conditions you to think in a more flexible way about what may be possible?
Well, yes, by taking AI to mean something different than the current meaning. Of course if you build something that's not an LLM you may find a way to fix the confabulation* problem, but that isn't fixing what we have now.
The history of 'AI' the term is one of continually redefining it to mean something different, as the winds of technology and business blow. In this case, the assertion is that, under what we currently call 'AI', there's not a fix.
* I prefer the term confabulation over hallucination. See <https://universeodon.com/@siderea/109883198218504351>
But since we're talking about an article in Fortune, maybe it's fair to think of this strictly from a lay-person perspective.
Also, to the extent that it matters, I'm not necessarily saying "build something (completely) different than an LLM". I can picture a world with systems that include LLM's that interoperate with other components that do more of the "world model" and "commonsense reasoning" parts, yielding an aggregate system that does more/better than any of the parts could do in isolation.
One way to think of LLMs as available now is as sort of thinking partner, much like bouncing ideas off another person, except LLMs use the ideas of entire communities of persons. Of course some of the ideas bounced back at you from thousands or millions of people are going to get things wrong. Don't be like the lawyers who blindly used ChatGPT to write their legal briefs without even checking the citations, and you'll be OK.
[1]: https://www.amazon.com/How-Build-Brain-Architecture-Architec...
[2]: https://www.amazon.com/How-Create-Mind-Thought-Revealed/dp/1...
> You need a system that has knowledge and a world model and common-sense reasoning, etc. When you're doing nothing but "predict the next token based on statistical patterns" then yeah, you're going to get hallucinations.
Arguably, LLMs already have this. They can do impressive reasoning already. The problem may be that the way they encode all that is untenable. Imagine trying to describe a car in terms of its individual pieces (tires, suspension, engine). Then imagine trying to do that in terms of its molecules. Uh oh. Then its atoms. Even worse.
> How much will it cost?
To describe a car in terms of atoms? Infinite cost. We can barely even get a baby fruit fly's connectome (total brain schematic) read and stored somewhere. A car is much bigger.
It may be that LLMs are structured such that they're trying to encode all that common sense reasoning in terms of atoms, and a fundamentally different grammar is necessary to go up to parts.
Whether this is the case is still being debated. If it is the case, you'll see LLMs near a ceiling of diminishing returns.
> How much will it cost?
To figure out how to do LLMs with more expressive grammars? Unknown. I think that'd take a qualitative breakthrough in AI, and I think LLMs are just the result of a quantitative breakthrough (we're still doing the old, non-clever stuff; we're just doing it faster than ever before thanks to continued growth in GPU compute power; more monkeys on typewriters than ever before!). I don't think we've had much in the way of qualitative breakthroughs since the AI winter.
I will freely admit, I'm torn on how to think about this. On the one hand, some of the kinds of mistakes LLM's make lead me to think that they aren't exactly "doing reasoning" as I think of it. But the things they can do hint at some sort of reasoning.
Part of me wants to think of it as "they can simulate reasoning", but there's a fair objection to be raised that "simulating reasoning is just as good as doing reasoning." Part of me wants to reject that, but then I remind myself that I've always claimed to be something like a "behaviorist" when it comes to evaluating AI's. That is, I've always tried to eschew those "it's not real thinking" arguments and focus on the observed behavior, holding to a "if it displays intelligent behavior then it is intelligent" mindset.
LLM's have me sort of vacillating on all of this.
By no conception of reasoning I'm familiar with do LLMs reason, at all. They are clever software systems of an advanced sort of autocorrect/predictive text kind.
In the optimal scenario, it would just be sufficient to massively increase the context space of current LLMs.
Maybe “one-size-fits-all” solutions won’t be fixed; but it seems that domain-specific systems should be easily tailored.
I'd say 2-3 years the money proposition will be a market of small, efficient models with clear domain bias and you build out by overlapping the domains till you cover the footprint.
These large models just interconnect too much to efficiently produce reliable output unless you don't care about accuracy(eg, art or far right propaganda)
Are there any good, maybe more gatekept forums for discussing AI news? It is great to have a community to discuss with, the field is so large and moving so fast, and there's so much to learn from other practitioners, but the SNR here (where i often find news) has just gotten so low.
From my perspective, this is easily the lowest SNR comment on this post.
There will always be a much greater effort required tackling bad information than it takes to post that bad information. This is why public forums are a terrible place for debates or discussions. It's also why you're considering their post a lower SNR when it's really not, certainly not in comparison to those just claiming that it's possible with a sort of blind optimism or ranting about bots or what have you. Important to signify here, I don't lean in any which direction on this specific topic, but just highlighting how the low effort posts being viewed as higher SNR to you is exactly what leads to more noise on public forums.
Early bird gets the worm, and all that fun jazz!