--
[0] - Or "something very convincingly pretending to think by parroting stuff back", if you're closer to the "stochastic parrot" view.
[1] - Per my hand-wavy hypothesis that the bulk of what we call thinking boils down to proximity search in extremely high-dimensional space.
Even GPT 3.5 has trouble following instructions, but I’ve found that GPT 4 is almost flawless. I can tell it the document uses Australian English but to preserve US spelling for product names and it’ll do it!
One quirk is that it’s almost too good at following instructions. You have to tell it to preserve product names, vendors names, place names, etc… otherwise it’ll “correct” the spelling of anything you forgot to list.
The idea is for it to automatically detect the "language" of each word based on the context and its own understanding of the world.
E.g., the following sentence:
"We deployed windows server data centre 2022 into our data center, which has no windows for physical security."
Will be corrected by GPT-4 to the following:
"We deployed Windows Server Datacenter 2022 into our data centre, which has no windows for physical security."
Notice that it combined "data" and "centre" into "Datacenter" and it corrected the second "center" into "centre", which is the British/Australian spelling of the word. It also correctly capitalised only the first use of the word "windows", etc...
That requires a level of understanding that GPT 3.5 just barely has, and no ordinary grammar checker tool has.
General purpose models containing significant overlap between Project Gutenberg and Github are unnecessary and don't scale. Moby Dick has little to do with C++ unless you're creating art for novelty's sake. This is entirely speculative, but I'm convinced ChatGPT is faking the appearance of a single oracle while delegating requests to specialized models under the hood. It scales better and makes sense than trying to serve a 1T model to address everybody's banal questions.
Like, at its core, for people who only want to write literature, give them a model with underweighed programming-related corpora. Writers don't need it, will never use it, and that space could be filled with training content relevant to literature. Anything else results in expensive, unscalable solutions or jack-of-all-trades, master-of-none outcomes.
In recent usage, GPT3.5 helped me hack my way through writing Pester tests for Powershell scripts for the first time, and I mean hack-- there were a lot of assumptions it made and things it got wrong. GPT4 did a much better job, but I couldn't help but think 3.5 probably has a ton of other training data in it that detracts from the specialization I needed from it in that context. For coding help, you don't want to ask some random librarian who occasionally recommends resources that don't exist; you ask someone who specializes in coding and trust they have familiarity with that domain.
Then we need a new system, because LMs, no matter if they are large or not, cannot do that, for a very simple reason:
A LM doesn't understand "truthfulness". It has no concept of a sequence being true or not, only of a sequence being probable.
And that probability cannot work as a standin for truthfulness, because the LM doesn't produce improbable sequences to begin with...it's output will always be the most (within heat settings) probable sequence. The LM simply has no way of knowing whether the sequence it just predicted is grounded in reality or not.
It can reason. To an extent.
Wrong statements about Python are simply less probable than wrong statements about Rust, since there is more Python than Rust in the training data.
That changes exactly nothing about the fact that the system isn't able to detect when it makes a blunder in Python.
That is not what we've observed though. Quite the opposite - we're seeing that the bigger LLM is and the more domain-specific material it digested, the more truthful it becomes.
Yes it can still make an error and be unable to spot it, but so can I.
No, that is not my claim. That is part of the explanation for it.
My claim is this: An LLM is incapable of knowing when it produces false information, as it simply doesn't have a concept of "truthfulness". It deals in probabilities, not alignment with objective reality.
And it doesn't matter how big you make them...this fact cannot change, as it is rooted in the basic MO of language models.
So, now that we have covered what my claim actually is...
> That is not what we've observed though. Quite the opposite - we're seeing that the bigger LLM is and the more domain-specific material it digested, the more truthful it becomes.
...I can ask what this observation has to do with it, and the answer is: Nothing at all. LMs with more params may produce untruthful statements less often, but what does this change about their ability to recignize when they do produce them? And the answer is: Nothing. They still can't.
a LLM can indeed know when it produces likely incorrect responses. Not a hypothetical.
What's the point of making claims you have no intention of rescinding regardless of evidence ? People are so funny.
I claim that the human brain doesn't understand "truthfulness" either. It merely creates the impression that understanding is taking place, by adapting to social and environmental pressures. The brain has no "concepts" at all, it just generates output based on its input, its internal wiring, and a variety of essentially random factors, quite analogous to how LLMs operate.
Do you have any evidence that contradicts that claim?
Empirical evidence? Yes I do.
The brain commands an entity that has to exist and function in the context of objective reality. Being unable to verify it's internal state against that, would have been negatively selected some time ago, because stating: "I'm sure that rumbling cave bear with those big sharp teeth is a peaceful herbivore" won't change the objective reality that the caveman is about to become dinner.
How that works in detail is, to the best of my knowledge, still the subject of research in the realm of neurobiology.
The concept of truth is notoriously hard for humans to grapple with. How do we know something is true isn’t just a neurobiological question, it’s been grappled with throughout the history of philosophy — including major revisions of our understanding in the past 80 years.
And for the record, rumbling cave bears are mostly peaceful herbivores.
For the record, all members of the Genus Ursus belong to the Order Carnivora, which literally translates to "Meat Eaters". And that includes Ursus spelaeus, aka. the Cave Bear.
And while it most likely, like many modern bears, was an Omnivore, that "Omni" very much included small, hairless monkey-esque creatures with no natural defenses other than ridiculously small teeth and pathetic excuses for claws, if they happened to stumble into their cave.
> The concept of truth is notoriously hard for humans to grapple with.
I am not talking about the philosophical questions of what truth is as a concept, nor am I talking about the many capabilities of humans to purposefully reshape others perceptions of truth for their own ends.
I am talking about truth as the observable state of the objective reality, aka. the Universe we exist in and interact with. A meter is longer than a centimeter, and boiling water is warmer than frozen water at the same pressure, whether any given philosophy or fabrication agrees with that or not, is irrelevant.
And as I have demonstrated above, humans, and for that matter other species on this planet featuring capable brains like Corvidae or Cetaceans, do in fact have a concept of truth: They are capable of recognizing false or misleading information as being incongruous with objective reality: A raven that sees me putting food into my left hand, will not jump to a patch of ground where I pretend to put food with my right hand.
This is despite the fact that my actions of "hiding the food" with the empty hand are stochastically indistinguishable from an action of actually hiding food from with my left hand.
I don't think those steps are out of the bounds of possibility, really.
The problem is what you mean when you say "consistency".
The LM checks if sequences are stochastically consistent with other sequences in the training data. Within that realm, the sentence: "In the Water Wars of 1999, the Antarctic Coalitions aramada of Hovercraft valiantly faught in the battle of Golehim under Rear Admiral Korakow, against the Trade Unions Fleets." is consistent. Because, while it is total bollocks, it looks stochasticaly like something that could be in a historical text.
So, in it's context, the LM does exactly what you ask for. It produces output that is consistent with the training data.
Truthfulness is a completely different form of consistency: Does the semantic meaning of the data support the statement I just made? of course it doesn't, there isn't an Antarctic Coalition, there were no Water Wars in 1999, and no one ever built an Armada of Hovercraft for any war against a "Trade Union Fleet".
But to know that, one has to understand what the data means semantically. And our current AIs ... well, don't.
Another wrong statement, you're on a roll today.
https://arxiv.org/abs/2305.11169
https://arxiv.org/abs/2306.12672
There's a word we would use to describe your confidently erroneous statements were it one of the outputs an LLM. Wonder what that might be..
Base GPT-4 was excellently calibrated. So this is just wrong.
“Hallucination” is part of thought. Solving a new problem requires hallucinating new, non existing, possible outcomes and solutions, to find one that will work. It seems that eliminating the ability to interpolate and extrapolate (hallucinations) would make intelligence impossible. It would eliminate creativity, tying together new concepts, creation, etc.
Is the goal AI, or a nice database front end, to reference facts? Is intelligence facts, or is it the flexibility and the ability to handle and create the novel, things that are new?
The ability to have confidence, and know and respond to it, seems important, but that’s surely different than the elimination of hallucinations.
I’m probably misunderstanding something, and/or don’t know what I’m talking about.
The latter given the kind of products that are currently being built with it. You don't want your code completion or news aggregator to hallucinate for the same reason you don't want your wrench to hallucinate, it's a tool.
And as for hallucinations, that's a PR friendly misnomer for "it made **** up". Using the same phrase doesn't mean it has functionally anything to do with the cognitive processes involved in human thought. In the same way a 'artificial' neural net is really a metaphorical neural net, it has very few things in common with biological neurons.
Obviously it's worth it to try and eliminate the incorrect information, but what grand-op is saying is we don't want to do that if it takes away some valuable emergent properties.
I've been thinking there's some parallels between how AIs hallucinate and how human toddlers do. If you ask a toddler/young child a question about a fact they don't know, they will usually say, "iuno" (even when they should), but depending on the child and the circumstances, they will sometimes just make up a story on the spot and sound as if they believe it. "Who invented ice cream?" "Santa Claus! Mommy left him milk and cookies and he turned it into ice cream." It doesn't make any real sense but it seems facially plausible in their universe.
But somewhere between first learning to speak and around 7ish, kids become markedly more accurate how they model the world, and their responses become correspondingly less fanciful. And they continue to improve beyond that point.
So how are kids doing what LLMs are currently incapable of? How do we teach ourselves not to hallucinate? Or do we, really? I mean, if I tell myself I'm going to make it through the intersection before the light turns red, but I end up running the red light, was I just mistaken, or was that a self-delusion, i.e., a mini-hallucination of sorts? Probably a self-driving car would be less likely to make that category of mistake, so maybe I shouldn't be so smug about being grounded in reality.
A more salient question would be, 'how do I know it ISN'T hallucinating'…
Any evidence to support this claim or just commentary ?
Disclaimer: I know little of this field.
[1] https://www.scientificamerican.com/article/perception-and-me...
[2] https://www.frontiersin.org/articles/10.3389/fpsyg.2021.7289....
It's simply when the predicted probable sequence isn't grounded in reality.
When I ask an LLM to summarize the great water wars of 1999, and how the Trade Union was ultimately defeated by the Antarctic Coalitions hovercraft-fleet under Vice Admiral Zagalow, it isn't "extrapolating" from knowledge of history, it is simply inventing a load of bollocks. But that bollocks will be dressed in fine language and probably mixed in with plausible-sounding references that have a somewhat-logical-sounding relation to the training data.
The problem is, the LM doesn't and cannot know when it produces bollocks.
All it can care about is if the sequences produced are probable according to it's model.
> What is heavier, a small floating passenger ferry or a two metric ton heavy rock that sinks to the bottom of the ocean.
> A two metric ton heavy rock would be heavier than a small floating passenger ferry. The weight of the rock is two metric tons, which is equivalent to 2,000 kilograms or 4,409 pounds. The weight of the passenger ferry would depend on its specific design and construction materials, but it is unlikely to be heavier than two metric tons. Therefore, the heavy rock would have a greater weight than the small floating passenger ferry.
It completely relies on surface information such as "small floating" and ignores the deeper "correlation" that all ferries are heavy.
A two metric ton rock weighs two metric tons by definition (or 2000 kilograms). However, a small passenger ferry, while it may look small compared to large ferries or ships, can weigh much more than two metric tons. Even a small passenger ferry can weigh dozens or even hundreds of tons, due to the mass of the hull, the engine, and other equipment on board.
So, without specific information about the ferry's mass, it's safe to assume that a "small" passenger ferry is likely heavier than a two metric ton rock. However, if the ferry is particularly small and lightweight, or the term "ferry" is being used to describe a very small watercraft (like a raft or dinghy), it's possible for it to be lighter. You would need the specific weight of the ferry to give a definitive answer.
"I'm sorry, but it appears there's a misunderstanding. As of my knowledge cutoff in September 2021, there were no events known as the "Great Water Wars of 1999" involving a Trade Union being defeated by an Antarctic Coalition's hovercraft fleet under Vice Admiral Zagalow. This might be part of a work of fiction, alternative history, or a future event beyond my last training cut-off.
My training includes real-world historical events and existing geopolitical structures, and as of 2021, Antarctica was governed by the Antarctic Treaty System, which prevents any military activity, mineral mining, nuclear testing, and nuclear waste disposal. It also supports scientific research and protects the continent's ecozone.
Please provide more context if this information is from a book, a movie, or a game, or if it refers to something else that I may assist better with."
The result were 2 very well written paragraphs, including the defeat of the Trade unions navy, a ceasfire agreement and a peace agreement ending the water wars.
I rephrased the entire thing as a question, asking the LM to tell me about the conclusion of the war. Again I got a pseudo-historical statement.
Data is still king.
In the real world, search reduces information acquisition costs as you only have to spend time and resources on finding an existing result rather than recreating it.
Unless you mean that Reddit is astroturfed with the SEO garbage you're trying to avoid, in which case this will definitely not help.
Is search on Reddit itself still useless?
Train on dataset A to learn to think, use thinking on dataset B to become an export in B's field.
But until "I don't know" comes out, rather than hallucinations, we're in trouble.
Yea that was what I was getting at with the "combination" of data. The publicly available data provides the base/primary education, then you specialize it with your proprietary data and bam, you have an AI model that nobody else can produce...an actual product moat.
Training on something huge like "the internet" is what gives rise to those amazing emergent properties missing in smaller models (including the recent Pi model). And there are only so many datasets that huge.
But its also indeed a waste, as Pi proves.
There probably is some sweet spot (6B-40B?) for specialized, heavily focused models pre trained with high quality general data.
Linear Regression is great for projections, and can even be fit to time series data using lagging.
By "cramming all of the web" on a model what is really going on is the hidden layers of that network are getting better at understanding language and logic. Imagine trying to teach a kid who doesn't know how to read to learn about a Science by only giving them science textbooks. Chances are they won't get very far.
Building little specialist model's don't really work either. It's like trying to train a parrot to do science, sure it can repeat some of the phrases that you give it, but at the end of the day it's not really making any new connections for you.
I wonder what people said about "bus" back in the day, especially those who knew Latin.
It seems that what we need to make a big leap forward is better reasoning. There is a lot of debate between the GPT-4 can/can't reason camps, but I haven't seen anyone try to argue that it reasons particularly well.
Building the data set for that should be quite trivial.
People who argue against GPT-4 reasoning at all are arguing against clear results. It extremely easy to show examples and benchmarks of 4 reasoning and understanding. The argument then turns into "well that's not "true" reasoning", whatever that means.
The other good use cases are using LLM to turn natural language prompts into API calls to real data.
Doing this a couple of times gives me 100% accuracy for my use case that involves some level of summarization and reasoning.
Hallucinations are not as big of a deal at all IMO. Not enough that I'll just sit there and wait for models that don't hallucinate.
for popular languages, though for JS it most of the time outputs obsolete syntax and code.
Today i tried to do a bit of scripting with my son in Garrys Mod, it uses Expression 2 for a Wiremod module. GPT hallucinated a lot of functions and the worst part it switched almost each time from e2 to lua.
It is good at solving homeworks for students, or solving popular problems in popular languages and libraries though it might give you an ugly solution and ugly code, it is probably trained on bad code too and it did not learn to prefer good code over bad code.
(And I don’t mean to be rude - I just really hate that syntax!)
i mean use at least "let"
What we call hallucination is just when the resulting text is wrong but the underlying probabilities could be high.