errr... all NNs are just optimisations of an associative probability objective: P(Y|X), they are by definition "statistical parrots". There isn't anything to prove or disprove.
People offering prompts as evidence are people who fundamentally do not understand the basics. NNs aren't strange empirical objects, they're specified by mathematical rules whose properties are known ahead of time.
Any property of a trained NN is derivative of a property of a formula P(Answer|Prompt, TrainingData)
This is an associative statistical relation, which by definition, selects elements of TrainingData by-association with the Prompt.
If the basis by which you understand LLMs is putting prompts into ChatGPT you're severely underqualified for drawing any conclusions about LLMs, and radically subject to confirmation bias.
>People offering prompts as evidence are people who fundamentally do not understand the basics
This I agree with, but I also don't see anyone doing that in this thread?
ie., it's just fitting capacity.
The "emergent boundary" is just an empirical measure of the necessary fitting capacity of these models on "everything ever digitised in english" given any particular functional requirement.
All the language around this area is not scientific, nor are these practices. This is superstitious neophyte engineers, hopped up on scifi, by giddy VCs who love to be told they're funding captin picard.
I've been to many academic conferences, and the fresh PhDs who pump out this BS are not, err, very credible seeming people. Yes, they're young and naive, and really desperate to make their career impactful -- etc.
But they're also not really the kinds of smart sceptics you'd hope for. Many are, though they tend to exit at masters level and go make money.
The academic publishing environment today is far far far away from epistemically hygienic, indeed, in these areas i could only imagine how insane credible sceptical empirical types would be.
Could you imagine being surrounded by this desperate need to call 'useless correlations' "hallucinations", and to call "model fitting", 'emergence' ?
I find it disquieting at a removed distance from it -- I'd be quite mentally ill within it.
I don't do this cultish game playing. And I wrote exactly the same here on AI, crytop and the reset of it many years ago when it would be downvoted; and the same today now it's upvoted.
This is the degree i'm inclined to make public contributions on these matters. I haven't the temperament to handle doe-eyed PhDs repeating universal function fitting theorems and turing-equivalence talking points without understanding a single iota of applied statistics, scientific epistemology, basic methdological scepticism etc.
Unless I am hired to do so, which occasionally I am. But thankfully in actual corp environments the people I meet are far more credible and sceptical than those writing hype papers
Whatever sort of internal structures that are being formed during the training process is somewhat evident when looking at the structure of a CNN… edge detect kernels emerge, etc.
Whatever sort of internal structures formed during the training of transformer based architectures are basically unknown at this point.
There’s your thesis, have at it! You won’t be ignored for your discoveries. Hell, send me a copy and I’ll make sure anyone who matters at UT or Baylor sees a copy!
The reason CNN weights obtain 'recursive-hierarchical representations' of pixel-pattern geometry in the training data follow from the recusrive-heirachical relationship of their weight matrices and from the geometry of the 'pixel space' from which the training data is drawn.
This is certainly interesting; and there's something magical feeling about 'principles of least action' at work. Indeed, many new physicists have a kind of schizophrenic reaction to discovering action principles -- since it imparts to nature a strange apparent conspiracy.
Of course, the job of any good physicist is to be sceptical of this conspiracy, and to get to the heart of how 'accounting tricks' performed by moving objects over time create this illusion.
Likewise this is the job of any good ML researcher; yet they do the oppoiste. Rather than get to the heart of this apparent conspiracy, they call it 'emergence' -- this offends my sense of what the virtues of a scientist ought be.
In any case, on the matter of the LLMs obtaining 'useful' weights for any given task here the job of the researcher is circumstantial, empirical, and sceptical: go and find those 'accounting tricks' within the training data that give rise to this apparent conspiracy of the system to acquire a useful state.
There is no emergence: there is just a set of weights which compress the structure of a target space. At some point this set is large enough, and 'lies across the space like a mental chain does a gate'.
Emergence is an ontological relation between parts and wholes whereby wholes arent reducible to their parts because of ontologically-relevant interaction properties between their parts which aren't intrinsic properties of them.
The fluidity of water emerges out of hydrogen bonding which does not occur when you isolate H20 alone. There is no such relationship here.
This ontologising of the formal, this language which gives a causal-physical semantics to purely formal properties of abstract models -- this is pseudoscience. It's done as part of a computational-idealist worldview in vogue because it's a helpful language for VC investment were-changing-the-world hype.
The formal properties of NNs cannot be described in these terms, because they do not have ontological relationship -- they have formal (mathematical, statistica, etc.) ones.
https://hai.stanford.edu/news/ais-ostensible-emergent-abilit...
and silly me! You are absolutely right, research should have a meta analysis ready of papers that came out… 6 months ago.
Absolutely ridiculous.
I am sure you have read this paper, and have an army of meta analyses and counter claims for this nascent sub field.
Yes there is, that's all there is.
The term “parrot” is used to imply inference by something akin to a look-up table, specifically it is used to indicate poor out-of-sample performance and a lack of a proper world model. The optimization objective is irrelevant when determining the generalization performance of a model and when judging whether it can reason beyond looking up answers in a table.
As the user above noted, it is now quite well established that GPT-4 has impressive out-of-sample performance which can be explained by it possessing an actual model of the world and not being a “parrot”.
Yes it’s impressive. Yes it’s got amazing zero shot performance in domains.
But there’s a pattern of failure in production which describe a limit, that shouldn’t exist if the emergent properties were stable.
You can build this right now and test it.
Build a sequence of agents to work on a domain you are not an expert in.
Let them loose. See what happens.
Do the same thing on a domain you have expertise in.
Assume the number of errors you find, the number of modifications you have to make are stable for other domains.
There may be a subtle correlation between properties needed to answer a specific out-of-sample request and in-sample features.
Unfortunately, prior to training/testing and without recognizing that correlation in the data set, I believe it's impossible to guarantee the model will include it. (Corrections welcome)
So claiming that out-of-sample performance is a mirage, would be a bridge too far?
Err... I can show this is false, kinda trivially. People who engage in prompt-confirmation-bias aren't aware of what the in-sample is.
It's basically everything ever digitised: you can ask it for the first paragraph of every dickens novel, to what the average petal length of an iris flower is -- etc.
How are you measuring the in-sample here?
If you engage in straightfoward reasoning from first principles, and are basically aware of what the training data is, you can show in 10 seconds critical failures of generalisation.
If you want a recipe: go find some fringe api docs. Establish that it has been trained on them. Then, since they're fringe there wont be much code on github, etc. Now ask it do something non-trivial with that API. It will fail, and the mechanism will be obvious: it'll jam in correlated code that lacks relevance.
Do the same on a popular API, and see it succeed.
The in-sample will be obvious for both, and the bounday of generalisation
I am sure you will continue to argue that this is still in line with everything-thats-ever-written prediction but my opinion is that at that point, it's a meaningless distinction. The human brain is also just a machine.
LLMs are enough to be a brain
LLMs are not enough to be a brain.
Being sceptical, as every person ought in these matters, I changed the finical data and performed the same analysis (both in a new tab, and within the same convo). The results were the same!
How strange?
Well, in being reference financial data ChatGPT was reporting prior reference summaries of it. When that data was changed it was reporting the very same reference summaries (which were now wrong).
Since it's incapable of actually summarising financial data. It's only capable of selecting combinations of pieces of its training set.
Now, is this distinction "meaningless" ?
No, it's the difference between this guy being fired for causing a massive loss on a major project; and this guy keeping his job and doing it well.
It's not, though. It is in fact able to summarize financial data, just as it's able to write code and diagnose a medical condition. It makes mistakes, yes, even grave ones, much more so than experts in those fields would.
Do you see a difference between the process of adding numbers and dividing by their count (taking a mean) and emitting numeric tokens which are most probable for a given input?
The former is called "taking a mean" the latter isnt. This system never engages in any method to summarise financial data. It's method is always the same: to emit tokens most probable given a set of historical tokens.
It's the difference between saying "the average of 1,2,3" is 2 because that sentence occurs 1,000,000 times and saying it's 2 because you've literally computed it.
This system does not run financial summary algorithms. It's a trick
Third completely off misconception from you today.
This is not at all what it is doing. "Supercharged Interpolation" is false and makes no sense. It's not a lookup table either. It doesn't memorize enough of what it needs to to make your assertion possible.
all statistical learning is a variation on k-nn (see the relevant paper on this) but likewise this is obvious a priori
k-nn is the ideal learner, and a good starting point for analysis
the question for any given system is: what is the learning space, what is the distance function, and how many points are being considered
NNs set up a compressed X,y space, in that space choose points via an empirical expectation, and obtain a weighted average as their prediction
That's just what they do -- there isn't any other mechanism here. The whole formal structure of the NN can be written down on a page of paper
your paper above doesn't deal with this -- it's a reply to the 'forced interpolation' view, which i haven't espoused. but often NNs are forced interpolated
'extrapolation' is of course a part of the possible predictive output of a statical learning system -- in that it's latent space is taken to be embedded in R^n and so one can 'veer off' into R.
Whenever you attribute a higher fidelity space to a small latent space you are, in effect, extrapolating
No you cannot.
>That's just what they do -- there isn't any other mechanism here.
That's not what they do. They are many papers now showing ICL demonstrating some kind of optimization method during inference which would not be happening if all they did was retrieval.
I'm come to realize you don't know what you're talking about. Your level of denial is scary to see.
more than all every written -- and so on
perhaps apply a single drop of scepticism to this credulity
even, just ask chatgpt to repeat the first paragraph of some book -- say, a dickens novel
In what reality is a character taking up only a few bits ? Certainly isn't this one.
Your denial is so weird. Why does ICL use Higher-Order Optimization Methods when it's just a lookup table ? https://arxiv.org/abs/2310.17086
How do you explain small language models that have much less space than the text they were trained on even with your nonsensical calculations ?
LLMs are not lookup tables and there's plenty evidence to support that. You look insane insisting they are.
>even, just ask chatgpt to repeat the first paragraph of some book -- say, a dickens novel
Not only will this not work for the vast majority of books it sees during training, Why only the first paragraph ? are you not insisting they memorize everything ? Why can't it repeat the whole page ? the whole chapter ? the whole book ?
Of course the first paragraph of any famous novel is in itself famous, so I also asked for the first paragraph from Chapter XL, from the same novel. Interestingly the first 12 words were correct, but the rest was hallucination.
Original: It was fortunate for me that I had to take precautions to ensure (so far as I could) the safety of my dreaded visitor; for, this thought pressing on me when I awoke, held other thoughts in a confused concourse at a distance.
ChatGPT 3.5: "It was fortunate for me that I had to take precautions to ensure the safety of my treasure, for with such a boy as Compeyson on the loose, I was not likely to find much mercy at his hands. I had not told Herbert about the money, and I was glad to have the chance to count it before he came back. I found that I had still twenty-one sovereigns, six shillings, and sixpence left."
No. If anything, we lose predictive power which is why it's extra silly.
If those skills were real, why do they fizzle out on production data ?
Me for starters. If it works, I can quit.
Next up are Karpathy and the CTO of OpenAI. Around July and September both talked about production challenges.
AI ops was the largest subcategory of fall YC startups.
Every single ml and LLM ops individual who gets far enough deals with evals.
I don’t know, but maybe - just maybe- the issue is that people don’t understand themselves enough, to avoid assuming too much of those emergent properties.
As I recall there was also a paper that pointed out the issues with how LLMs are measured, and that the emergence of properties was not a step change once the tests were updated.
Edit- found it: https://hai.stanford.edu/news/ais-ostensible-emergent-abilit...
Yes. We can infer from this that out of sample hallucinations will be closer to the desired productions when measured by token productions metrics than by any other more informative domain relevant metrics.
Which is exactly the case.
If you ask llm to solve a math problem it haven't seen its response will be closer to the desired solution in its linguistic form rather than in mathematical meaning.
This isn't true. Wild how people will confidently say nonsense about things they obviously haven't actually tested.
GPT-4 can manage arbitrary arithmetic calculation closer to the real value than you could ever do without an external tool.
Like those: https://www.mathschool.com/locations/andover/news/prepare-fo...
Here's it solving the last two problems on the Grade 7-8 Russian Math Olympiad from the linked page.
https://chat.openai.com/share/e90d4711-aa38-45c5-97e7-a9a04e...
https://chat.openai.com/share/0c9c579a-ca9f-4aa6-bac6-624372...
The two answers (84 and 225,792) agree with the answer key. I only gave it the questions and didn't give it the PDF either, so it didn't cheat and just read the answer off the answer key in the PDFs.