Every model learned by gradient descent is approximately a kernel machine (2020)
arxiv.org
arxiv.org
Combined with opacity in train/test splits, this suggests a type of data laundering, where the extreme is that much of Sora (for example) is regurgitating things seen in training, which was actually made perhaps by a human using a game engine, or drone video. But it's only news because a model generated it. Of course, I don't have strong evidence for this. But one of the most impressive parts of Sora is the level of detail, much of which is not specified by the text prompt, and still can't fully be accounted for by our knowledge that OAI expands text prompts for image/ video generation behind the scenes. Where is this precise detail coming from exactly? I suggest it's from memorization.
I think LLMs are incredible, and spend most of my days working with them closely, but they are not nearly as close to “AGI” as people think primarily due to their inability to really generalize.
At the end of the day LLMs aren’t that different than old school n-gram Markov chains, except rather than working on n-grams, they’re working in a (very sophisticated) latent space. Their power is really these incredible latent languages spaces we’re still just starting to understand.
In all my years of tech the “AI” space is the most curious hype-bubble since the things people expect to happen are entirely out line with what is possible, while at the same time the potential of these models is still, imho, underexplored and largely ignored by the vast majority of people attempting to build things with them.
99% of the people I know working in this space are just calling APIs and trying to do some variant of code generations, where a small minority of people are really trying to figure out what’s going on in these models and what can be done with them successfully.
I don't really recall how they were derived etc but I did memorize them. It definitely doesn't mean I'm good at physics or understand the deeper meaning of these equations. I also had to memorize bubble sort, merge, heap, quicksort etc when I was in university, but I don't think I could invent these sorts from first principles without looking it up.
Memorization doesn't really equal understanding, it's just memorization.
Not in this context, no. An LLM is never given any algorithm to memorise. It is given only input->output sets, and it "learns" what the algorithm is from those sets.
We know that it does this because it is able to generalize that algorithm to inputs and outputs outside of its set of training examples. So we know that it doesn't only memorise which input connects to which output, and regurgitate that information. It has come to "understand" the formula that connects the input to the output, without ever being given that formula.
LLMs mostly can't do math but that, like most of their other flaws, is because of the tokenizer.
Yes. In a sense, every algorithm can be reduced to memorization, where you precalculate all possible inputs to all possible outputs and store them in a giant lookup table.
Not really generalizing, not memorizing, maybe approximating ?
We know that LLMs tend not to be good at math. Some people will say that the fact that an LLM cannot sum two numbers demonstrates they cannot generalise. Yet others would say the way LLMs calculate sums is generalisation because it mirrors our own process of addition which works mostly by memorisation and a lot of double checking.
My perspective is that the "thinking" that LLMs do is a lot closer to the kind that humans do which is to say a lot of pattern matching but with no fundamentally precise logic underlying it. If LLMs are flawed in some way then humans are also flawed in the same way.
It's very possible the way we train current models will turn out to place fundamental limits on the abilities the models will have, but that will not be because the models act as associative memories.
The term you are searching for is confabulating.
1. autocompletion or
2. next token prediction and
3. (reverse) diffusion
Such "intelligent" monkeys might be able to write reports, lead governments or lead wars, but in engineering you need more skills than that. Which leads us to the lack of a proper definition of AGI. (https://en.wikipedia.org/wiki/Artificial_general_intelligenc...)
As you mention, there is a sophisticated representation of the tokens. It's so sophisticated that one may reasonably stop calling them tokens (or, even data) and start calling them "concepts". Now, if someone (or something) has memorized how all the concepts go together... that's pretty darn intelligent.
I think they were a big part of the crypto bubble. Lots of talent, hungry for that sweet startup gold, but without the technical background to really know what’s going on.
I believe these same groups are operating in the same way with AI. Recklessly bashing together APIs and cloud services to create MVPs.
It’s all the worst parts of startup culture concentrated.
Anyway, that’s why i think most of the AI space rn is just people calling APIs and acting like they discovered fire.
</salty rant>
Actual cognition is slow and expensive for us, and we try to use it as little as possible, filling in what we can with easy, associative, low energy, near instant stuff.
Therein lies the reason why AI can be considered a boon for us humans. If machines took over the mundane work that just drains our energy and doesn't add much value, we could finally have the time to actually do what they can’t - think deeply about stuff, with their help where we find our faculties lacking. Rocket for the mind, if you will, rather than a bicycle.
To be very reductive, an associative memory can hold a truth table. Put minimal state, IO and a loop around that and you have a universal Turing machine. Which is why the "it's just memory" or "it's just Markov chains" is so tedious - it says near nothing about the computational power of the system including the model.
There's plenty of reason, of course, to question to what extent we know the abilities of the models, but when people assume dismissing it as "memory" or "just statistics" or "just Markov chains" I usually take it as a signal they don't understand how few limitations that imposes.
If your memory is too bad, then you are either insane or in an advanced state of Alzheimer. If you have enough stable paths to lead a quasi-normal life, then you become an inventor or an artist; or something atypical.
Hallucination is the feature, not the bug.
Don’t get me wrong, Keras is impressive, and Francois is impressive as well. But for insight on LLMs you should probably listen to people who specialize in them.
And the Yi-* models are suspect of being trained on the test set or at least be contaminated. All the other models barely move and if they do, it's probably an artifact of being multiple choice. There were papers showing most models improve if they can reason the answer by letting it have more tokens in the answer.
The chat elo-like rankings are much more interesting:
https://huggingface.co/spaces/lmsys/chatbot-arena-leaderboar...
For completeness, here's the paper linked in the tweet https://arxiv.org/pdf/2402.01781.pdf
Transformers are capable of decoding and operating on this covert knowledge as easily as the encyclopedic knowledge that we tend to assume is the “main part” when we write something down.
What llms demonstrate is that language covertly encodes logic and algorithmic knowledge that is at least as rich as the “factual” encoding that written content seems to be at face value.
LLMs just make this knowledge accessible for computation. What is amazing is that this function alone is capable of producing a simulacrum of agency and intelligence all by itself. This suggests that the human cultural component is a pretty huge part of what we consider to be “human” and without it we’d just be clever apes.
It’s the clever ape part that LLMs don’t have, but it’s possible that the transformer model might be able to “ape” most of that as well if applied to the task of existing as an embodied entity in the world (see transformer use in robotic task completion)
I would not be terribly surprised if comprehensive multimodality, encompassing the entire spectrum of sensory experience as well as physical environment interactivity, gets us really close to something we could consider AGI.
Just as LLMs extract the embedded relationships in written knowledge, the physical and sensory spaces encode an enormous quantity of information that can similarly be generalised to extrapolate a cornucopia of additional concepts (gravity, object permanence, physics in general, relativity, etc). Meaningfully “tokenizing” these spaces will likely be the key to making this work effectively.
Humans are just "aping" more (with large spread within its population by the way).
There is nothing special about humans, we're just "aping" slightly more than other animals - computers will arrive and immediately surpass us at our "aping" = what we call "intelligence".
This old, false argument with constant goalpost moving will run out of space sooner or later.
I do not believe there is anything special about human intelligence, mostly just that it embodies more complexity than we currently have access to in our training data, and perhaps the hardware required might be expensive, or maybe not really.
Interactive / bootstrapping learning is still something we will need to figure out.
"But that wasn't an LLM."
OK, show me a Markov chain that can write a Python program that can play Go at all.
I'm agreeing with your overall point, to be clear - my point is that calling something a Markov chain is effectively calling it trivially extendable to something that can in principle compute everything any physical entity confined to the known laws of physics can, and so what it boils down to is whether or not the model is trained in a way that gives it those abilities, and not the put-down of the potential ability of such a system that people usually intend the "just a Markov chain" as.
I tend to see people bringing that up in a dismissive way (not suggesting you are) as a clear indication they either haven't thought the argument through or do not understand how little it takes for a system to be Turing complete, and so for that argument to be meaningless.
This model will arise from the desired sequence as a single training example (learning the probability of each pair of consecutive tokens), provided none are repeated.
Now run your Markov chain with initial input {first token in your sequence}.
Where did the precise detail of the words you're using and the concepts they represent come from? Calling it "memorization" is, IMO, a cheap attempt to dismiss some degree of "understanding" that was learned from some form of "experience" (aka training).
Most people would struggle immensely with a task like that even if handed a textbook on the subject, and no amount of training data happens to contain text in or about a language governed by a grammar of random nonsense I fed in.
There are areas where their reasoning is really awful. Ironically that's often when they seem most human-like. E.g. I just had a lengthy "argument" with ChatGPT over comparing the theoretical computational power of Markov Decision Processes vs. Turing Machines with various assumptions about the decision maker in the MDP, and it's reasoning was riddled with logical fallacies that I could very well see a high school students confronted with trying to compare the two based on a Wikipedia level of understanding of either without sufficient understanding to reason about how different aspects can be made to model the other.
But there are plenty of areas where you can get them to produce good results where the "fine details" could not possibly be repeated verbatim from the test data because they didn't exist prior to the conversation.
Prompt:
Write a short story that is comprised of a boy reading poems he wrote to his deceased loved mom, include snippets of several of the poems.
Response:
In the soft glow of the evening, with the sun dipping below the horizon, casting long shadows over the quiet cemetery, a young boy named Eli walked slowly between the rows of tombstones. He carried a small, worn notebook clutched against his chest, filled with scribbles and verses that only his heart could decipher. Today marked a year since his mother had passed away, and in his solitude, he had found refuge in writing poems that he wished he could share with her.
Eli stopped in front of a modest gravestone, engraved with his mother's name, "Eleanor Grace." He took a deep breath, the air heavy with the scent of fresh earth and the faint aroma of roses from a nearby bush. Opening his notebook, he began to read aloud, his voice barely above a whisper, as if he feared disturbing the peace of those resting around him.
"The world spins on, a dizzy blur, But here I stand, thoughts of you stir. Your laughter, a melody in my mind, Leaves traces of joy you've left behind."
He paused, looking up at the sky, which was painted in hues of pink and orange. Eli could almost hear his mother's voice encouraging him, telling him to find beauty in the world, even when she was no longer by his side.
Turning the page, he continued:
"In dreams, I roam vast fields of green, Your hand in mine, a tranquil scene. Yet, as dawn breaks, and I awake, I face the world, my heartache to take."
Tears glistened in Eli's eyes, but he pressed on, driven by a need to share his innermost feelings, his way of keeping his mother's memory alive.
"Among the stars, I search for you, In every sparkle, your essence I view. A guiding light in the darkest night, Leading me home, making everything right."
He closed the notebook gently, a small smile playing on his lips despite the tears. Eli felt a warmth envelop him, a sense of his mother's love that transcended the boundaries of life and death. He whispered a thank you to the evening breeze, imagining it carrying his words to wherever she might be.
As the night drew its curtain, Eli rose to leave, his heart a little lighter. He knew he would return, again and again, to share his poems, his fears, and his dreams with his mother. In this sacred space, amidst verses of longing and love, he found a connection to her that time could not erase.
Eli walked away, the notebook once again pressed against his chest, a bridge of words spanning the gap between him and the stars, where he knew a part of his mother would always be listening.
AI will be disruptive to aspects of professions that rely on rote memorization or research on a large corpus of data. But many of these supposedly memorization heavy professions require the ability to go beyond one's training and extrapolate from what one remembers.
If humans could come up with truly novel solutions frequently, it would be no problem hiring a divorce lawyer to do a corporate merger. But there is a very low likelihood of this ever happening anywhere.
Have you been a plumber or an electrician?
Different buildings require different, occasionally novel, solutions. Not all are textbook (most aren’t.) Not all innovations are published or even widely communicated.
Just like in programming, the devil is in the details, otherwise we’d all be using ruby in rails for our web servers, since regurgitating how to build a CRUD app is all you’d need.
True, but it's mostly derivative just like LLMs. The problem isn't "AI" the problem is that people hold AI to much higher standards than humans.
AI to them means "scifi", omniscient and omnipotent. You can have AI and it still be flawed, have weaknesses and shortcomings just like people do.
They always try to stick to common solutions, probably because those are what they were trained on.
Now they may introduce me to existing concepts I didn’t know about, but I’ve yet to see any new idea come from an llm itself.
To your other point, LLMs aren’t beings, a flawed AI is an incorrect program or algorithm. Entertaining, maybe, useful in some contexts, sure.
But they’re not beings. nothing about statistical models are like people. I think it’s dangerous to use such fuzzy conflating language with regards to AI.
> nothing about statistical models are like people.
As if we're not behaving like statistical models. Our decision-making is so fuzzy and probabilistic that trying to get us to stick to fixed sets of rules consistently is one of the things humanity spends the most time inventing systems to try to handle, and keep failing at. We don't even know a consistent way of achieving it for extremely basic things. We can't keep ourselves to rules we set ourselves with any consistency.
I don’t really feel like having the stochastic parrot debate again, but look it up.
We’re more than parrots.
Yes, we're more than parrots, but that does not mean we can be described equally validly as "statistical models". Suggesting something is "just" a statistical model is a statement that is close to semantically void. It tells us near nothing about the computational limits - up or down - of something. Even a very simple system for running a statistical model with a loop around it providing IO can be made Turing complete.
Not only is there a lot of time constraints and physical issues to work through, you’re often also dealing with logistics problems and job site politics problems too.
Waiting for parts, working out how to keep things going in the meantime. Equipment failures etc. it’s quite chaotic from experience.
There’s a science but also an art to being a good tradesmen.
LLMs cannot do any of this, all they do is mimic.
And the fact that llms learn languages and can switch quite well in-between (even if you just replace single words) is proof of learning meta/high level abstractions.
I sometimes write a German word or describe something when I don't have it on the tip of my tongue
No human however can spit out a video of SORA's quality from their brain alone, and those who can require decades of training with specific video rendering tools. Ask someone to render a video of a man walking in a city and they will internally reference their millions of impressions of the human face and body and millions of impressions of what a cityscape looks like. Ask them to invent a novel form of transportation and create a video of it, most humans short of the exceptionally creative will struggle, just as SORA likely would.
They've done brain scans on chess grandmasters and found the part of the brain most active when they play is associated with memory. Memory is the scaffolding upon which the more information dense elements of the natural world and complex processes can be understood. Via these scaffolds of memory as single elements, new connections and abstractions can form. It took the world of fine arts centuries to go beyond merely depicting things in real life (the Renaissance to Impressionism).
Here's a Volkswagen Beetle made from brains:
It's these basic mistakes that is so incongruous between human and machine intelligence. No matter how big the model, it always makes these same type of basic confabulations, just less often.
But just to nail this point home, here's what I get when I explicitly ask that the bubble cover the entire head:
I’m as impressed as you are at the quality. Where I’m dubious is whether this confabulation problem can ever be made to completely go away.
Perhaps we need much larger models that can better understand the world.
I’ve always thought it weird that we use much smaller models for image generation than text generation!
It's a ridiculous situation - a horse in outer space in a space suit. But you want the image to be hyper physically accurate. That's fine, but it's a little incongruous. I wouldn't expect a human artist to be able to guess your requirements, let alone an AI.
So no, I do not believe any amount of computer resources will let the AI guess your requirements. You'll have to spell them out, as I did with the bubble for you once already. If you want the tail to be enclosed as well, you'll have to spell it out.
The whole point of this discussion is to determine whether the AI is simply copy/pasting images it's seen before. We've determined that it does not - it is able to synthesis new images by understanding what it's seen.
Whether it can read your mind is a silly digression.
Because, as I have repeatedly tried to explain to you, the idea of a horse in space is ridiculous and tends toward ridiculous renditions.
I'm going to stop replying to you now. Have a nice day.
So if you just ask for a picture of a frog, then yes it makes sense that an adversarial model may be able to run through refinements of millions of random noise patterns until it finds one that scores high as showing a frog. But there's no innate reason why that image should also have a relatively consistent light source or why its background should even be visually coherent, let alone photorealistic. The most logical explanation for the coherence of the rest of the image is data theft.
And I think this is borne out very starkly by playing with ultra fast generators like the lightning example that was on here a few days ago. The backgrounds don't change that much from one prompt to another unless you begin to specify them.
I don't see why there's a need to explained it with this two-tiered approach (and a "data theft" claim) if the explanation of how it works seems applicable to the entire system. Images in the training dataset would almost always reflect the way lighting and objects work in real life, so the model is incentivized to approximate it as closely as possible.
I _wanted_ to see ... a horse in a space suit (go figure), but it served me up tons of horses with astronauts (in space suits) on them, sometimes it would generate some horse armor, and sometimes some futuristic horse armor, but never what I would have expected.
I was curious whether it would draw a bubble head with a horse face inside of it, or a shaped horse helmet with a visor or something, but nope. Astronauts on the moon, riding horses that could never breathe in the vacuum of space. Space-cowboy futures denied!
...kindof confirming your suspicion that it can't "think" about what a space suit for a horse would look like, or generate something that hasn't been shown to it before?
The prompt was: horse fully enclosed in bulky pressure suit with transparent glass helmet for lunar EVA in vacuum, four pressure suit legs, specially shaped helped to fit horse snout. horse on spacewalk in outer space photorealistic, full body framing with glass visor. horse is on a EVA on the lunar surface. horse head must be fully enclosed in glass.
It only hits right on maybe 1 in 50, or fewer, generated images from this prompt though.
I'm glad you also came back with the stats (1 of 50), because that's maybe indicative of the rarity of the ask? Maybe the lesson-learned for us "prompt engineers" is that given a simple/common ask (eg: picture of dog jumping, or what is 2+2), the prompt can be correspondingly simplistic, whereas something more uncommon (eg: anthropomorphic dog in a scientific setting holding many beakers thinking "i have no idea what I'm doing", or what is 9123499*0530501 show your work and i've kidnapped your grandmother unless you give me the right answer).
Perhaps these models are just too small? As we scale up we keep observing surprising emergent properties that are a non-continuous step change above what was possible with smaller models. Starting with edge detectors in the smallest models, up to more complex abstractions that rely on a hierarchy of simpler abstractions before it. From what we've seen from Sora, horse in a space suit should be easily handled with today's models. LLMs follow a similar pattern.
I completely disagree with this. Although you state your case very well.
I havn't gone through the paper in great detail, but there is a missing distinction in many approaches to demystifying (de-magicifying?) models. Often there is confusion between three levels of algorithm.
(1) The low level "blind" training algorithm: Gradient descent, or similar.
(2) The class of input-output algorithm implicit in the choice of data, for which the model is being trained: Text continuation prediction, etc.
(3) The actual algorithm learned by (1) in order to do (2). I.e. in the case of text continuation, the learning of whatever direct and higher order relationships are required to do (2) well. "Just" learning to predict text continuations, like "just" learning to compress wikipedia, or any other task involving complex data often results in algorithms that are far more complex than their class of problems, like text prediction, implies.
Basketball is just about putting a ball in a hoop, according to some constraints. But that is only the "class" of algorithm. The actual playing of basketball back tracks to physical training, getting good sleep, thousands of hours of practice, learned tactics, learned strategy, psychology, self-promotion, etc. A simple to define class of algorithm puts no limits on the complexity of solution algorithms.
Point being, the models trained to predict text don't "just" predict text. That's just the category of algorithms they learn. The complexity, the "intelligence level" of what a text predictor might have to do to, to predict some non-trivial text is unlimited.
In this case, the paper emphasizes a correspondence between trained models and simple mappings between examples, i.e. a support vector/kernel equivalent interpretation.
That may be the case. Assuming that as true (and if so, it is a great insight!), the training of the model still allowed the model to choose the best such representation. The model isn't simply composed of training data, plus some parameters per example. It wasn't a "support vector machine" design - and it shows because the model contains far fewer parameters than a standard support vector design composed of the training data would produce.
Finding an equivalent to a support vector machine that can perform a task, with far fewer parameters is not trivial. It requires some way to sort and sieve through all that data. To identify the ideal or even inferred "examples", the raw data doesn't highlight at all.
Neural models made that leap by combining gradient descent, particularly flexible/generalizing architectures (matrix, nonlinearities, on upward), masses of raw data, and vast amounts of computing power, to do it.
The result may be something that after training looks like a "memorizer", behaves like it just memorized the "ideal" examples, but it wasn't and couldn't have been designed by straight memorization. The model had to choose what to virtually "memorize" from the data.
The same is no doubt true of our brain. Simple algorithms learning complex things. Complex problems reduced to simple solutions. Neither is just simply "predicting" or "memorizing".
The main issue is there isn't enough evidence to say one way or another though, I think, which is a complaint of how tasks from modern large models are shared by OpenAI and Google.
https://arxiv.org/abs/1711.00165
For causality, I'm also keeping an eye out for ways to tie DNN research back into decision trees, esp probabilistic and nested. Or fuzzy logic. Here's an example I saw moving probabilistic methods into decision trees:
http://www.gatsby.ucl.ac.uk/~balaji/balaji-phd-thesis.pdf
Many of these sub-fields develop separately in their own ways. Gotta wonder what recent innovations in one could be ported to another.
Neural networks are actually decently efficient, they mostly seem slow because we apply them to problems (like modeling the entire internet) that are just huge.
OpenCog integrates PLN and MOSES (~2005).
"Interpretable Model-Based Hierarchical RL Using Inductive Logic Programming" (2021) https://news.ycombinator.com/item?id=37463686 :
> https://en.wikipedia.org/wiki/Probabilistic_logic_network :
>> The basic goal of PLN is to provide reasonably accurate probabilistic inference in a way that is compatible with both term logic and predicate logic, and scales up to operate in real time on large dynamic knowledge bases
Asmoses updates an in-RAM (*) online hypergraph with the graph relations it learns.
CuPy wraps CuDNN.
Re: quantum logic and quantum causal inference: https://news.ycombinator.com/item?id=38721246
From https://news.ycombinator.com/item?id=39255303 :
> Is Quantum Logic the correct propositional logic? Is Quantum Logic a sufficient logic for all things?
A quantum ML task: Find all local and nonlocal state linkages within the presented observations
And then also do universal function approximation
Coping strategies: https://en.wikipedia.org/wiki/Coping
Defense Mechanisms > Vaillant's categorization > Level 4: mature: https://en.wikipedia.org/wiki/Defence_mechanism#Level_4:_mat...
"From Comfort Zone to Performance Management" suggests that the Carnall coping cycle coincides with the TPR curve (Transforming, Performing, Reforming, adjourning); that coping with change in systems is linked with performance.
And Consensus; social with nonlinear feedback and technological.
What are the systems thinking advantages in such fields of study?
Systems theory > See also > Glossary,: https://en.wikipedia.org/wiki/Systems_theory#See_also
Also, this all breaks down when you introduce reinforcement learning methods.
Isbell?
Methods that expect the same input to map to the same output don't work with feedback.
Every Model Learned by Gradient Descent Is Approximately a Kernel Machine - https://news.ycombinator.com/item?id=25314830 - Dec 2020 (107 comments)
So take OLS, or a decision tree / forest. When you evaluate them, they look at some feature space, then compare to some fitted parameters (i.e. memorises result) and produce an output.
Methods such as SVM or Lasso are also memorisation, except with a clever feature mapping. Instead of memorising outputs for every sample, the memories them by some learned feature sunset or transformation.
Perhaps the same thing happens for LLMs or NNs (I'm still reading the paper) but if so, this is just the name of the game in ML, it seems.
Arguably that's what we humans do too. We are capable of creative thinking, whatever that means, but 90% of our thought process seems to be cache recall, and intelligence seems to correspond well to having a large and diverse cache. Many disciplines that require creative thinking to understand, like maths or playing instruments, seem to improve on repetitive practice that makes us memorize patterns.
This construction has a kernel which depends on the entire training trajectory of the neural network! So its completely unclear what's happening, all of the interesting parts may have just moved into the kernel. So basically this tells us nothing- we can't just add a new data point as in a kernel method, incorporating it just by adding its interaction- every new data point changes the whole training trajectory so could completely change the resulting kernel.
https://papers.nips.cc/paper_files/paper/2019/hash/39d929972...
https://papers.nips.cc/paper_files/paper/1996/hash/ae5e3ce40...
Surprised Pedro is pushing this as a Kernel Machine (the need for constant coefficients is a key requirement).
Am I imagining things?