Let's try to understand AI monosemanticity
astralcodexten.com
astralcodexten.com
Human brains are also a "black box" in the sense that you can't scan/dissect one to build a concept graph.
Neural nets do seem to have some sort of emergent structural concept graph, in the case of LLMs it's largely informed by human language (because that's what they're trained on.) To an extent, we can observe this empirically through their output even if the first principles are opaque.
Alternatively, what you're seeing are the structures inherent within human culture as manifested through its literature[1], with LLMs simply being a new and useful tool which makes these structures more apparent.
[1] And also its engineers' training choices
I mean we don't live in the chaos realm, atoms make molecules, those molecules make structures, those structures make cells, organs, lifeforms.
yes, sure: you should expect a well trained NN to have an internal structure that approximates that of the statistical distribution of encoded features of the training data into the latent space of the model.
that’s very nearly the definition of “trained”.
Knowledge is compression https://en.wikipedia.org/wiki/Hutter_Prize
To be fair, I don’t see a categorical difference between human intelligence and compression either
Whether and to what extent human culture itself has emergent properties, including some sort of emergent intelligence, is a different question; indeed, I think a more interesting question than that of LLMs, though the dispute over the behavior and nature of LLMs certainly suggests these questions generally, and begs the question of what we mean wrt emergence, intelligence, etc. (My understanding of emergence comes from reading the seminal book, Complexity: The Emerging Science at the Edge of Order and Chaos, and some related literature. As far as I understand emergence, it's not really a close question of whether LLMs exhibit emergence.)
However, I think that real neurology and machine-learning can be mutually reinforcing fields: structures discovered in the one could be applied to the other, and vice versa. But thinking we can create "AGI" without first increasing our understanding of "wet" neural nets is the height of hubris.
But given the exponentially greater chemical, biological, and neurological complexity of human brains, the millions of years of evolutionary "pre-training" that they get imbued with, and the years of constant multimodal sensory input required for their culmination in what's called human intelligence, it takes an extremely bold assumption to believe that equivalent ends can be met with a stream of fmadds driving through in an array of transistors.
It's hard to overstate how big a gulf there is in both the hardware and software between living systems and what we currently see in our silicon projects. Neural network learning in general and LLM's in particular are like crude paper airplanes compared to the flight capabilities of a dextrous bird. You can kinda squint and see that they both move some distance through the air, and even marvel at it for a bit, but there's a long long long long way to go.
However I concur that we may end up creating an AI which thinks nothing like us, has no sense of morality, and yet far exceeds us in intelligence. And that would be somewhat unsettling.
We may build something that far exceeds us in capability, but without us understanding it, or it understanding us. This is the alignment issue.
Yeah, nature does it bottom-up through evolution of self-assembling, self-replicating machines. A lot of "design constraints" that go into living things has to do with... keeping them alive. Compare an electrical wire or silicon trace with a biological nerve or neuron - the former are just simple traces of inert metal, the latter are complex nanomachines mostly dedicated to self-maintenance, and incidentally conducting electricity.
Point being, we can't compete with nature across every dimension simultaneously, but we don't need to, and we shouldn't, because we don't need most of the features of natural solution, not at this point. We don't need Concordes hatching from eggs - we have factories that build them.
That depends on the landscape of the solution space. If the techniques that work much better than anything else happen to be the same techniques used within our own brains, then we might find them, through aggressive experimentation and exploration, before we even realize that it's also how our own brains work too.
[1] possibly bhudist monks excepted.
The plane is in many ways lesser than birds, and we barely understood aerodynamics when the wright brothers gave us a working plane: While the bird is efficient, we had a whole lot more thrust. Our planes are far better than they were back then, but not because we understand bird better, but because we focused on efficiency of the simpler designs we could build. The Formula 1 car doesn't come from deep understanding of the efficient, graceful running movements of the cheetah.
Our results in medicine have been far and ahead of our understanding of biochemistry, DNA, and in general, how the human body works. The discovery of penicilin didn't require a lot of understanding: It required luck. Ozempic and Viagra didn't come from deep understanding of the human body: Resarchers were looking for one thing, and ended up with something useful that was straight out unintended.
We might be able to build AGI with what we know of brains, we might not. But either way, it's not because we need to understand the human mind better: Do we have enough compute for a very crude systems that we know how to build to overcome our relative lack of understanding?
So if you ask me, what is the height of hubris is to think that working engineering solutions have to come from deep understanding of how nature solves the problem. The only reasonable bet, given the massive improvements in AI in the last decade, is to assume that the error bars of any prediction here are just huge. Will we get stuck, the way self-driving cars seem to have gotten stuck for a bit? Possibly. Will be go from a stochastic parrot into something better than humans, the same way that Go AIs went from pretty weak to better than any human in the blink of an eye? I'd not discount it either. I expect to be surprised either way.
It is basically understood how "nature solves the problem". It does that by evolution. And what is machine learning and neural networks but artificial evolution?
Also known as throwing shit at a wall to see what sticks. Evolution is dumb, but it brute-forces its way down the optimization gradients by sheer fucking scale of having turned the entire surface of this planet into tiny, imperfectly self-replicating machines.
The former is where as you describe where better understanding of biology might help but is not a prerequisite for progress, but in the latter it is not just needed but the goal.
Now I know this is a bit of a caricature as both of these disciplines are in practice poorly deliniated and often intermixed. It's easy to find hubristic examples to mock on both sides, but there's value and brilliance to be found in each respectively as well.
You say Go later so I think you know this and it was just a thinko, but chess fell easily to pre-neural techniques that are fully logic based and understandable. It was Go that required a big neural network trained through self play.
Already ~40 years ago I read in a popular science book that scientists had tried to apply genetic algorithms to finding a better nozzle for a given purpose (jet engine? I don't remember that). That was then done by prefabricating a number of ring-shaped metal parts the inside of which had different widths and different slopes; a stack of these could form a wide variety of shapes for the conduit. They started with an initial configuration, did a measurement, then threw the dice to determine what part should be changed; if the new measurement was better (according to their metric) than the preceding one, they kept it, otherwise reversed it. They did find a nozzle shape that was not only significantly better than any known configuration; it was also significantly weirder, seemingly completely ad-hoc.
There's no computers in this experiment, no intelligent design; people were only necessitated to procure the initial parts and setup, do the dice-throwing and so on; the design process proper was a very basic, randomized, mechanical procedure, no thinking required, no understanding necessary.
Each ReLU layer is just a (quasi-)linear transformation, and a pass through two layers is basically also a linear transformation. If you say you want some piece of information to stay (numerically) intact as it passes through the network, you say you want that piece of information to be processed in the same way in each layer. The groups of linear transformations that "all process information in the same way, and their compositions do, as well" are basically the Lie groups. Anyone else ever had this thought?
I imagine if nothing catastrophic happens we'll have a really beautiful theory of all this someday, which I won't create, but maybe I'll be able to understand it after a lot of hard work.
Just dropping the reference here, I don't grok the literature.
And a possibly relevant paper from it:
In one sense, that seems rather close to being linear. If you take a random point (according to a continuous probability distribution) , then with probability 1, if look in a small enough neighborhood of the selected point, it will be indistinguishable from linear within that neighborhood.
And, for a network made of ReLU gates and affine maps, still get that it looks indistinguishable from affine on any small enough region around any point outside of a set of measure zero.
So... Depends what we mean by “almost linear” I think. I think one can make a reasonable case for saying that, in a sense it is “almost linear”.
But yes, of course I agree that in another important sense, it is far from linear. (E.g. it is not well approximated by any linear function)
???
Are you writing off all abstract mathematics as nomenclature gymnastics, or is there something about this connection that you think makes it particularly useless?
1. Proving that a thing or set of things is part of some grouping
2. Proving that a grouping has some property or set of properties (including connections to or relationships with other groupings)
These are extremely powerful tools and they buy you a lot because they allow you to connect new things in with mathematical work that has been done in the past. So for example if the GP surmises that something is a Lie group that buys them a bunch of results stretching back to the 18th century which can be applied to understand these neural nets even though they are a modern concept.
If they have inner symmetries we are not aware of, you can avoid waste in searching in the wrong directions.
If you know that some concepts are necessarily independent, you can exploit that in your encoding to avoid superposition.
For example, I am using cyclic groups and dihedral groups, and prime powers to encode representations of what I know to be independent concepts in a NN for a small personal project.
I am working on a 32-bit (perhaps float) representation of mixtures of quantized Von Mises distributions (time of day patterns). I know there are enough bits to represent what I want, but I also want specific algebraic properties so that they will act as a probabilistic sketch: an accumulator or a Monad if you like.
I don't know the exact formula for this probabilistic sketch operator, but I am positive it should exist. (I am just starting to learn group theory and category theory, to solve this problem; I suspect I want a specific semi-lattice structure, but I haven't studied enough to know what properties I want)
My plan is to encode hourly buckets (location) as primes and how fuzzy they are (concentration) as their powers. I don't know if this will work completely, but it will be the starting point for my next experiment: try to learn the probabilistic sketch I want.
I suspect that I will need different activation functions that you'd normally use in NN, because linear or ReLU or similar won't be good to represent in finite space what I am searching for (likely a modular form or L-function). Looking at Koopman operator theory, I think I need to introduce non-linearity in the form of a Theta function neuron or Ramanujan Tau function (which is very connected to my problem).
I disagree with how you came to this conclusion (because it ignores non-linearity of neural networks), but this is pretty true. Look up gauge invariant neural networks.
Bruna et al. Mathematics of deep learning course might also be interesting to you.
I guess we could also look at it the other way; embedding spaces work this way because the underlying neurons work this way.
1. Neurons encode a concept and activate when it shows up.
2. No, it's way more complicated and mysterious than that.
And now this seems to add:
3. Actually, it's only more complicated than that in a fairly straightforward mathematical sense. It's not that mysterious at all.
I suspect this means that either I'm not picking up on subtleties in the article, or Scott is representing it in a way that slightly oversimplified the situation!
On the other hand, the last quote in the article from the researchers does seem to be hitting the "it's not that mysterious" note. A simple matter of very hard engineering. So, I dunno. Cool!
> Researchers simulate a weird type of pseudo-neural-tissue, “reward” it a little every time it becomes a little more like the AI they want, and eventually it becomes the AI they want.
This isn't the only way. Back propagation is a hack around the oversimplification of neural models. By adding a sense of location into the network, you get linearly inseparable functions learned just fine.
Hopfield networks with Hebbian learning are sufficient and are implemented by the existing proofs of concept we have.
I feel the same way about transformers vs RNNs: even if RNNs are more “correct” in some sense of having theoretically infinite memory it takes forever to train them so transformers won. And then we developed techniques like Long LoRA which make theoretical disadvantages functionally irrelevant.
How’s that?
The reason huge context windows weren't possible in the past is that memory requirements were quadratic with input length. Long LoRA lets us use less memory for our context windows, or use the same memory footprint for larger context windows.
Not sure about actual implementation, but at least for us concepts or words are not pure nor isolated, they have multiple meanings that collapse into specific ones as you put several together
There is a distinction to be made in "knowing how it works" on architecture vs weights themselves.
Or, is it enhanced cognition, on the part of the interpreter having to unpack much from little?
I would recast this: any thinking is a linear superposition of weighted tropes. If you read TVTropes enough you'll start to realize that the site doesn't just describe TV plots, but basically all human interaction and thought, nicely clustered into nearly orthogonal topics. Almost anything you can say can be expressed by taking a few tropes and combining them with weights.
If there's any room for alternatives or unconventional thought, it must now be assigned within an established hierarchy. Or, it must be shoehorned into invisible guardrails.
This is not bad in itself. Of course, useful ideas must transmit successfully to realize progression over decades.
But if one has a choice to focus attention, there also ought to exist the option to ignore it.
I don't mean to "ignore all frameworks" or "everything original must come from the self." I just mean I avoid TVTropes--yet I also embrace Dramatica.
No doubt TVTropes is a form of design patterns for writing, for which it has incredible value. It's just not something I want to "shape my mind around."
As I write this, I know it doesn't make much sense. It may just be a silly thing, and I have yet to "mature to realize the concept of" a shared, tropic mindview.
That said, the post is itself clearly summarizing much more technical work, so my analogy is resting on shaky ground.
Also I can think of some counterpoints to yours: the people who bred teosinte into corn (or any wild grain into a domesticated one) appear to be making conscious choices or direction- that is, they used their intelligence and reasoning from observed examples of pairings to conclude that they could make improved specimens based on selective breeding (without knowing about random mutations of natural selection!).
And if we start to modify human germline then would also be an example of evolution with conscious choice or direction (assuming the modifications became fixed in the population).
Just wondering if I understood you, I don't know anything on the subject.
- a fully deterministic machine (even if the interface and the way OpenAI let people access ChatGPT make it seem like it's non-deterministic, there are fully deterministic models out there, who not only only respond to inputs but also always respond the exact same thing given the same inputs [query, seed, ...]),
and:
- god exists
There could be chaos at work when human thinks. There may be interferences at play, say because whatever element that traveled trillion of kilometers just traversed our brain.
While a fully deterministic machine that always respond to the same input in the same way is just that: a deterministic machine.
P.S: I don't know about other LLMs like Falcon 180b but image-generation models like StableDiffusion are fully deterministic. I think a model is using a broken design and shall quickly hit limitations if it cannot be queried in a deterministic way (and its usecases are certainly limited if repeatability is not achievable). If you want different answers, use a different seed or a different query. But the same query+seed should always give the exact same output.
Are these different ideas of God entirely? Yes: in India there are however many gods and Brahman says these are lesser precisely because they don’t include the whole universe.
Terms like “free will” and “intelligence” are too fuzzy to talk about precisely unless we’re on the exact same page regarding what we mean. And applying our imprecise definitions to machines is not doing us any favors.
The neural network training process wants to minimize the neural network's loss function. The neural network, if it "wants" anything, will have such wants as were embedded in its weights through the process of minimize its loss function, which will mostly be "wants" whose satisfactions correlate with reduced loss function. Of course this is using the term "want" in a behaviourally-descriptive sense, not a subjective-experience sense.
For example, AlphaGo's training routine wants to minimize AlphaGo's loss function. AlphaGo wants to beat you at Go.
It's another thing to anthropomorphize by accident or by illusion, as per pareidolia: https://en.wikipedia.org/wiki/Pareidolia . Just as in pareidolia, where the human brain is primed to "see" a human face in a certain pattern of light and shapes, it seems that human brains are primed to "see" a human intelligence in the output of an LLM, because our brains are pattern-matching on "things that look like human speech". But that's a reason to not anthropomorphize LLMs, precisely because people are inclined to do so without thinking.
Given that, I honestly can't find anything too upsetting.
In any case, anthropomorphism is something I don't mind, mostly. Is it misleading? For the layman. But the domain is one of modeling intelligence itself and there are many instances where an existing definition simply makes sense. This happens in lots of fields and causes similar amounts of frustration in those fields. So it goes.
I feel this is an abuse of the language. Biological neurons and ANN neurons aren’t the same or even all that similar. Brains don’t do backprop for example. Only forward passes. There’s a zoo of neurotransmitters which change the behavior of individual neurons or regions in the brain. Unused neurons in the brain can be repurposed for other things (for example if your arm is amputated).
They're not the same but they're definitely similar.
>Brains don’t do backprop for example.
We've developed numerous different learning algorithms that are biologically plausible, but they all kinda work like backpropagation but worse, so we stuck with backpropagation. We've made more complicated neurons that better resemble biological neurons, but it is faster and works better if you just add extra simple neurons, so we do that instead. Spiking neural networks have connection patterns more similar to what you see in the brain, but they learn slower and are tougher to work with than regular layered neural networks, so we use layered neural networks instead.
The only reason modern NNs aren't closer to their biological counterparts is because they genuinely suffer for it.
The secret of bird flight was wings. Not feathers. Not flapping.
And yet, there's no ANN that's as good at interacting with the real world as the simplest worms we've studied, despite having many times more neurons than those worms have cells.
We are clearly still missing some key pieces of the puzzle for intelligence, so claiming that the difference between ANNs and biological neurons is irrelevant is quite premature. We are far away from having an airfoil moment in AI research.
Silicon simulations of brains may “suffer” from being faithful but this also discounts the advantages that brains have. As I mentioned for example, brains can repurpose neurons for other tasks. Brains can also generalize from a single example, unlike neural networks which require thousands if not millions of examples.
Brains also generally do not suffer from catastrophic forgetting in the same way that our simulations tend to. If I ask you to study a textbook on cats you won’t suddenly forget the difference between cats and dogs.
There is not a single brain on earth that is the blank slate a typical ANN is. "Brains generalize from one example" is pretty dubious. Millions of years of evolution matter.
>As I mentioned for example, brains can repurpose neurons for other tasks.
Isn't this just a matter of the practical distinction between training and inference and not some fundamental structural limitation ?
>Brains also generally do not suffer from catastrophic forgetting in the same way that our simulations tend to.
This suggests CF may well be a simple matter of scale - https://palm-e.github.io/
Since individual anns are much closer to synapses, we don't have anything near the scale of the brain yet.
Of course structure matters, but biological neurons have far more degrees of freedom than those in ANNs. The fact that we even need to keep differentiating between the two is an indication that classifying both as “neurons” is not accurate.
> Isn't this just a matter of the practical distinction between training and inference and not some fundamental structural limitation?
It’s a difference in capabilities of the things themselves. A biological neuron organically seeks out new connections. Sure we could program that into an ANN somehow but the fact that nodes in an ANN don’t have this capability out of the box is a fundamental difference.
> CF may well be a simple matter of scale
For a moment, a big enough network might be able to mirror an entire brain with the lottery ticket hypothesis. But if it takes two or ten or a thousand ANN neurons to simulate the degrees of freedom of a biological neuron, are they really the same?
You are the one saying that biological neurons and ANN are similar...
Since the other comment already went for the evolution and structure angle, I'll go for the other part. What single example? What test have you seen done on the brain capacity of few weeks old fetuses? Our brains start learning patterns in the world before we are even born. How much input does a baby receive every single second from it's eyes and ears and every other sense?
Even when you are "analyzing" a new object for the first time, you receive a continuous stream of sensory input of it. Our brain even requires that to work, if you put a single frame different in a fast enough display, most times you won't even notice the extra frame and your brain will just ignore it.
> Most animals are goal-directed, intentional, sensory-motor agents who grow interior representations of their environments during their lifetime which enables them to successfully navigate their environments. They are responsive to reasons their environments affords for action, because they can reason from their desires and beliefs towards actions.
In addition, animals like people, have complex representational abilities where we can reify the sensory-motor “concepts” which we develop as “abstract concepts” and give them symbolic representations which can then be communicated. We communicate because we have the capacity to form such representations, translate them symbolically, and use those symbols “on the right occasions” when we have the relevant mental states.
(Discrete mathematicians seem to have imparted a magical property to these symbols that *in them* is everything… no, when I use words its to represent my interior states… the words are symptoms, their patterns are coincidental and useful, but not where anything important lies).
In other words, we say “I like ice-cream” because: we are able to like things (desire, preference), we have tasted ice-cream, we have reflected on our preferences (via a capacity for self-modelling and self-directed emotional awareness), and so on. And when we say, “I like ice-cream” it’s *because* all of those things come together in radically complex ways to actually put us in a position to speak truthfully about ourselves. We really do like ice-cream.
I would also like to add that the subject of conversation is artificial and natural neurons, which humans, though contain some, are not.
If a NN is trained to do something, it can be equally considered as "wanting" to do that thing within the autonomy it is afforded, as much as any human.
Human wants are driven by instinct though - i.e. our preferences; if you like women, you like women, if you don’t, you don’t.
Our “output” is in service to those wants.
Current AI doesn’t have instinct / preprogrammed goals - except for goal-driven AIs but the hyped up LLMs aren’t such AIs. Their output isn’t motivated by any goal - a LLM can’t deliberately lie to you to get you to do something; it lies because it doesn’t differentiate between what’s true and what’s false.
What is instinct physically?
> a LLM can’t deliberately lie to you to get you to do something; it lies because it doesn’t differentiate between what’s true and what’s false.
A LLM also can't do multi-step reasoning, yet here we are.
Does it matter? As long as it conceptually exist that’s all that matters.
LLMs don’t seek any goal. It’s advanced autocomplete.
You can’t give it a bunch of facts and a goal then expect it to figure out how to achieve said goal. It can give you an answer if it already has it (or something similar) in its training set - in the latter case, it’s like a student who didn’t study for the exam and tries to guess the right answer with “heuristics”.
> A LLM also can't do multi-step reasoning, yet here we are.
Where’s “here”?
It does matter. Where is it? By evading this you have basically described "instinct" as something computers simply cannot have, axiomatically. That's boring.
> LLMs don’t seek any goal. It’s advanced autocomplete.
These two sentences are contradictory.
> You can’t give it a bunch of facts and a goal then expect it to figure out how to achieve said goal.
Have you used a language model lately? It sounds like you're saying things you think a LLM shouldn't be able to do as if they can't do them. Giving it a bunch of facts and a goal and expecting it to figure out how to achieve the goal is something you can do. It's not perfect, but they can be surprisingly good.
> Where’s “here”?
At a point in time where LLMs can do multi-step reasoning.
Fish don't like ice cream and we don't feel the need to spawn. It's because of how we are built.
Weird, because I'm pretty sure I had the choice whether to respond to this comment.
It wasn't the light waves hitting my retina from an HN post, leading to nerves firing and neurotransmitters all coming together to post this.
I posted it because I have free will. I almost didn't.
Unless you truly feel that there is no free will, and that reality is just a bizarre movie we have to experience .... well, then ... we'll disagree.
Yes well, your neurons don't "want" to do anything either.
>Maybe humans are the same, but in the case of artificial neural networks we at least know it's a simple mathematical function
So what, magic ? a soul ? If the brain is computing then the substrate is entirely irrelevant. Silicon, biology, pulleys and gears. all can be arranged to make the same or similar computations. If you genuinely believe the latter, it's fine. The point is that "simple" mathematical function is kind of irrelevant. Either the brain computes and any substrate is fine or it doesn't.
>Also, an artificial neuron is nothing like a biological neuron.
They're not the same but "nothing like" is pushing it a lot. They're inspired by biological neurons and the only reason modern NNs aren't closer to their biological counterparts is because they genuinely suffer for it, not because we can't.
>Biological neurons fire because of their internal state, state which is modified by biological signaling chemicals
Brains aren't breaking break causality. The fire because of input.
> The fire because of input.
No they do not fire because of input, they modulate their firing probability based on input, and there are different modalities of input with different effects. Neurons are self-contained biological units (descended, let me remind you, from standalone unicellular organisms, just like the rest of our cells), which actually have an independently developing internal state and even metabolic needs; they are not merely a system of logic gates even if you can approximate their role with a system of equations or an ANN. This is very different, mechanistically and teleologically. Hell, even spiking ANNs would be substantially different from currently dominant models.
> So what, magic ? a soul ? If the brain is computing then the substrate is entirely irrelevant
Stop dumbing down complex arguments to some low-status culture war opinion you find it easy to dunk on.
Computation is substrate independent. I'm not saying neurons and ANN weights and «pulleys and gears» are the same. I'm saying it does not matter because what you perform computation with does not change the results of the computation. If the brain computes, then it doesn't matter what is doing the computation.
>No they do not fire because of input, they modulate their firing probability based on input, and there are different modalities of input with different effects. Neurons are self-contained biological units (descended, let me remind you, from standalone unicellular organisms, just like the rest of our cells), which actually have an independently developing internal state and even metabolic needs; they are not merely a system of logic gates even if you can approximate their role with a system of equations or an ANN. This is very different, mechanistically and teleologically. Hell, even spiking ANNs would be substantially different from currently dominant models.
Yes, a neuron is firing because of input. To suggest otherwise is to suggest something beyond cause and effect directing the workings of the brain. If that is genuinely not the case then feel free to explain why, rather than an ad hominin attack on someone you don't even know.
> So what, magic ? a soul ? If the brain is computing then the substrate is entirely irrelevant
>Stop dumbing down complex arguments to some low-status culture war opinion you find it easy to dunk on.
I personally don't care if that's what anyone believes. The intention is not to attack anyone.
If you believe in a soul or the non religious equivalent, that's fine. We just have different axioms.
If you don't believe in a soul(or the equivalent) but somehow think substrate matters then you need to explain why because it makes no sense.
* Analogies aside, neurons are quite different than NN nodes, because each neuron has an incredibly complex internal cellular state, whereas an NN node just has an integer for state.
* A brain is not a "function" in the way that a trained LLM model is. Human life is not a series of input prompts and output prompts. Rather, we experience a fluid stream of stimuli, which our brain multiplexes and reacts to in a variety of ways (speaking, moving, storing memories, moving our pupils, releasing hormones, etc). That is NOT TO SAY a brain violates causality; it's saying that the brain is mechanically doing so much more than an LLM, even if the LLM is better at raw computation.
None of this IMO precludes AGI from happening in the medium term future, but I do think we should be careful when making comparisons between AGI and the human brains.
Rather than comparing "apples to gorillas", I'd say it's like comparing a calculator to a tree. Yes, the calculator is SIGNIFICANTLY better at multiplication, but that doesn't make it "smarter" than a tree, whatever that means.
It's a tautology. If the substrate did change the computation, then it wouldn't be the computation.
Claims where it isn't possible for you to be incorrect may be less impressive than they seem.
Human cognition would be a good example of substrate dependent computation I'd think....it even varies per instance of substrate.
You can move your pointer anywhere you'd like, it is ultimately tautological. Infinite regress is a bitch lol
Say, have you taken into consideration the role consciousness and culture are playing here? Like this "reality" you are describing, do you know what the actual, biological/scientific source of it is? :) But now I'm kind of cheating, aren't I...I think we're not supposed to say that part out loud! ;)
It's not dumbing down. It's extracting the crux of the matter that the complexity of arguments is trying to hide, perhaps unintentionally. Either the brain implements a function that can be approximated by a neural network thanks to universal approximation theorem, or the function cannot be approximated (you need arguments for why it is the case), or magic.
This is not to say that the human brain leverages quantum effects. It's just a well known example where the hardware and a specific algorithm can be shown to matter.
I also think it's strange to describe the brain as implementing a function. Functions don't exist. We made them up to help us think about building useful circuits (among other things). In this scenario, we would be implementing functions to help us simulate what is going on in brains.
I suspect there's some fundamental metaphysical framework protection in play here, the sort of language being used is pretty common, I believe it to be a learned cultural behavior (from consuming similar arguments).
Neurons also don't "respond" to specific input either. They can't speak or provide an answer to your input.
These are all just abstract metaphors and analogies. Literally everything in computer science at some point or another is an abstract metaphor or analogy.
When you look up the definition and etymology of "input", it says to "put on" or "impose" or "feed data into the machine". We're not literally feeding the machine data, it doesn't eat the data and subsist on it.
You could go on and on and nitpick every single one of these, and I don't think the use of "want" (i.e. anthropomorphizing the networks to have intent) is all that bad.
Anyways, this is Scott’s writing style. I recall an earlier ACX post on alignment that was real heavy on ascribing desires and goals to AI models.
This seems wrong. God-zilla is using the concept of God as a superlative modifier. I would expect a neuron involved in the concept of godhood to activate whenever any metaphorical "god-of-X" concept is being used.
Some more:
cheaper sensors (happening now)
better sensor integration (happening now, kind of)
better tools for ml grokking and intermediate engineering (this article, kind of)
better tools for layering ml (probably the same thing as above)
a new model for insurance/responsibility/something like this (unsure)
better communication with people inside and outside the car (barely on the radar)Using models to generate/score/rank/modify data to be more useful as training data is a very interesting angle.
One particularly interesting topic is the "theories of superposition" section, which gets into how LLMs categorize concepts. Are concepts all distinct or indistinct? Are they independent or do they cluster? It seems that the answer is all of the above.
This ties into linguistic theories of categorization[2] that I saw referenced in (of all places) a book about the partition of Judaeo-Christianity in the first centuries CE.
Some categories have hard lines - something is a "bird" or it is not. Some categories have soft lines - like someone being "tall." Some categories work on prototypes, making them have different intensities within the space - A sparrow, swallow, or robin is more "birdy" than a chicken, emu, or turkey. Apparently Wittgenstein was the first to really explore with Family Resemblances that a category might not have hard boundaries, according to people who study these things.[3] These sorts of "manifolds" seem to appear, where some concepts are not just distinct points that are or aren't.
It's exciting to see that LLMs may give us insights into how our brains store concepts. I've heard people criticize them as "just predicting the next most likely token," but I've found myself lost when speaking in the middle of a garden path sentence many times. I don't know how a sentence will end before I start saying it, and it's certainly plausible that LLMs actually do match they way we speak.
Probably the most exciting piece is seeing how close they seem to get to mimicking how we communicate and think, while being fully limited to language with no other modeling behind it - no concept of the physical world, no understanding of counting or math, just words. It's clear when you scratch the surface that LLM outputs are bullshit with no thought underneath them, but it's amazing how much is covered by linking concepts with no logic other than how you've heard them linked before.
[1] https://transformer-circuits.pub/2023/monosemantic-features/...
[2] https://www.sciencedirect.com/science/article/abs/pii/001002...
σ(C*σ(B*σ(A*x)))... = y
repeated matrix-vector multiplication, separated by nonlinear functions, σ(). A,B,C etc here are your weights, x is your input vector, y is your output vector. Any nonlinear function will work, some behave better than others, but they must be present, otherwise you could simplify the problem by apply the associative property to the weights side of things, and you'd effectively have a single layer that can only simulate linear functions. A hidden layer is just any vector that is intermediate to that calculation, for example the result of σ(A*x) would be the first hidden layer.
So right off the bat, you obviously have a problem: "Examining the code" means examining the weights, which are just gigabyte/terabyte sized matrices. Opaque is putting it mildly. The only sane approach is to start from one of our known-meaningful vectors (the input or output), and work our way inwards from there, either seeing what elements of hidden layer vectors have the most significant value when a certain input is applied, or determining what hidden layer vector values produce values closest resembling a desired output.
Then you start running into the problems described in the article.
But as a simple example, convolution would be a little tedious to describe in the general notation you wrote above without getting into the image dimensions, stride, padding etc., not to mention residual layers or norm layers that are commonly used. Then there are things like stop gradient, dropout or even other training targets which are only used during training but not inference.
It's pretty much the quintessential example of an enforced symmetry, in that it introduces a symmetry against translation.
Saying that understanding that code implies understanding how the network works, is like saying “John wrote this emulator for GBA games, therefore he must understand how the game [some GBA game] works”.
Except, instead of the game being written in assembly, it was written in malbolge.
“Still, last month Anthropic’s interpretability team announced that they successfully dissected of one of the simulated AIs in its abstract hyperdimensional space.
(finally, we’re back to the monosemanticity paper!)
First the researchers trained a very simple 512-neuron AI to predict text, like a tiny version of GPT or Anthropic’s competing model Claude.
Then, they trained a second AI called an autoencoder to predict the activations of the first AI. They told it to posit a certain number of features (the experiments varied between ~2,000 and ~100,000), corresponding to the neurons of the higher-dimensional AI it was simulating. Then they made it predict how those features mapped onto the real neurons of the real AI.
They found that even though the original AI’s neurons weren’t comprehensible, the new AI’s simulated neurons (aka “features”) were! They were monosemantic, ie they meant one specific thing.
Here’s feature #2663 (remember, the original AI only had 512 neurons, but they’re treating it as simulating a larger AI with up to ~100,000 neuron-features).”
Oh and it doesn't always use logic. So there are no ORs and ANDs and IFs, just +, *, exp, max, etc.
Thanks for the heads up! AI is better at making up stories than humans already. Hodl my beer while I go buy some BSitCoins.
So what exactly is the point, except "look at us, we are so clever"?
It's when you read it with the "at last, I can understand how reasoning and inference with meaning is going to emerge from this" you have a problem.
It's a great read but what you bring to it, informs what you take from it.