We know how common molecules can interact with each other. Does this mean that anything built on top of them is not intelligent?
Everything in life could just be "statistics".
We know how common molecules can interact with each other. Does this mean that anything built on top of them is not intelligent?
Everything in life could just be "statistics".
People who thing organic adaption, sensory-motor adaption, somatosensory representation building... ie.., all those things which ooze-and-grow so that a paino player can play, or we can here type... are these magic?
Well I think it's exactly the opposite. It's a very anti-intellectual nihilism that all that need be known about the world is the electromagnetic properties of silicon-based transitors.
Those who use the word "magic" in this debate are really like atheists about the moon. It all sounds very smart to deny the moon exists, but in the end, it's actually just a lack of knowledge dressed up as enlightened cynicism.
There are more things to discover in a single cell of our body that we have ever known; and may ever know. All the theories of science needed to explain its operation would exhaust every page we have ever printed. We know a fraction of what we need to know.
And each bit of that fraction reveals an entire universe of "magical" processes unreplicated by copper wires or silicon switches.
As a result of the combination of this method of thinking and the Dunning-Kruger effect, people in our field tend to apply this to the entire world, even where it doesn't fit very well, like biology, geopolitics, sociology, psychology, etc.
You see a lot of this on HN. People who seem to think they've figured out some very deep truth about another field that can be explained in one hand-waving paragraph, when really there are lots of important details they're ignoring that make their ideas trivially wrong.
Economists have a similar thing going on, I feel. Though I'm not an economist.
And that's the heart of the problem. The CSci crowd have a somewhat well-motivated inclination to treat abstractions as real objects of study; but have been severely misdirected by learning statistics without the scientific method.
This has created a monster: the abstract objects of study are just the associations statistics makes available.
You mix those two together and you have flat-out pseudoscience.
I am not sure what more scientific method you could propose. And we can, in this field produce actual reproducible experiments. Really, more so than any other field.
There are no experimental conditions, no causal properties, no modelled causal mechanisms, no theories at all. "Replication" means that you can reproduce an experiment designed to validate a causal hypothesis.
Fitting a function to data isnt an experiment, it's just a way of compressing the data into a more efficient representation. That's all ML is. There are no explanations here (of the data) to assess.
Take the research into Loras for example. Surely the basic scientific method was followed when developing it. You can see that from the paper.
Obviously the results can be reproduced. Unlike in many other fields, reproducibility can be pretty trivial in CS.
Training a model isn’t really a science, but the work gone into creating the models surely is.
This can be plainly seen when trying to get a model to replicate its input.
Some models perform better in fewer steps, some perform worse for many steps, then suddenly much better.
How is discovering these properties of statistical models NOT science?
"the reason prompt Q generates A1..An is because documents D1..Dn were in the training data; these documents were created by people P1..Pn for reasons R1..Rn. The answer A1..An related to D1..Dn in so-and-so way. The quality of the answers is Q1..Qn, and derives from the properties of the documents generated by people with beliefs/knowledge/etc. K1..Kn"
This explains how the distribution of the weights produces useful output by giving the causal process that leads to training data distributions.
The relationship between the weights and the training data itself is *not* causal.
Eg., X = 0,1,2,3; Y = A,A,B,B; f(x; w) = A if x <= w else B
w = 1 because the rule x <= 1 partitions Y st. P(x|w) is maximised. These are statistical and logical relationships ("partitions", "maximises").
A causal relationship is between a causal property of an object (extended in space and time) to another causal property by a physical mechanism that reliably and necessarily brings about some effect.
So, "the heat of the boiling water cooked the carrot because heat is... the energetic motion of molecules ... and cooking is .... and so heating brings about cooking necessarily because..."
heating, water, cooking, carrot, motion, molecules, etc.. -- their relationships here are not abstract; they are concretely in space and time, causally effecting each other, etc. etc.
Was physics not actually a science until we uncovered quarks, since we weren’t sure what caused the differences in subatomic particles? (I’m not a physicist, but I hope that illustrates my point)
Keep in mind most ML papers on arxiv are just describing phenomena we find with these large statistical models. Also there’s more to CS than ML.
I need to use my hand, a pen and paper to draw a mathematical formula. That formula (say, 2+2=4) expresses no causal relationships.
The whole field of computer science is largely concerned with abstract (typically logical) relationships between mathematics objects; or in the case of ML, statistical ones.
Computer science has no scientific methodology for producing scientific explanations -- it isnt science. It is science in the old german sense of just "a systematic study".
Scientists conduct experiments in which they hold fixed some causal variables (ie., causally efficiacious physical properties), and vary others, according to an explanatory framework. They do this in order to explore the space of possible explanations.
I can think of no case in the whole field of csci in which there are cases where causal variables are held fixed; since there is no study of them. Computer science does not study even voltage, or silicon, or anything as physical objects with causal properties (that is electrical egnineering, physics, etc.).
Computer science ought just be called "applied discrete mathematics"
So if I observe some phenomena in a bit of software that was built to translate language, say the ability to summarize text.
Then I dig into that software and decide to change a specific portion of it, keeping same all other aspects of the software and its runtime, then I notice it’s no longer able to summarize text.
In that case I’ve discovered a causal relationship between the portion I changed and the phenomenon of text summarization. Even though the program was constructed, there are unknown aspects.
How is that not the process of science?
Sorry if this is just my question from earlier, rephrased, but I still don’t see how this isn’t a scientific method.
I aspire, at best, to be one of the children outside the zoo laughing. I fear I might be the monkey who stole the key...
its a purely emotional provocation and a universe sized leap, not an argument for LLM's having intelligence or sentience. Anything could be anything, wowww! This goes back to what the other person was saying, "I cannot reason about what is going on behind the curtain, therefore..."
Whether LLMs have intelligence depends on your definition of intelligence.
Could connections of artificial neurons arranged in certain way as a result of training on data yield in human level intelligence?
Remember also, that companies like OpenAI are tracking what prompts fail and adding it to their datasets. So their initial data is ever-more just a record of questions-and-answers.
Given this, we should expect that the vast majority of questions we have, of ChatGPT will be answered very very well.
What has this to do with intelligence? Nothing at all.
Intelligence is not answering questions correctly. It's what you use when you dont have the answers and arent even clear on the question.
What you're describing sounds like RLHF, which changes the style of responses and impacts things like refusals but does not add to a model's intelligence (in fact it reduces model intelligence).
An LLM's intelligence comes from pretraining in which there are no prompts, or answers, only corpus and perplexity.
Statistics is an analysis of association, not of causation. The frequency of (Q, A) pairs follows a distribution that is not constrained, or caused by, or explained by, how Q and A are actually related.
For example, recently there was some scandal at microsoft that if you used "pro choice" in prompts you got "demonic cartoons". Why? Presumably because "pro choice" are symbols that accompany such political cartoons in the data set.
So does Q = "pro choice", and A = "cartoons of hell" occur at notable frequency because hell has caused anything? Or because there's a unique semantic mechanism where by "pro choice" means "hell" and so on.
NO.
It is absolutely insane to suggest that we have rigged all our text output so as to align one set of symbols (Q) alongside another (A) such that, Q is the necessary explanation of A. I doubt this is even possible, since most Qs dont have unique As -- so there is actually *no function* to approximate.
In any case, your whole comment is an argument from ignorance as I complained in mine. What you don't know about life, about machines, about intelligence justifies no conclusions at all (esp., "everything in life could be").
And let's be clear. Lots of people do know the answers to your questions, they arent hard to answer. It's just not in any ad companies interest to lead their description of these systems by presenting good-faith research.
Everything printed in the media today is just a game of stock manipulation using the "prognosticator loopwhole" whereby the CEO of nvidia can "prophesy the future" in which his hardware is "of course" essential -- without being held to account for his statements. So when that stock hits its ATH and crashes, no one can sue.
I think we should change this; remove this loophole and suddenly tech boards and propagandists will be much much more reserved.
OK, here's something much stranger. Suppose you see your friend touch the fireplace, he recoils in pain. Do you touch it? No.
Hmm... whence statistics? There is no frequency association here, in either case. And in the second, even no experience of the fireplace.
The entire history of science is supposed to be about the failure of statistics to produce explanations. It is a great sin that we have allowed pseudosciences to flourish in which this lesson isnt even understood; and worse, to allow statistical showmen with their magic lanterns to preach on the scientific method. To a point where it seems, almost, science as an ideal has been completely lost.
The entire point was to throw away entirely our reliance on frequency and association -- this is ancient superstition. And instead, to explain the world by necessary mechanisms born of causal properties which interact in complex ways that can never uniquely reveal themselves by direct measurement.
But consider these LLMs and such as extremely crude simulations of biological neural networks. They aren't just any statistics; these are biomimetic computations. Then we can in principle "do science" here. We can study skyscrapers and bridges; we can study LLMs and say some scientific things about them. That is quite different than maybe what is going on in industry, but AFAIK there are lots of academic computer scientists who are trying to do just that, bring the science back to the study of these artifacts so that we can have some theorems and proofs, etc. That is - hopefully - more sophisticated a science than trying to poke and prod at a black box and call that empirical science.
For some ways in which these differ:
- real neurons are vastly complex cells which seem to perform significant non-trivial computations of their own, including memory, and can't be abstracted as a single activation function
- real neural networks include "broadcast neurons" that affect other neurons based on their geometric organization, not on direct synapse connections
- there are various neurotransmitters in the brain with different effects, not just a single type of "signal"
- several glands affect thought and computation in the brain that are not made of neurons
- real neural networks are not organized in a simple set of layers, they form much more complex graphs
And these are just talking about the structure. The way we actually operate these networks has nothing in common with how real neural networks work. Real neural networks have no separate training and inference phases: at all times, they both learn and produce actionable results. Also, the way in which we train ANNs, backpropagation with stochastic gradient descent, is entirely unrelated from how real neural networks learn (which we don't really understand in any significant amount).
So no, I don't think it makes any sense to say that ANNs are a form of biomimetic computation. They are at best as similar to brains as a nylon coat is to an animal's coat.
I'm coming from the context of theoretical models of computation, of which there are only so many general ones - Turing machines, lambda calculus, finite state machines, neural networks, Petri nets, a bunch of other historical ones, ... etc. etc. Consider just two, the Turing machine model, versus the most abstract possible neural network. We know that the two are formally computationally equivalent.
Abstractly, the distinguishing feature of theoretical neural networks is that they do computations through graphs. Our brains are graphs (and graphs with spatial constraints as well as many other constraints and things). The actually-existing LLMs are graphs.
Consider, C++ code is not only better modeled by the not-graph Turing machine model, it is also easily an instance of a Turing machine. These man-made computers are instances as well as modeled by von-Neumann architectures, which can be thought of as a real implementation of the Turing machine model of computation proper.
I think this conceptual relationship could be the same for biological brains. They are doing some kind of computable computation. They are not only best modeled by some kind of extremely sophisticated - but computable - neural network model of computation that nobody knows how to define yet (well, Yann LeCun has some powerpoint slides about that apparently). They are also an instance of that highly abstract, theoretical model. It's a consequence of the Church-Turing thesis which I generally buy (because of aforementioned equivalence, etc.): if one thinks the lambda calculus is a better model than neural network for the human brain, I'd like to see it! (It turns out there are cellular models of computation as well, called membrane models.) But that's the granularity I'm interested in.
In different words, the fact that many neural network models (rather, metamodels like "the category of all LLMs") can be bad models or rudimentary models is not a dealbreaker in my opinion, since that is analogous to focusing on implementation details. The goal of scientific research (along the neural network paradigm) would be to try sort that out further (in the form of theory and proofs, in opposition to further "statistical tinkering"). Hope that argument wasn't too handwavy.
Now, whether CPUs are an instance of a Turing machine or not is again quite debatable, but it's ultimately moot.
I think what matters more for deciding whether it makes sense to call a model biomimetic or not is whether it draws more than a passing inspiration from biology. Do practitioners keep referring back to biology to advance their design (not exclusively, but at least occasionally) or is it studied using other tools? Computers are obviously not biomimetic by this definition, as, beyond the original inspiration, no one has really looked at how mathematicians do their computations on paper to help build a better computer - the field evolved entirely detached from the model that inspired it.
With ANNs, admittedly, the situation is slightly murkier. The majority of advancements happen on mathematical grounds (e.g. choosing nonlinear activation functions to be able to approximate non-linear functions; various enhancements for faster or more stable floating point computations) or from broader computer science/engineering (faster GPUs, the single biggest factor in the advancement of the field).
However, there have been occasional forays back into biology, like the inspiration behind CNNs, and perhaps attention in Transformers. So, perhaps even by my own token, there is some (small) amount of biomimetic feedback in the design of ANNs.
>After all, Turing very explicitly modeled it after the activity of a mathematician working with a notebook. The read head represents the eyes of the mathematician scanning the notebook for symbols, the write head is their pencil, and the tape is the notebook itself.
My feeling on this is complete opposite to yours. To me, this is completely valid mode of discovery, and possibly even what led to the thought of the Turing machine. We are after all, interested in mimicking/reproducing the we way think. So it's perfectly sensible that one would "think about how we think" to try and come up with a model of computation.
I dont care at all about this argument of whether to call something biomemic or not. Thats just semantics. What you associate with meaning "biomemic" is subject to interpretation and one can only establish an objective criteria for it by asserting ones own mental model is the only correct one.
I'm not sure if you thought I was being sarcastic, but what I was describing there is literally how Turing came up with the idea, he describes this in the paper where he introduces the concept of computable numbers [0]. I just summarized the non-math bits of his paper.
If you haven't read it, I highly recommend it, it's really easily digestible if you ignore the more mathematical parts. This particular argument appears in section 9.
[0] https://www.cs.virginia.edu/~robins/Turing_Paper_1936.pdf
Who said that? You make it sound like this was some important trend in the past, that got derailed by the evil statisticians (spoiler: there never was such a trend that was big enough to have momentum).
Your rant against statistics is all nice and dandy, but when you have to translate that website from a foreign language into English automatically, when you ask ChatGPT to generate you some code for a project you work on, or when you are glad that Google Maps predicted your arrival time at the office correctly, you rely on the statistics you vilify in essential ways. You basically are a customer of statistics every day (unless you live under a rock, which I don't think you do).
Statistics is good because it works, and it works well in 90% cases which is enough. What you advocate for so zealously (whatever such a causally validated theory would be) currently doesn't.
Since Hume it become popular to somehow rig measurement to make it necessarily informative (Kant), or to claim that measurement has no necessarily informative relation to reality at all (in the end, Russell, Ayer et al.).
It took a while to dig out of that nightmarish hole that philosophers largely created, back into the cold light of experimental reality.
It wasnt statisticians who made the mess; it was first philosophers, and today, people learning statistical methods without learning statistics at all.
Thankfully philosophers started actually reading science, and then they realised they'd go it all wrong. So today, professional research philosophy is allied against the forces of nonsense.
As far as the success of causal explanations, you owe to that everything, including the very machine which runs ChatGPT. That we can make little trinkets on association alone pales in comparison to what we have done by knowing how the world works.
citation needed
All measurements and all experiments ever done with matter and fields confirm that it behaves according to the laws of quantum mechanics. Those laws leave absolutely no room for a self that is not an emergent phenomenon of some kind. They also don't leave room for something like a free will that allows "you" to control "your body" by your "will", which is what I assume you might mean by a soul. That is, they clearly show that me writing this reply could have been (in principle) foretold, or at least had a calculated probability.