We most certainly can, just not as well or as strongly as when we're able to influence the system under observation. You're speaking way too strongly and simplifying a complex mechanism down past anyone's expertise.
If you're saying that interaction in any sense is important, I'd very much agree that unsupervised learning and supervised learning aren't equipped to handle reinforcement learning problems. Correct framing of a problem is necessary to achieve a desired property like causality.
That is the only way we can learn anything, by observing. And what else is there to observe than the "world"? We can observe our own thinking process but that is part of the world too. I would say. It is definitely not "out of this world" :-)
That is a very, very strong statement that requires some proof to go with such a strong statement of certainty.
While we can learn that the ball's change of movement just fine without swinging the bat, we can only do that because we are generalizing from a large body of knowledge that we devoloped by experimenting on the world using our body.
I am not aware of a single piece of evidence that an agent can use purely observational learning to ever aquire the causal knowledge of the real world to sufficient level to make those sorts of inferences with any sort of reasonable accuracy.
In my view, a lot of things that are noninteractively inferred are compositions of more fundamental things that required empirical experience. When you've had the causality of gravity thoroughly beaten into you at a young age, a lot of other things seem intuitive that would otherwise completely fall outside a framework for being unempirically learned.
Do you have a specific counterexample of causality you can infer without interaction or empirical experience of something related?
Caveats: I'm not a neurologist or psychologist, so this is mostly philosophical speculation on my part.
Whether you swing the bat or just watch it hit the ball into the sky, you have the prerequisites needed to reason about the interaction.
A more entertaining question is how a system comes to believe causality (i.e. comes to believe that things can and must have causes)
That's actually the point!
For example, if you're allocating patients in control and treatment group based on a roll of dice - the dice don't have free will, but they are influencing the "system" of the patients and their treatment, while that system is not causually influencing the dice rolls.
For another example, if a baby is "experimenting" by babbling (a key part of language acquisition, https://en.wikipedia.org/wiki/Babbling), it's not necessary or relevant to decide whether free will is involved, the baby obtains useful experimental results of what sound experiences are caused by which attempts to move the tongue/mouth/etc, even if there's no free will and the attempts to move the tongue/mouth/etc are also deterministically caused by the sensory experiences of the baby.
A human can never be smarter than said human: Our brains are not connected and we can't share capacity with others.
So, discovering causality is always an individual experience. And that happens likely by "playing" with "the world".
I think it's noticeable that smarter animals are more playful. Which is also a hint that points to the fundamental importance of interaction with the world as a prerequisite for "smartness". Additionally the capabilities of the "sensors" and "actors" that make interaction with the world possible in the first place seem to be crucial to develop "smart behavior".
The part about the "sensors" seems quite obvious. I think one can gain a better general understanding of some thing if one can "experience" it in more than one "dimension".
And the "actors" allow one to perform "experiments" with the things around one, and find out this way how that thing "works" or is supposed to be "used".
That's actually the behavior that can be observed in children of all kinds of "smarter" species. So it seems to be at least linked somehow to "smartness".
On the other hand, the starting point for an ML system interpreting that same image is essentially a stream of scalar values that tend to demonstrate multiple layers of periodicity (3-4 byte intervals for RGBA and then another layer per line of rasterized image data and yet another per frame if it's a video).
Here's a quick experiment. Let the video linked below play for a five count (sound is essential but a bit intense so maybe moderate the volume first) so you have some confidence it's not just me playing a rude trick, then close your eyes for a five count. There's going to be a major change in the sound when you get near 'five', now try to imagine how the scene changed before opening your eyes again:
https://youtu.be/qnL40CbuodU?t=25
I think without any reference from an embodied perspective, we're asking ML systems to understand the sounds (which are also streams of scalar values that demonstrate periodicity) the same way we interpret the representation visually.
(Also if you enjoyed the example above check out these two channels, some of them are mindblowing)
https://www.youtube.com/user/jerobeamfenderson1
https://www.youtube.com/c/ChrisAllenMusic
And a fun video explaining it all - https://www.youtube.com/watch?v=4gibcRfp4zA
I mean there was some stuff about instrumental variables but it seems the theory is a bit incomplete in that area.
What distinguishes an intervention vs just normal observable randomness? Does it have to do with the complexity of the entity performing the "intervention", with the fact that this entity can observe and act on knowledge? I guess it's kind of the debate about where determinism ends and free will begins. Are there mathematical bounds to help us sort it out though? Maybe there is something information theoretic? It's very unclear in my mind.
This is all about fitting generative models that are robust to counterfactual changes, that remain predictive even if you run your models with data you've never observed, this beyond simple interpolation/extrapolation. Are there priors in model structure that tend to naturally make models much more robust to counterfactual changes and that make models work well beyond the data? Do these priors get more effective when they include some latent variables that distinguish between observations and interventions? How to you train these latent variables?
I think the question that is really on your mind is the one of (causal) identification.
For ML researchers, this may be explained as follows: How, and to what degree, can we find a "deep" parameter that is a general or constant relationship between our inputs and outputs? Is this possible? When? Under which assumptions? If we do this, THEN we can run counterfactuals with our model.
But as it turns out, to convince yourself that your model really generalizes, you will have to prove to yourself that there exists a set of general relationships (parameters) in the data - which may for example be general because they are causal parameters - and that your model is sufficient to "identify" these parameters given the data source. You will quickly find that you need to examine both the model, and the data (and assumptions that you have about it). Generally, no "model" itself is ever causal. It's always the combination of model and DGP.
One way to have a causal model is to simply assume some physical relationship exists, and estimate parameters from data. This is called structural modeling in econometrics and scientific ML in, well, physics mostly. You bring in prior science and knowledge - like how weather fronts move, or how prices relate to costs - and use this structure to contrain your model. If you do this well, then this may be enough to identify causal parameters! Or ranges thereof. Here is a technical treatment of such issues:
http://fmwww.bc.edu/EC-P/wp957.pdf
Another way is to find some measure of interest that you can derive without a complete model. For example, you may be able to "identify" so-called average marginal/treatment effects of some input variable without a parametric model. Then, as long as the changes in your inputs are "exogenous" in some sense, you can get a causal parameter.
In experiments, the differences in inputs are randomized. What remains as group differences between placebo and treatment are the causal effects.
In observational data, you may look for "natural experiments". Here, some exogenous variation can be found that identifies a similar treatment effect between individuals that are affected, and those that are not. Depending on your issue at hand, further techniques may be necessary to identify causal effects.
Let's make an example:
You built your start up to a huge and successful company, made up of programming teams. As CEO, you finally think about realizing your dream: Firing every non-technical team lead (say, everyone with an MBA). What effect would this have on productivity in the short term?
Well, you have lots of data about your teams, and have certainly fired a lot of people, so you build a model. That model tells you, that firing such team leads is super good for productivity.
However, after you do the firing, you find that productivity drops. What happened?
Your model did not identify a causal effect. In this case, the firings you have observed probably occurred because the team lead was simply a bad manager. However, that data does not identify the "casual effect" of firing team leads. The "intervention" or the "treatment" you observed was not "exogenous".
Okay, let's do more science here. Let's say you try to find out the effect of technical ability of the team lead on team productivity. You run your fancy ML models, and again conclude that more technical ability of the manager makes for a better team. But when you implement your effective training measures (or hiring measures, for that matter), the benefit is less than you expect. Again, what happened?
Well, you might have a selection effect. Technically competent team leads may select into productive teams. Or, there is endogenity: productive teams mean that the team lead learns technical stuff, instead of managing chaos. The effect is actually reversed: good teams make leaders more competent. Your estimation - your model - overestimates the causal effect. It is not sufficient to create a counterfactual, simply because you did not find the real causal chain.
Some key words to look into that are probably simpler than Pearl's full framework (and more practical, as they are applied every day):
Natural Experiments, Instrumental Variables, Difference-in-Difference, Synthetic Controls, Matching, Heckman Selection
If you are a person into technical papers, here's an approach that combines ML with causal analysis
The difference is your knowledge that value of some variable is controlled and independent from anything else in this situation. If causal model is some kind of a puzzle, then you could see how it would work if some parts of model are made dysfunctional by your intervention.
X causes Y means that changes in X lead to changes in Y. If you changed X than Y would change also. By observation you can witness correlation, which says nothing about causes. You could add temporal element, like "X happened earlier than Y" it means that Y couldn't be the cause of X. But it doesn't answer the question about cause, it just rejects some hypothesis about cause, and is all uncertain: we could easily devise an example when such a rejection would be a mistake. For example people could act because they are able predict something will happen. So something is not happened yet, but already is the cause of some events.
But when changes to X were made by you, then you know exactly when X was changed, you know that it was no changed to changes of Y or some Z which is a confounder. Here we also could make a mistakes, for example, because we do not know exactly how we make decisions and our behavior could be a confounder or mediator or something like. So we devised randomized controlled trials, double-blind experiments to make experiment really independent.
> I guess it's kind of the debate about where determinism ends and free will begins.
No need to get to that length. If your behavior is independent from the process which you research, than you'll get a "practical free will": free from influences of studied variables. This practical and situational free will definition is enough for learning about causes.
I think metaphysically speaking the approaches (observing interventions vs causal calc) aren’t meaningfully different in terms of inferences you can make with infinite data, see my similar observations to yours: https://vladfeinberg.com/2019/12/01/metaphysics-of-causality...
But if you can presume a fixed DAG you can get away with fewer observations bc then you can derive some minimal/cheap set of vars to randomize over such that the resulting experiment measures a causal effect. All causal calc does is give you a framework for clarifying assumptions necessary to derive such a set.
In high dims performing randomization is exponentially costly.
Memory in these models is used as afterfact, or some side utility for complex iterative routines based on calculus of function optimization. While in living organisms memory and its "hardcoded" shortcuts allow to cut through the search space quickly as in a large database index.
Speaking in database terms we have something like "materialized views" on acquired and genetically inherited knowledge, built from compressed and hierarchically organized sensory data and prior related actions and associations, including causal links. Causality is just a way to associate items in the memory graph.
Error correction doesn't play as much role in storing and retrieving information and pattern recognition, as current machine learning models may lead you to believe.
Instead, something akin to self-organized clustering is going on, with new info embedded in the existing "concept" graph via associations and generalizations, through simple LINK and JOIN mechanisms on massive scale.[1] The formation of this graph in long term memory is tightly coupled with sleep cycles and memory consolidation, while short term memory serves as a kind of cache.
Knowledge is organized hierarchically starting from principal components [2] of sensory data from e.g. visual receptive fields, and increasing in level of abstraction via "chunking", connecting objects A and B to form a new object C via JOIN mechanism, or associating objects A and B via LINK mechanism. Both LINK and JOIN outputs are "persisted" to memory via Hebbian plasticity.
All knowledge including causal links are expressed via this simple mechanism. Generating a prediction given a new sensory signal is just LINKing the signal with existing cluster by similarity.
Navigation in this abstract space is facilitated via coordinate system similar or perhaps identical to the role hippocampal place & grid cells play in spatial navigation. Similarity between objects is determined as similarity between their "embeddings" in this abstract concept space.
It's possible that innate structures are genetically pre-wired in this graph which represent high level "schemas", such as innate language grammar which distinguishes e.g. verb from noun, visual object grammar which distinguishes "up" from "down", etc. It is also possible these are embodied, i.e. connected to some representation of motor and sensory embeddings. And serve to bootstrap the graph structure for subsequent knowledge acquisition. I.e. no blank slate.
The information is passed, stored and retrieved via several (analogue) means both in point-to-point and broadcast communication, with electromagnetic oscillations playing primary role in synchronization in neural assemblies, facilitating e.g. speech segmentation (or boundary detection in general), and coupling an input signal "embedding" to existing knowledge embeddings in short term memory; while neural plasticity/LTP/STDP as storage mechanisms on single neuron level.
[1] See Leslie Valiant "neuroidal" model and his book https://www.amazon.com/Circuits-Mind-Leslie-G-Valiant/dp/019...
[2] See Oja Rule http://www.scholarpedia.org/article/Oja_learning_rule
and Olshausen & Field classic work on sparse coding http://www.scholarpedia.org/article/Sparse_coding
Kids are often an unpleasant annoyance in restaurants, and many people that don't like that annoyance try to convince restaurants and lawmakers to ban them from restaurants. The problem with those ideas is that by banning kids from restaurants, you are just going to create annoying adults in restaurants over time. Kids are annoying in restaurants, but they are also learning how to interact with the world. If you don't find a way to let them explore boundaries, they never learn, and they'll become obnoxious restaurant patrons even as fully grown adults.
Which kind of goes back to ethical AI. You can't unleash unbounded AI on the world, or else you'll cause chaos. And you can't sandbox AI, or it will never truly learn. What are you supposed to do then? I don't know, but the answer isn't firing the ethical AI department because you don't want them criticizing your ad empire ;)
The world must have seen some wild, explosive action in its day.
Immortal organisms still exist (two-headed planaria!), but on the whole the ecosystem seems pretty well calibrated by now.
I think besides the sheer amount of work put in to program something like that, the main limitation would be processing power, that sort of thing would take an immense amount of it.
Now that I think about it, isn't this exactly what Tesla and other self-driving companies is trying to do?
So, it's not like it's sandboxed, it's just very hard to make this "kid" play with things.
There’s a hell of a lot of observation happening in kids minds very early on.
No doubt it’s easier once you can pick up a baseball bat yourself, but I have no doubt a young kid would understand basic objects connected together without ever having used them.
And there is incredible amount of experimentation at that age. They dont learn things by observation alone for sure.
The illusion of causality is incredibly strong -- so much so that it's really hard to get people to give it up, even when faced with paradoxes (First Mover, quantum mechanics, multiple causation, etc.)
We don't come into the world with a fully formed sense of causality, and it appears over the course of months. But at least some of it may be wired into the hardware, like the language instinct. It's just that the wiring isn't done the day we leave the womb, and is influenced by what comes after.
You could observe the results of others acting, but it means that the questions you're getting answers to are outside of your control. So if you need to know the answer to a particular question, you either need to test it yourself, or hope that whoever you're watching will test it for you.
We can learn well from a demonstration, where someone else is acting and we correctly understand the intent and expected results (perhaps they've explicitly communicated that, perhaps we just know implicitly through our previous shared experience) and observe what happens.
However, if we observe someone else - at an entirely different skill level - who is essentially testing a hypothesis (and we don't know that, we have no observation of the agent's internal) and getting useful data about where their world model diverges from reality, then it's not nearly as helpful for us to determine where our world model diverges from reality, that would require different actions and different data.
The article: "“Machine learning often disregards information that animals use heavily: interventions in the world, domain shifts, temporal structure — by and large, we consider these factors a nuisance and try to engineer them away,” write the authors of the causal representation learning paper. “In accordance with this, the majority of current successes of machine learning boil down to large scale pattern recognition on suitably collected independent and identically distributed (i.i.d.) data.”"
The key words are "interventions in the world." The article goes on to say, "“Generalizing well outside the i.i.d. setting requires learning not mere statistical associations between variables, but an underlying causal model,” the AI researchers write." The point being that, whether or not acting in the world is an essential condition for learning causality, current machine learning approaches are not even trying for causality.
Also crucially, we learn these inferences by acting on the world and knowing something about why we acted. "I was playing with it" is a conditional independence statement that we use all the time while learning how things work, we just usually don't use the mathy language to describe it. We're running randomized controlled trials constantly, but implicitly.
Coincidentally, it's a common anecdote that people who seem to learn things quickly and deeply have this curiosity and will play with a thing/twiddle the knobs while they're learning how it works. When you're playing a video game and someone says "hang on, let me figure out the controls for a sec" they're changing the conditional independence structure of their observations and running an RCT.
Held, R., & Hein, A. (1963). Movement-produced stimulation in the development of visually guided behavior. Journal of Comparative and Physiological Psychology, 56(5), 872–876. https://doi.org/10.1037/h0040546
Maybe direct acting is a quicker learning method. Or maybe seeing others' learning processes and instructions allows one to leapfrog ahead by not duplicating mistakes.
Maybe always learn off a good dataset if it exists first?
If you don't test it then you don't know if it works, and you can't improve on it.
Knowledge is derived from experience.
Observation is a type of experience.
If you require actual experience or direct observation to learn, then you're not using your brain to its full potential.
Would a child who has never experienced the human body's pain response be able to infer the causal connection between the heat of a stove and the response after touching it?
Arguably, language is a tool that allows us to generalize the direct experience of other agents. It is unclear if it is possible to remove direct interaction from a learning system and still reach the same level of understanding.
I do agree that observation is a type of experience, but a model that is meant to guide action (basically any useful model) needs to be tested in action. I can't learn to juggle only by watching other people juggle. I can only develop a hypothesis about how one juggles, but to test (and refine) it is to try the hypothesis out.
> a model that is meant to guide action (basically any useful model) needs to be tested in action
No, it doesn't. For example, the vast majority of work on modeling the stock market is done on machines completely sandboxed from any ability to make trades and are owned by companies who will never make a trade themselves but instead return an API response with a yes/no. Whether that is fed directly into some sort of automated action is largely irrelevant as the ability for an individual trade to cause a measurable impact on the market is negligible until it isn't. So, these systems are built separate from the system they model and learn entirely through observation.
tl;dr: weather forecasting models don't have an action to take and also can't influence their system. And yet they learn and grow more accurate.
Except that isn't how that really works right? A mental model explains part of the reasoning that led to an outcome but never all. After all, the map is not the territory. A mental model is grounded in trivial assumptions. Ultimately, your brain produces inferences in a way that is incredibly hard to couple to a specific logical processes. There has been a lot of research on how experts think and reason and none of it is compatible with having a mental model whatsoever.
I'm a constructivist so I'm under no assumption that we have access to the territory at all except modulated by our subjective perception.
I'm confused, you're saying you watched others learn and then managed to get on a bike for the first time and properly ride it like an experienced rider with no practice whatsoever?
We see what did we really do to get this. We often don't have any idea. And we put any label we find most appealing and plausible based on our previous assumptions.
Machines can learn to do the same by going back one step by answering "What could have happened before" in terms of probability.
But for the purposes of machine learning, it seems like it should be possible to learn from observations of someone else’s interventions?
This is important because "learning without explicit instructions" in ML speak is Unsupervised learning (clustering and dimensionality reduction). There are no labels except the ones that you decide upon yourself (cluster membership). Unsupervised Learning is still far in its infancy in effectiveness compared to supervised systems, and it's no surprise that its algorithms are generally extremely easy to implement from scratch (e.g. K-means or DBSCAN) compared to relatively difficult work like automatic-differentiation in neural networks.
Learning by reading information in a book or by direct didactic teaching would be supervised learning. Learning through a dialectical format would be reinforcement learning . Self-supervised learning would be equivalent to autodidactic learning and the creative act upon the world. (maybe the distinction between self-supervised and reinforcement is arbitrary)
The point is that we want to learn as much as we can given the information available to us. We should not rule out the role that the biological analogue to unsupervised learning plays in human development.
It is for all of these reasons that I become far more excited when a new clustering or dimensionality reduction algorithm comes out than I am when a new neural network architecture becomes state of the art.
I have always felt that the signifance of the "social software of human culture" in our general intelligence and learning capacity was underestimated by the AGI community.
So personally, I see more potential in communities of learning agents than any developments in the underpinnings.
Tangentially, wouldn't sampling based ML methods like particle filters / kalman filters or other randomized state space exploration algorithms be analogous to the person learning by acting on the world? In this case, the "action" would be bouncing the radar off the object being tracked.
Of course these models are far more limited than a child in the way they can act on the world, and in the number of aspects of reality they can model.
And furthermore, they have no concept of causality, and represent only the current state of knowledge they are modeling.
- https://www.routledge.com/The-Ecological-Approach-to-Visual-...