Roboticists discover alternative physics
phys.org
phys.org
I did see a presentation a few years ago on a similar topic by Erwin Couwans, author of Bullet physics engine, who was discussing doing neural inference of physics. He basically wanted to replace Bullet with a neural network, which I thought was kind of funny at the time, but cool if it worked. (Physical laws at this level of mechanics are mostly quite well understood, but the devil is in the details and actually solving time-discretized non-smooth dynamics, as well as performing good contact detection is less than obvious the more you get into it, not to mention friction, so I could see why a "learned" solution could be attractive, but at the time I found it funny that he was proposing to learn the physics that were already simulated by Bullet.) Looking back, it appears that those ideas culminated in a paper [2] two years ago, which I'll have to read now -- from the abstract, the differentiability of the simulation has benefits not only for learning the physics (mostly friction models it appears), but for learning controllers for plants within the simulated physics. It seems to be not cited in the linked article but maybe it is a slightly different, but related, topic. In any case, fascinating stuff.
[0]: https://www.science.org/doi/10.1126/science.1165893
On the other hand, there are lots of places where "functions" are useful, that have nothing to do with "intelligence", but where actually writing the function with full fidelity can be difficult or intractable. These are also opportunities where learnable function approximation can provide some great benefit, provided you can figure out how to pose the problem such that data is available for the learning part.
A good example is in physically-based rendering, you have lots of complicated aspects of light transport for which we can model the physics quite well, but when you get into complicated reflections and scattering, etc., this can all be modeled by a complicated function called the BRDF [0]. Hand-writing a good BRDF is possible, and quite typical of high-end renderers in fact, but it's no surprise that there's been research in replacing it with a "neural BRDF" [e.g. 1].
That's just to give an example of a place where a single, very targeted and small (but complicated) part of a larger framework can stand to benefit from data-driven modeling, and a neural network can be one good way to do that. Another example is similar usage in computation fluid dynamics [2], where we can hand write a pretty good model, but to capture what is missing, it can be useful to have an approximator. The problem there is that having approximated the function well, it doesn't necessarily lead to better human understanding of the phenomenon. Discovering the true, sparse latent variables in a way that is interpretable, instead of a black box, is a useful step towards that. (Which I guess the current article is aiming at, but I haven't read it in full yet.) But sometimes all you want is results, as in the case of synthesizing a good controller. For example in CFD, if you can use the blackbox model to generate a good stabilizer for an ocean platform, you don't really care about the physics, as long as it's accurate enough to be trustworthy. So the utility of these methods, like most things, is relative to what your goals are.
[0]: https://pbr-book.org/3ed-2018/Color_and_Radiometry/Surface_R...
[1]: https://arxiv.org/abs/2111.03797
[2]: https://github.com/loliverhennigh/Computational-Fluid-Dynami...
We know the leaves of a house plant will bow in the wind, because you have observed it in the wild, as a child several times. Even given non-still image training material, the nn would have to learn all of the physical properties of all the things from scratch, like a human.
To train a physics model like this to derive the properties of material from the world is going to be bound to hilarious failures. A mountain with smoke hanging over it, is obviously a house plant.
In this case, you wouldn't be aiming to identify the material properties specifically, but given a model that's tightly fitted enough, there should be constants that stand in for the missing material properties when it's all said and done.
Just an initial thought, I'd love to hear why I'm wrong from those in the field.
Lets take another extreme, a floaties palmtree. How does on derive from the observed visual differences, that this palmtree is stiff and bouncy.
Or circumstantial decorations as pysic change indicator. If there is snow on the palmtree, the tree might break in strong winds.
To develop this physics reasoning is quite a larger part of childhood and i think the idea, that one can train this into a Neural Net beyond simple "all things are solid and bound to gravity" with references alone is quite a challenging endavour.
For my own curiosity, I wonder if some of the variables they're missing are based more on the context of the video, rather than the content. Like, colors, for instance. Maybe if you ran everything through in black and white you could get it closer to an integer lower bound?
[0] https://en.wikipedia.org/wiki/Takens%27s_theorem
[1] https://clgiles.ist.psu.edu/papers/NC-2000-learning-chaos-nn...
If anyone else is interested in this line of work I recommend checking out Kathleen Champion, Steve Brunton, and J. Nathan Kutz's work on Discovering governing equations from data by sparse identification of nonlinear dynamical systems(https://www.pnas.org/doi/full/10.1073/pnas.1517384113).
Also this intro video is great! https://youtu.be/Z-l7G8zq8I0
More importantly, though, they could use this NN on systems that have not yet successfully been modeled, perhaps complex dynamical systems, to discover good parameters and conserved quantities.
That would only make sense to try if the model could do this for systems we already understand. By the sound of the article, it can't even do that. Despite many efforts the researchers couldn't even understand the second pair of parameters. That doesn't correspond to my understanding of "good parameters".
I just read about Langrangian and Hamiltonian mechanics. I didn't encounter those at all in my EE physics, and they are fascinating. Great examples! Are you a physics professor, or is this stuff undergrad physics majors learn?
There's a good series of videos on YT, with the title Variational Calculus and the Euler-Lagrange equation on channel Structural Dynamics. I have only seen the first few. This first video should give you the full playlist:
> [software returns 4.7]
> Oh my GOD, it's discovered new physics!
I suppose anything nonlinear could invite multiple terms incidental to a particular local fit.
https://en.wikipedia.org/wiki/Conservation_of_energy#History
> code to reproduce our training and evaluation results is available at the Zenodo repository and GitHub (https://github.com/BoyuanChen/neural-state-variables).
It's amazing that the AI found two variables that corresponded to actual physical phenomena just from video footage but it seems highly unlikely that it found an alternative description of physics given how far off the answer was. The NN could be modeling any kind of weird "fantasy" for how the visual signal it was being fed could be produced that might not even correspond to 3D space in a meaningful way.
The "AI" came up with a slightly off number and the researchers struggle to understand why. They struggle because "the AI" is a black box that lacks explainability and cannot produce an explanatory model, only a predictive model [1]. The simplest conclusion is that their "AI" has not learned any useful models, and has only learned to accurately predict their test set, to which the authors had acces throughout the training of "the AI" [2]. The simplest explanation is that their model isn't very good at doing what they wanted it to do, but they opt to explain it as a scientific mystery, instead. Why?
It seems the only reason to try and see a scientific mystery where a simple failure would suffice as an explanation is because the model is "an AI". There seems to be some kind of expectation that "an AI" must have some deeper understanding of physics than humans, even when we don't understand what it's doing, even whe it's not doing very well.
In short, this seems to be based on very wrong assumptions and to be coming up with very wrong conclusions, but then again it makes for a great headline and so here we are, discussing it on HN.
____________________
[1] To clarify: a predictive model is one that can predict novel events. An explainable model is one that can explain how it made a prediction. An explanatory model is one that explains how the world works and why certain predictions are true, or not. Predictive and explainable models are useful, but most scientists aim to build explanatory models, because an explanatory model is necessarily also explainable and ultimately better at predictions, while a predictive model is not necessarily explainable or explanatory and an explainable model is not necessarily predictive or explanatory. Deep neural nets can only build predictive models, which are sometimes explainable, but they can't build explanatory models. That's because they can only identify correlations in data, but not explain those correlations with generalised theories, based on previous knowledge.
To put it plainly, neural nets can't come up with scientific theories, but they can estimate probabilities of things happening. But scientists can and want to come up with scientific theories, not just predictions.
[2] When the experimenter has access to the test data and can tune the learner's model until it scores highly on the test data, that leads to a model that overfits to the test data. Such a model is useless for prediction over unseen data (i.e. data not available to the researchers during training).
Sorry but that distinction doesn't mean anything falsifiable. Did you mean an explanatory model is one that's easier to understand for a human? If so, that's certainly useful, but it may well turn out that the simplest working model of physics is incomprehensible.
The theory of epicycles is predictive because it predicts the motions of the planets, as they are observed in the night sky. It is explainable because anyone can perform the necessary calculations and understand how the predicted motions are, well, predicted. The theory of epicycles is not an explanatory model because it does not explain why the planets should move in circular orbits with epicyclical sub-orbits. Kepler's laws are explanatory because they explain planetary motion as a result of Newton's law of universal gravitation, are predictive because they can be used to predict the motion of the planets and are explainable because anyone can plug in the numbers and see how the results are calculated.
So an "explanatory" model is a theory that explains why things happen they way they happen. A predictive model only predicts that some things will happen. An explainable model explains why it made a prediction, but it does not explain why this prediction should hold or with what frequency.
Another way to see this is that an explanatory model explains past observations and predicts future observations, while a predictive model only predicts future observations and has no explanatory power, cannot explain why past phenomena were observed.
An "explainable" model is just a model that people can understand. It's not more complicated than that.
So "explanatory" means that the theory references some other theory?
But just to be clear, I understand there's a colloquial meaning of "theory" as in what people mean when they say "that's just a theory". What I mean by "theory" is an epistemic object with either explanatory, or predictive power, or both. A theory can include multiple laws and hypotheses etc. And a theory "is just a theory" only until we can refute it, or find a better theory that does a better job at explaining things.
That's not a qualitative difference, it's a quantitative difference. Here you somewhat implicitly claim that models of lower Kolmogorov complexity for same, or better, accuracy are better, but how is that 'explanatory'? The term has no meaning. Ironically Kepler's laws are a better equivalent of epicycles, because they reduce to an imperfect approximation of Newtonian mechanics. Predicting actual orbits in a Newtonian Solar system of perfectly spherical objects with fully known matter distribution can only be done with an iterative solution.
The actual, physical reality continuously diverges from all current models, and all we can try to achieve is to reduce that error. It can't ever be eliminated unless the universe is in fact a simulation, and whatever is controlling it decides to copy humans into the parent world, then gives them external access to the full underlying state.
I just re-read your comment and I notcied that.
I think by "the simplest working model of physics" you mean quantum mechanics? I hope I clarified how I mean "explainable" vs. "explanatory" but yeah, I think that's absolutely spot on. Quantum mechanics is predictive, but not explanatory. I do think it's "explainable" though in the sense I sort-of define it, because it's a bunch of formulae that anyone can plug in the numbers to, and see how they come up with results. I don't reckon there's many theories in the sciences that are not explainable. The lack of explainability only becomes an issue with black box models like neural nets.
Perhaps I should have used the word "comprehensible" or "interpretable" instead of "explainable" since it causes confusion by being too close to "explanatory". But those words have their own problems ("comprehensible"... by whom?)
As an aside, I'm on the camp that hopes that quantum mechanics is somehow fundamentally wrong and we'll figure it out in the next big paradigm shift. It bothers me deeply that we seem to have hit a wall where we can predict, but don't have a clue and can't understand why. I think every other big leap in scientific ... understanding (hint) has come from the ability to make explanatory theories, that tell us how things work and why.
The human mind really wants to understand things in terms of the paltry slow coarse-grained 3D world we inhabit, despite 100 years of QM and QFT-based experiments showing that the world simply doesn't work that way. Can't other realms have other "why"'s? Tens of thousands of professional solid state physicists and QM researchers would probably not refer to their field as "not having a clue" for example, any more than anybody working in a field dominated by Newtonian mechanics would have a clue at least.
In fact, if you make a simulated world in a computer game, it's trivial to come up with algorithms that do alternative physics in the game that can create interesting behaviour, despite being far from explanatory for a being inside the sim.. I'm rather amazed the reality we inhabit is so easy to comprehend as it is even despite QM/QFT being pretty far from Newtonian.
I think what you're saying is that it may be impossible to fully understand the world in the same way that we have tried to, in all of science, until now. I agree, in fact I suspect we probably can't, because it's obvious to me our intelligence is limited (as a species) and we're going to hit our limit sooner or later, if we haven't indeed hit it already.
But that's not for me a reason to stop trying to understand as much as we can, neither is it a good reason to replace understanding with ... something else. Because the day when we finally hit our limit of understanding is the day our civilisation stops dead in its tracks.
I for one am not prepared to go gently into that good night, any time soon.
We can resort to math (like we have done in the last 100 years with QM) because while we rely on math being stringent we don't rely on it having to be intuitive on the level of our own experiences. On the other hand, even in a mathematical treatment of QM/QFT a lot of "classical" intuition and terminology is still used even though it, in my opinion, is hurting.
For example the insistence to apply properties like position and momentum to particles in the same physical framework even though QM/QFT has consistently shown for 100 years that conceptually you should pick one of them, the other is a dual, and failure to get rid of that baggage leads you into "weird" things like Heisenberg's uncertainty relation which is only mysterious when you insist on mapping properties to the classical world one by one. That is certainly one aspect where humans just can't seem to shake off the classical notions..
Current science is famous for "shut up and calculate" principle.
- Foo, bar, baz, ...
- But why?
- Shut up and calculate!
Your comment comes across as arrogant and ignorant. I have not formally studied quantum mechanics, but I do know theoretical physics is all about identifying and testing theories to better explain phenomenon. And the tests are statistically rigorous.
This is a thread specifically regarding an AI getting a wrong answer and practitioners suddenly deciding it's discovered new physics. Yet you claim CS is rigorous and quantum mechanics is not.
There is a tendency of someone well learned in one subject (eg CS) deciding that makes them an expert on any subject they casually observe. This is very off-putting.
I also absolutely did not try to claim that "CS is rigorous and quantum mechanics is not", as you took my comment to mean. The conversation here is not about rigour, or at least I didn't realise it was about rigour. Is the "shut up and calculate" dictum about rigour? In the context of the conversation and without looking it up I took it to mean that we should not try to understand and only try to calculate. I didn't put any more meaning on it.
I think you (I don't know about the other users who downvoted my comment) misunderstood both the meaning and the intent of my comment. May I suggest that clarifications should be sought, in the future?
Not that it will kill me to have my comments downvoted once in a while, it's just that this time I really didn't understand what I got so wrong.
EDIT:
>> There is a tendency of someone well learned in one subject (eg CS) deciding that makes them an expert on any subject they casually observe. This is very off-putting.
I'm with you on that, but I think if you look through my comments you'll see that I generally only discuss things I know. In particular with physics, I read but not comment on such conversations because I'm really not an expert.
I wasn't telling you what your comment meant to you personally -- only you know that. I was telling you how it came across and likely the reason it was downvoted significantly. I think a lot of people here don't have a lot of time to give feedback when they see an off-putting comment. And usually if someone has a counter-argument they'll say something instead of downvote. It's a good rule of thumb to consider how a tweet comes across that you may not have meant if it is just downvoted and not replied to.
This gave me a Hitchhikers Guide chuckle.
It does make sense though. Answers a much easier to arrive at than questions or processes.
indeed! i'm assuming it finds the same variables over multiple runs with different rng seeds. if that is true, it seems the next interesting question is if multiple kinematic problems can be used for the training of the same net. one key feature of physics is that it was invented to describe general principles that apply to multiple problems, so it doesn't surprise me that it might find some more optimal representation that doesn't necessarily generalize to other physical problems.
So that's why we put dangly things above cribs for babies to focus on.
\•••••/
\•••/
\•/
-Or maybe not. Maybe the "AI physics" has predictions for phenomena which our physics never bothered to figure out, whereas it would ignore large chunks of our physics area as irrelevant.
As far as I know, the article claims that the AI has discovered new physical variables, yet the researchers are unsure as to what they are. For all we know, these variables are the distance of objects from the edge of the image frame.
Is not most physics based on very extreme oversimplifications?
They are still very useful because errors average out, but for example: we don't know why fluid simulations work, in fact aerodynamics is an entire empirical field.
This was seen in F1 this year. All cars bouncing on straights, it was even given the name porpoising by the pilots. The teams never expected it, because the simulations never discovered anything about it.
Also: all the work of AlphaFold predicting protein structures.
Physics are awesome at predicting planetary motion, but stuff at smaller scales require a lot of empirical engineering to figure out.
And that made the fortune of CAD software designers.
Engineering needs engineering,. Physics does not. The interesting thing about aero in physics is the emergence of extremely complicated phenomena, not the equations as per se.
As-is, it's pseudoscience. What happens when you do a factor analysis on images? You get some measure of the axes of geometrical variance across those images.
Are those axes "related" to any physical variables, sure -- but almost never directly. To suppose the system itself had these properties is to suppose, for example, constellations actually exist and cause your personality traits.
Everything we want to know is what phyical properties of the system give rise to the observed consistent correlations in geometrical properties. *THAT* is physics.
Showing these geometrical properties exist and are consistent is just what we're trying to explain.
You cannot go from images to the domain of physics -- there are an infinite number of theories consistent with these images domains. And this is pseudoscience.
There's no such thing as "physical properties of the system" other than measurable quantities that can be used to make predictions, which is what this does. There's no reason to be sure that temperature, for example, is a "real" physical property of a system rather than just one of many variables that would help us model it and understand it.
Do you think it's pseudoscientific because there's no theory-ladenness in the predictions?
From the article, it doesn't work. They found on known physics it gives 4.7 dimensions, of which only two are explicable -- 4 is correct; the others have no known physical interpretation. No surprises: those two are just the geometric properties of the system (angles) which are actually properties of the image. The others are pure bullshit.
Since, of course, the real physical parameters of the system we take to have generated those images are not present in them. The images are distal effects of these things
Only in cases where the geometric properties of the target system are causally relevant to its actual causal properties will this work -- ie., only when "angles matter"
Thinking you can infer laws of nature from images is pseudoscience, and these guys need to think more carefully about why we experiment in the first place
eg., Consider that if mass is a relevant causal property, there'd be no way of inferring it from images: two objects can be visually identical whilst having radically different masses... making images *OBVIOUSLY* not a measure of mass...
this project almost defines the modern kind of schizophrenic pseudoscience born of this wave of AI
So for example, take an extreme case, suppose that somebody says he wants to eliminate the physics department and do it the right way. The “right” way is to take endless numbers of videotapes of what’s happening outside the video, and feed them into the biggest and fastest computer, gigabytes of data, and do complex statistical analysis — you know, Bayesian this and that [Editor’s note: A modern approach to analysis of data which makes heavy use of probability theory.] — and you’ll get some kind of prediction about what’s gonna happen outside the window next. In fact, you get a much better prediction than the physics department will ever give. Well, if success is defined as getting a fair approximation to a mass of chaotic unanalyzed data, then it’s way better to do it this way than to do it the way the physicists do, you know, no thought experiments about frictionless planes and so on and so forth. But you won’t get the kind of understanding that the sciences have always been aimed at — what you’ll get at is an approximation to what’s happening.
Unless you have videos of experiments designed to observe measurement devices we have created, on systems we have designed, it's all useless.
The only useful thing in figuring out how nature works is creating truely novel experimental circumstances and measuring them with novel devices created for that purpose.
You cannot do science as a statitics of images; that's pseudoscience. And chomsky is here only half-right; it's actually much worse than he's sayign.
And maybe the 4.7 is actually more correct? The 4-parameter model is an approximation that neglects friction and air resistance. Moreover the double pendulum is a chaotic system and chaotic systems sometimes have dynamics described by laws with non-integer exponents such as Lyapunov dimension. I'm just spitballing, but the point is that it's not a priori ridiculous.
It's definitely possible to estimate mass from images. How do you think we know the masses of asteroids and planets? No-one put them on a scale, we just record their motion and work out which value fits best.
A concept that says gravity is the result of bending spacetime, with the speed of light being constant. It's not just a model, it's saying the universe is 4D spacetime, which explains why GR is so predictive.
I think a counter argument would be that if there is SOME signal in the photos AND there's enough training data that does have the correct ground truth signal that the scientists are matching up then you can have SOME level of accuracy. If the training set can reasonably cover the space of possibility that we're interested in then we can get reasonable interpretations.
However in this case the insane number of physical phenomena will always be larger than any training set so this approach should NEVER generalize there will always be way too much noise which is what the scientists have figured out here. So I agree with you that it's extremely limited but I don't think I'd call it pseudoscience there might be very limited domains where for example the only data we have available are images and so such a tool may be appropriate.
I definitely share your frustration though since any half way decent scientist should have just done a thought experiment instead and figured that this wouldn't work well. This smells like BS academic marketing where they always inflate their own impact and significance.
Perhaps they could leverage their lifelong training set which correlates scenes that look like they have bowling balls with scenarios that have a high mass movable sphere.
Perhaps we could have a good laugh together by painting a bowling ball to look like styrofoam and painting styrofoam to look like a bowling ball- then we could watch the silly ai/human apply an incorrect mental model and fail to grasp the causal reality! Ohohoho
Granted, the reason why people did Astronomy was because they believed in Astrology, but it's no longer been the case since a while.
I didn't see any attempt to formulate laws. The researchers trained a neural net model to predict the next event in a sequence. That is not a natural law, it's a maximum probability estimator.
To clarify, a natural law would be a formula with variables that one can plug in numbers to, in order to predict the behaviour of a system. For example, Newton's law of gravitation is a natural law, Kepler's laws of planetary motion are natural laws, the laws of thermondynamics are natural laws. But a neural net model trained to predict the next frame in a video? How is that a "law"?
As a for instance, if I train a neural net to predict the motions of the planets, the trained model is a law of planetary motion, like Kepler's laws of planetary motion? Is that correct?
Then the first half of the network (before the low-dimensional layer) will learn how to "encode" the state of the system in the video in as few variables as possible, such as the orientations and angular momenta of the double pendulum. This is equivalent to what humans do when we look at a messy physical system like the Solar System and model it with a few quantitative parameters.
The bottleneck layer will represent the handful of state variables, and then finally the other half of the network will learn the mathematical function that predicts the system's evolution. This is equivalent to what humans do when we work out physical laws and equations of motion.
I can agree that a neural net can learn a model that can predict the behaviour of a system, to some extent, within some margin of error.
That's not enough for me to see neural net models as (scientific) "laws". For the sake of having a common definition of what a scientific law is, I'm going with what wikipedia describes as a scientific law: a statement that describes or predicts some set of natural phenomena, according to some observations (paraphrasing from: https://en.wikipedia.org/wiki/Scientific_law). Sorry for not introducing this definition earlier on. If you disagree with it, then that's my bad for not estabilishing common terminology beforhand.
In that sense, neural net models are not scientific laws because, while they can predict (but not describe) they are not "statements". Rather they are systems. They have behaviour and their behaviour may match that of some target system, like the weather say. But like a simulation of the economy, or an armillary sphere are not, themselves "laws", even though they are possibly based on "laws", a neural net's model can't be said to be a "law", even if it's based on observations and even if it has an internal structure that makes its behaviour consistent with some (known or unknown) law.
There is also the matter of usability: neural net models are, as we know, "black boxes" that can't be inspected or queried, except by asking them to analyse some data. While useful, that's not a "law", because it does not help us understand the systems they model. If this sounds like a semantic quibble, it isn't. To me anyway it doesn't make sense to base scientific knowledge on a bunch of inscrutable black boxes. Scientific laws and scientific theories are not black boxes.
As an aside, neural nets fall short of what Donald Michie (father of AI in the UK) called "ultra-strong machine learning" [1]. That's the property fo a machine learning system that improves not only its own performance, but that of its user, also. Current techniques aren't even close to that.
____________________
[1] Machine Learning: the next five years, Donald Michie, 1988
But I would argue that this parsimony is illusory. There's a lot of implicit knowledge needed for the interpretation of physical laws. The laws are written using specialized mathematical notation such as special functions, partial differential equations, in a certain conceptual framework such as Lagrangian mechanics. You need to understand the concept of abstracting and quantifying a dynamic system (most people wouldn't imagine you can do this) and then you have to learn all the tips and tricks of how to reformulate and solve systems.
For example, I could write a mathematical representation of quantum electrodynamics (the theory of how electrons and photons interact) on a single index card. However, I would need to dig into my two shelves of QFT textbooks to actually make any quantitative experimental predictions, on top of my degree, PhD and post doc experience, which I need to even be able to read the textbooks (and I would still mess up the minus signs).
I think it's important to remember that these neural networks are doing all of that - not just finding the physics, but also all the abstraction, calculation and interpretation that is usually taken for granted but actually very non-trivial.
The tools of physics have a lot of implicit assumptions that guide the end result in ways that I would describe as parsimonious in terms of how much the output state space must be reduced. They are much more free, which is why they can be amazing for some very hard shit, but proving they're behaving exactly in "physical" way is very hard.
"Time is defined so that motion looks simple" is my favourite quote from MTQ for this reason. It's intuitive and yet also very physically "rigorous" in a way that people don't necessarily realize is a thing in physics beyond just using mathematics.
Maybe we can just train the AI to do the maths for us, dunno, but I think currently this tabula Rasa approach will inform the physics-y-ness. I still call it physics personally, but I don't really think it's interesting from a purely physical perspective.
There have been some works deriving conservation's laws and so on from empirical motion, which I think is very impressive at scale, but I don't know what that does for physics as opposed to the applications of said physics.