‘Machine Scientists’ Distill the Laws of Physics from Raw Data
quantamagazine.org
quantamagazine.org
[0] https://www.researchgate.net/publication/286905402_Symbolic_...
[1] https://towardsdatascience.com/symbolic-regression-the-forgo...
The Julia automated discovery of physical equations work showcases these symbolic regression techniques quite heavily, for example in the State of SciML talk (https://www.youtube.com/watch?v=eSeY4K4bITI) we discuss how this has been used now across hundreds of different scientific use cases. The universal differential equations paper which describes the Julia SciML organization (https://arxiv.org/abs/2001.04385) demonstrates these cases where neural networks in differential equations are mixed with symbolic regression to allow for mixing prior physical knowledge with data for discovering just the unknown higher order physical equations. And there's a lot more directions we're going next.
And quick mention to bring it back to the main thread here, the DataDrivenDiffEq symbolic regression API gives back Symbolics.jl/ModelingToolkit.jl objects, meaning that the learned equations can be put directly into the simulation tools or composed with other physical models. We're really trying to marry this process modeling and engineering world with these "newer" AI tools.
There is a segment in Adam Curtis's "All Watched Over By Machines of Loving Grace" which describes scientists trying to model the ecology of a prairie. As described by Curtis, the more data they collected and more complex their model, the worse the predictions of the model became. It feels like the lessons of the past couple generations (many decades) have been that symbolic models don't compose together easily; symbolic analysis only works in simple systems and has hit diminishing returns; numeric and ML methods work well enough; etc. I'm curious if better tooling (augmentation and/or automation) can push through some of these challenges and yield real understanding. Or if there are some fundamental scaling problems that can not be overcome.
I don't think I've seen a Python wrapper on Julia code before.
Yes, that and that the tree data structure used by genetic programming tends to have functions in every non-leaf node, and static values or variables in the leaf nodes.
You don't have to use genetic/evolutionary algorithms to search the space of programs, it's just the most popular method.
You can even try pure random search if you're feeling particularly lucky:
It sounds like this symbolic regression approach can potentially find correct PDE solutions directly from data:
"In February, they fed their system 30 years’ worth of real positions of the solar system’s planets and moons in the sky. The algorithm skipped Kepler’s laws altogether, directly inferring Newton’s law of gravitation and the masses of the planets and moons to boot."
Note that Newton’s law of gravitation can also be derived from a differential equation. Also note that simply knowing the exact solution to the PDE does not mean that the PDE model itself is an exact representation of reality. For example, we know from Einstein's laws that Newton's laws are not correct at astronomical scales. In the symbolic regression approach, the solution would be limited by the number of different observable variables available to train the model on.
Many PDEs in real world problems also do not have any known exact symbolic solution, so approximate symbolic formulas are used routinely anyway but are just painstakingly discovered and derived by humans with "deeper understanding" (actually, just hoards of PhD students and their advisors throwing everything they can think of at the problem until one of them lands on a unusable formula).
What do you mean here? Aren't all approaches limited by the number of observables available (except speculatively)?
“The algorithm picked up on additional terms,” Zanna said, producing a “beautiful” equation that “really represents some of the key properties of ocean currents, which are stretching, shearing and [rotating].”
We don’t know if they break down, because we have no means of measuring them more accurately on earth with all the statistical fluctuations happening.
In practical terms, some of them are in the sense that it is impossible to measure their inaccuracy with common equipment and our primitive human senses. In the end, their purpose is to make predictions and help us make sense of what’s around us, they do not need to be exact.
It's trivial to find formulas that fit data to within the variable's uncertainties. You can easily drown in such output. There are two main challenges with this approach: 1. make sure you have the right "ingredients" (factors) for the search, and 2. filter the output so that human reviewers don't waste time on obvious nonsense equations.
Filtering output is not terribly difficult. Some assessment of complexity is required and depending on the formula a symmetry score might be useful. Some actual physics equations are quite complex so you have to be expecting a simple equation or filtering output will be less effective.
Making sure you have all of the right potential factors available is very hard. Expanding the input factor set extends the search space and time.
Wikipedia page:
https://en.wikipedia.org/wiki/Robot_Scientist
Nature publication:
https://www.nature.com/articles/nature02236
I guess US media don't realise there's science and technology outside of the Americas eh?
It’s very cool witnessing the consequences of Moore’s law; as the cost of billions of calculations goes to pennies, this becomes easier and easier.
Since the complexity of nature is unbounded, I wonder if the slow down of Moore’s law will represent a plateau of what can be symbolically discovered.
For example, will the next set of scientistic discoveries require quintillions of combinatorial checks that can’t be accomplished in a human lifetime?
Very interesting article. Thanks for posting.
I've been thinking about the physics equation and its relation to proportionality. Not all physics equations are proportionalities because in physics the equality sign is loaded.
To me, proportionality is fundamental not the physics equation. So I would have written the last sentence of the quote as "...it stopped when it found [a proportionality].
--Dr. Henry Walton Jones, Jr., 1936Wikipedia contains equations from many different branches of math.
Aren't mathematical equations written in all sorts of notations (really obscure notations sometimes, for obscure branches of math), depending on the branch of math they're used in?
How would this program even be able to use equations written in a notation it doesn't understand? Even if it somehow understood the notation, that doesn't mean the algorithm understood the branch of math the equation was for.
This all sounds unworkable to me, unless they're somehow limiting the equations they use to some branch of math they already understand.
Are you saying that different branches of math don't have their own special notations?
Sure, they have notation in common, but they have their own notation too, to express objects, relations, or operations that are of special relevance to them.
That's not to mention that any given symbol could mean different things depending on the branch of math it's used in.
"..the algorithm evaluated candidate equations in terms of how well they could compress a data set. A random smattering of points, for example, can’t be compressed at all; you need to know the position of every dot. But if 1,000 dots fall along a straight line, they can be compressed into just two numbers (the line’s slope and height). The degree of compression, the couple found, gave a unique and unassailable way to compare candidate equations."