No physics? No problem. AI weather forecasting is already making strides
arstechnica.com
arstechnica.com
You have a model that predicts with great accuracy that rowdy teens will TP your house this Friday night, so you sit up late waiting to scare them off.
You have a model with less predictive power, but more discernible parameters. It tells you the parameter for whether or not houses have their front lights turned on has a high impact on likelihood of TP. You turn your front lights on and go to bed early.
Sometimes we want models that produce highly accurate predictions, sometimes we want models that provide mechanistic insights that allow for other types of action. They're different simplifications/abstractions of reality that have their time and place, and can lead you astray in their own ways.
Sadly, insight is always lost. In a noisy world where even with the best regularization, some fitting on it, or higher order features that describe it, is inevitable for maximizing prediction accuracy, especially if you don't have the right tools to model it (like transformers adapting to lacking registers [1]) and yet a lot of parameters within chosen architecture.
What's worse, bad expectations are often much worse than none. If your loan had been denied by a fully opaque black box, you may be offered recourse to get an actual human on the case. If they've trained an interpretable student [2], either by intentional manipulation or by pure luck, it may have obscured the effect of some meta-feature likely corresponding to something like race, thus whitewashing the stochastically racist black box. [3]
[0] "Interpretability in ML: A Broad Overview" https://www.lesswrong.com/posts/57fTWCpsAyjeAimTp/interpreta... [1] "Thread: Circuits" https://distill.pub/2020/circuits/ [2] "Why Should I Trust You?": Explaining the Predictions of Any Classifier" https://arxiv.org/abs/1602.04938 [3] "Fairwashing: the risk of rationalization" https://proceedings.mlr.press/v97/aivodji19a
I think having multiple layers of abstraction can be really useful and have done it myself for some agent-based models with high levels of complexity. In some sense, these approaches can also be thought of as "in-silica experiments".
You have a model that is complex and relatively inscrutable, just like the real world, but unlike the real world, you can run lots of "experiments" quite cheaply!
Interestingly though, physical models are usually expressed as mathematical equations. Which is an arbitrary way of modeling
A Neural Network could technically “discover” different models, just through optimizing predictions for whatever we want the model to do
We might not be able to distill the NN into nice compact equations, but they might still form a pretty good model of whatever phenomenon is fed to it through observational data
Note that just by picking the input data and the output expectations, we are already defining a model
What is “the underlying process”?
For example, Newton was able to model gravity quite successfully without ever being able to “understand the underlying process”. In fact, physics today still doesn’t have a good grasp on what gravity is. Yet we use the models and equations all the time
In a way, physics is also a collection of black boxes, perhaps just seemingly more elegant boxes
IMO -
The simple ones: advection, latent heat release/absorption from water changing phases, and the Coriolis force. If you need an AI for this, please take a course on differential equations.
The hard ones: droplet/ice crystal formation, cloud feedback on radiative transfer, evaporation at air-sea boundaries. If you can train a model for these processes, please, please tell someone.
Newton realized that these phenomena could be explained as arising from the same underlying process of an inverse square law. This is a much more useful model, and allows predictions that allow us to do things like space flight, even if it is not complete.
It's not useful to draw a false equivalence between AI-style "the model predicts, that's good enough" and science as a whole which cares very much about the underlying structure.
If anything, AI researchers are digging deeper into the models too
And people in physics are starting to use AI tools to model physical phenomena
I think that it’s a never ending task to understand all the black boxes. Definitely not possible by a single person. But also at some level you get to circular references. There is no fixed point in the universe, there is no point 0 or origin that we can find. Everything is relative to something else
What if the hammer gets angry and starts hitting us in the head? How can you not see the danger in getting hit in the head by a hammer?
This is all just short term noise and no one will ask these stupid questions soon enough. In the meantime, it makes for good theater on podcasts.
The training data could be real physics in a simulator held up against evolutionary driven AI logic that competes against it with various goals that are then evaluated and if given a high score then marked as isomorphic and given enough runs you'd get a dataset.
Having models doesn’t imply what you probably mean by “knowing about physics”, it just means having a representation of something that is not that something (like a map of the world is a map, not the world)
So, depending on how the dog’s understanding of the world works, it maybe doesn’t need any models, I can’t really know
But I do know, that if I want to describe anything, using any symbols whatsoever, then I’m implicitly creating a model of what I’m trying to describe with the symbols
So, if I’m trying to communicate and understand physical phenomena using AI, then I’m implicitly creating a physics model, whether I want to call it that or not
No. That's precisely the issue here. NN's do not "model the world" in anything remotely like the traditional meaning of that term.
Modelling the world historically means identifying objects/phenomena and proposed causal relationships between them. Without the relationships - e.g. if we add heat to the system, it will move more - there's no model. You may still be able to get predictions, and they may still be useful, but you are not defining a model.
Now our lead experimenter asks this person "what will happen if the global average temperature increases by N degreesC?" and they get an answer.
Can we way that the lead experimenter has built a model? They have not, certainly not in the sense that they have any access to it. The person who replaced the NN may have (and indeed, probably has) built some sort of model, but that's a very different claim.
Explainability in NN/ML systems is a hot topic, and many people (not all!) would say that if the NN/ML system cannot explain why adjusting parameter X will cause changes in parameters A, M and T, then you have no access to anything that merits being called a model.
A consequence of this is that if the person who replaced the NN can explain themselves (e.g. answer the X -> A,M,T coupling), then even the experimenter can probably be said to "have a model". But if all that can be said is "I don't know and/or I can't explain, you just need to trust me that this coupling is real", then the claim that a model has been built is on unstable ground.
What symbols or language did the lead experimenter use to ask the person the question? And what does degrees, temperature and global mean?
All of those things require models to be communicated between the components of your system
Any symbolic communication is necessarily a model of what it is trying to represent
Of course, if there isn’t someone to interpret it, it’s just symbols. But to interpret a meaning behind symbols, then it implies the symbols represent a model of the meanings that are being communicated
The insight gained by rigorously modeling a system in computer code produces a person (the modeler) who can provide valuable insight when asked questions about the system. In policy analysis, the modeler’s insight can often provide quick and dirty and auditable (and often correct) analyses/answers about the modeled system without ever running the developed formal computer model. The exercise of the formal development of a computer model credentials the modeler as having gained a level of rigorous systems-level expertise. And the scope and detail of that modeler knowledge is certified in the depth and breadth of the computer model itself (and the currency and accuracy of the input data sets).
Nice to have such an human analyst around when important policy decisions need to be made, since such policy decisions should be made and implemented by humans who can explain the confidence that exists regarding the knowledge that supports the given decision. The decision makers can then point to the analysts for the estimate of the degree of confidence that can be ascribed to the policy analysis that supports the decision. That’s how it’s supposed to work, and that philosophy is formalized in existing decision processes for complex technical systems such as transportation, telecommunications, power, military systems, etc.. You know, the important stuff….
To train a NN you need a training set of data, that data follows a certain order or pattern that represents the model that the person has
As long as the NN behaves as expected by the person setting it up, then it is useful as a model
How useful? That depends on what you are trying to do, your model and the data
If you want to take it to its extreme, all language is a model. Even what I’m expressing right now with this text. It’s just a model of what I’m thinking and the meanings I’m trying to convey
I have been repeatedly assured by people who manage to sound authoritative about NN's and ML here on HN that, despite my instincts to the contrary, this claim is no longer true. I continue to doubt it, but there you have it.
Your claims about language are quite interesting, and quite controversial.
What is the goal of that NN? How are inputs and output of that NN defined?
If you define any structure or properties of the NN, or any description at all, you are implicitly creating a model. Any NN is already a model
Not sure why that would be controversial
Not exactly. The particular symbols used in mathematics may be arbitrary but the mathematical structures and relationships are not.
There’s some interesting research on how the limits of predictability may come down to what is mathematically provable in a formal system. That may sound kind of weird, but consider that any computable model can be modeled as a Turing machine, which is essentially a process that manipulates symbols according to a set of rules—not too different than mathematics itself. The difference is that based on certain assumptions of internal consistency (that cannot be proven if the system is in fact consistent), the mathematical manipulation of symbols can be used to make predictions about the behavior of Turing machines. There’s a very deep connection between the two and neural networks are still just another form of this, perhaps not as human-interpretable however.
[0]: https://www.microsoft.com/en-us/research/blog/introducing-au...
> The first step is potentially changing the way data is assimilated into AI-based models. At present, they almost universally use a set of initial conditions produced by a physics model. That is, a model like the ECMWF spends an enormous amount of computing power to collect data from buoys, surface stations, weather balloons, airplanes, ships, satellites, and many other sources and then synthesizes a set of initial conditions for grid points across the planet. All models then take this as the beginning "state" of the planet's weather and forecast from that.
So this is essentially learning the time-stepping part of the physical model, not deriving predictions from raw data. While still interesting and probably still complex, this is far less impressive than the title lead me to believe.
For context, intensity changes are the current Big Problem. Otis [1] had its track predicted almost exactly, but its explosive intensification from a tropical storm to a Cat 5 was totally unpredicted. Possibly some of the ~$12bn damage could have been avoided if Mexico had known that in advance.
I've said this before in another comment some months back, but I'll repeat: my worry is that these models aren't learning some comprehensive new climate dynamics model with parameterisations [2] , but only fitting what the Earth has historically done. And if AI weather prediction is only learning what climate dynamics do 95% of the time, it's almost by definition not useful for predicting extreme weather and it will get less accurate the more the climate changes. You're just going to get more Otises.
[1] https://en.wikipedia.org/wiki/Hurricane_Otis
[2] much as I would welcome, with open arms, some accurate AI-generated black-box parameterisations for e.g. subgrid precipitation - might be more explainable than the FORTRAN black-box parameterisations we have now :)*
My naive understanding is that the majority of temperature data comes from where humans are: the surface. Hurricanes are 3d, extending up for miles. The models go almost entirely off the surface temperatures, with very very sparse balloon data (which is a poor sample, since a balloon will follow the air it's put in). Wouldn't the whole volume, or at least a little of it, need to be observed, since the energy in that volume is what's powering the hurricane, not the energy on the surface? I would assume this is why the models have trouble.
Hurricanes are powered by air, not clouds. Often, there are only specific heights of clouds in their path.
> but they give good data in their climb/descent.
That seems incredibly sparse, with a very small coverage in very specific places.
> Eh, atmospheric CO2 (and other GHGs) is just another input parameter
only works for CO2 inputs outside our measurements, if the climate response to those inputs is linear (and thus predictable from the responses we have already seen)?
I claim that the climate response to CO2 forcing is, in fact, strongly nonlinear, and further that it's nonlinear for other "unusual inputs" - not just CO2 - things like sea surface temperature or unusually low pressure troughs. So-called extreme weather. I can't bring up good citations at the moment, sorry, but here's a somewhat exaggerated thought experiment:
Take an AI model trained on weather measurements ~1970-2024. Also take a model of the primitive fluid equations on a rotating sphere. What predictions might you expect from each one for an asteroid hitting an empty patch of the West Pacific?
A sensor that’s off by some constant factor is feeding bad data to a physics model and resulting are deemed incorrect if they don’t match the future value of that sensor. AI on the other hand could self correct for such issues because the data doesn’t mean anything only the patterns.
I can only assume current models include everything even tangentially relevant from albedo to topology and ocean currents. But that doesn’t mean they include everything relevant just everything people consider relevant.
https://www.faa.gov/air_traffic/publications/atpubs/notam_ht...
The standard model is considered to have too many free variables and those are few enough that you can memorize and describe them all.