How to solve computational science problems with AI: PINNs
mertkavi.com
mertkavi.com
See for example https://www.nature.com/articles/s42256-024-00897-5
Classical solvers are very very good at solving PDEs. In contrast PINNs solve PDEs by... training a neural network. Not once, that can be used again later. But every single time you solve a new PDE!
You can vary this idea to try to fix it, but it's still really hard to make it better than any classical method.
As such the main use cases for PINNs -- they do have them! -- is to solve awkward stuff like high-dimensional PDEs or nonlocal operators or something. Here it's not that the PINNs got any better, it's just that all the classical solvers fall off a cliff.
---
Importantly -- none of the above applies to stuff like neural differential equations or neural closure models. These are genuinely really cool and have wide-ranging applications.! The difference is that PINNs are numerical solvers, whilst NDEs/NCMs are techniques for modelling data.
/rant ;)
Edit: this is the Yao lai paper I’m talking about:
https://www.sciencedirect.com/science/article/pii/S002199912...
PINNs have serious problems with the way the "PDE-component" of the loss function needs to be posed, and outside of throwing tons of, often Chinese, PhD students, and postdocs at it, they usually don't work for actual problems. Mostly owed to the instabilities of higher order automatic derivatives, at which point PINN-people begin to go through a cascade of alternative approaches to obtain these higher-order derivatives. But these are all just hacks.
The best part about PINNs is that since there are so many parameters to tune, you can get several papers out of the same problem. Then these researchers get more publications, hence better job prospects, and go on to promote PINNs even more. Eventually they’ll move on, but not before having sucked the air out of more promising research directions.
—a jaded academic
http://www.incompleteideas.net/IncIdeas/BitterLesson.html
I would love to hear HN's take on this argument.
It used to be almost everyone worked in agriculture, now about 1% does, so the others are free to do something else. Prior to the microprocessor making a computer required manual assembly of thousands of parts, early microprocessors contained thousands of parts manufactured by a small number of photographic and chemical steps, and the number of parts has grown into the billions without the number of steps expanding millions of times.
For medical imaging your solver is most likely "physics based" in any case. PINNs want to encode the physics inside a neural network, instead of developing an appropriate algorithm (which obviously has to also incorporate the physics) which solves the problem and which can be analyzed mathematically.
On the contrary, neural networks will never give you a better way to convert to the frequency domain than the Fourier Transform. At best, they might approximate it.
No. That is the use case for numerical analysis, where you can develop a high performance, accurate algorithm based on the mathematical analysis of the problem.
It is really kind of silly to want to encode the solution in some neural network.
Maybe there are some very special problems, without well developed silver where they work.
See, the reason for the Bitter Lesson being a lesson is that Neural Nets, that Sutton is mainly writing about, are pretty crap at representing background knowledge. The only way you can store expert knowledge in a neural net is to modify its structure and its weights. The weights you can only modify by some kind of learning procedure like backprop, in practice. Very limited forms of background knowledge, like convolutions can be encoded in a neural net's structure, but imagine trying to represent, I don't know, the last ten lines of code you wrote today as a bunch of neural net connections. Continuous functions is just not the right kind of notation for that sort of thing.
If neural nets were any better at encoding background knowledge, they would use it, but they can't so they have to rely on data. And that's why they need so much of it. Background knowledge functions as a strong inductive bias- it directs the search for a hypothesis to hypotheses that we know make sense (again, think of convolutions). Without background knowledge, or with only a little background knowledge, you need tons of examples to learn anything useful.
So the Bitter Lesson is basically making a virtue out of necessity. In any case, it's not a prescriptive thing, only descriptive.
Now, Sutton argues that some of those systems at least did not rely on domain knowledge. They all did: Monte Carlo Tree Search, used for board game-playing AI agents, is nothing else but an encoding of domain knowledge - specifically, domain knowledge about the structure of two-player, complete information games. It's the same domain knowledge that was used to create the minimax-based DeepBlue software that won against Kasparov.
Neural nets have still not managed to win against human players in board games without incorporating such a strong, knowledge-dependent component, as a game-tree search.
So sutton is fudging the details. Not on purpose. He's an RL person. In RL, as in planning and other disciplines, folks tend to forget all the knowledge they put into their systems in the form of inductive biases, or auxiliary (but can't-do-without) algorithms like MCTS. He's like the proverbial fish that don't know what water is, because they swim in it.
I think someone who cared about the specific content would at least note that linear PDEs like the heat equation often have closed-form solutions and/or efficient algorithms for solving any particular problem, so aren't likely to be usefully solved with PINNs.
As a researcher currently working on a project involving PINNs for cardiac biomedical applications, in my field industry players generally steer clear of PINNs. The core issue lies in the approach of approximating physical laws through a loss function.
Basically: instead of leveraging mathematically robust solvers, PINNs attempt to encode the underlying physics (e.g., conservation laws) as constraints within the loss function of a NN. Clearly this is a problem for systems requiring high fidelity in physical accuracy...
It is very telling that the motivating example in https://physicsbaseddeeplearning.org/intro-teaser.html totally fails. If you read further you will notice that all other examples also fail and I mean "fail" as in the result is literally useless. (Reading this was genuinely upsetting)
The one legitimate area where I can see them being used is for interpolation in complex engineering/physics applications. Surrogate models actually can be a worthwhile endeavor if you need to solve a complex problem over a big parameter space. But of course this is fundamentally different than training a neutral network to solve a single PDE.
Even the idea is fundamentally flawed, ODE solvers are well developed and are based on sound mathematical theory. The idea that you can outperform them, by a neural network is on its face pretty silly.