Per stephen wolfram [1]
> It’s worth emphasizing that there’s no “theory” being used here; it’s just a matter of what’s been found to work in practice.
https://writings.stephenwolfram.com/2023/02/what-is-chatgpt-...
Per stephen wolfram [1]
> It’s worth emphasizing that there’s no “theory” being used here; it’s just a matter of what’s been found to work in practice.
https://writings.stephenwolfram.com/2023/02/what-is-chatgpt-...
The formula of a neutral network is not hard to see. People would not be able to write code to do training or inference if they didn't know what is happening in the network.
>the number of parameters, the number of layer
You can measure the effectiveness of a neutral network to know how tweaking the number of parameters affects it. There are also restrictions like hardware or execution time that can place an upper bound on these aside from overfitting.
>the choice of temperature
Similarly, you can measure the quality of output to find the temperature that seems to work the best.
The theory is that it's all one big optimization problem. Where the values of the parameters of the NN aren't the only thing you can tweak.
We can throw darts to the board and measure how close we are to the bullseye. Then we can tune parameters like throwing force, elbow angel, etc. These parameters have physical meaning. We can reason that by throwing harder, the trajectory of the dart can be so and so affected, hence causing the dart to fall closer to or further from the bullseye. Regarding "People would not be able to write code to do training or inference if they didn't know what is happening in the network", I understand from an operation perspective what the neurons are doing, but I also don't get what they are doing, if that makes sense.
For example in each layer of a CNN, is it doing Fourier transform, low pass filtering, rotating, scaling or whatever? And even if they are, why does it work? Why does it give the result that we want? Can we reason that varying our throwing force will change the trajectory of the dart? Shouldn't we do that? Or we are happy with getting results of hot dog/not hot dog [1]. There lies my uneasiness with ML.
I agree with cycrutchfield's comment that practice outrun theory though. We will learn more as we progress.
This is how Newtonian physics is derived. Similar to neural networks you create an approximation based off your observations. The actual reason why those approximations work is not necessary to be known to utilize Newtonian physics.
>in each layer of a CNN, is it doing Fourier transform, low pass filtering, rotating, scaling or whatever? And even if they are, why does it work? Why does it give the result that we want?
What the layers themselves are actually doing doesn't really matter and each one is just a part of a larger system. What is important is finding a set of weights that approximates a function that minimizes loss. People care about the ability to find this approximations. What exactly it is found is less of a concern.
>Or we are happy with getting results of hot dog/not hot dog
Neural networks are not limited to binary classification.
1) They’re arbitrary function approximators.
2) We have a great way of fitting them to data (backpropagation).
As an aside, it’s just an accident of nature (or is it?) that neural networks have passing similarity to biological systems.
Anyways, it doesn’t make much sense to try to reason about what a neural net is doing inside. What it’s doing is trying to approximate the function it was trained to. That’s it.
It has nothing to do with a poor misunderstanding of the model. Temperature is not that hard of a concept. A strong understanding leaves you understanding that such a parameter will be kind of fluffy. That's the nature of creative processes.
CSI is low temp. Twin peaks is high temp.
The theory usually catches up, belatedly.
Sometimes there just may not be a simpler theory of why something works. Or the 'simple' theory is still massively complex. Reality doesn't owe us an easy explanation.
Neurosciences are helping understand the brain better, but we don't much yet.
AI is developing more capability, while we are losing understanding of how it is working.
Convergence would happen at some point, where for a given feature/capability, we may have an equalish understanding of the brain and AI. Subsequently understanding of both may increase together.
This may rinse and repeat for more desirable features of the brain.
We may find that the brain does not have too much more to put into AI. What if the hard problem of consciousness is not actually is.
Even so, it’s escaped the lab already. Which model will make the better products?