An idea from physics helps AI see in higher dimensions
quantamagazine.org
quantamagazine.org
1) The article doesn't say this, but dimensions don't always have to do with locations in space and time, you can treat any value that can continuously vary as a dimension -- for example, a person might have dimensions for personality type, age, hair color, etc, etc.. Seems like they could use this technique to better train CNNs to recognize patterns in a lot of data besides imagery -- fraud detection based on credit card transactions, for example. 2) There are a lot of local and global symmetries in physics -- i wonder what new capabilities adding them to a CNN would enable?
Without repeating the article too much, this is important because it can be used to learn very complex systems from a series of lower dimensional projections. Such as, creating a 3d map of a dog from a collection of 2d images of dogs. The resulting system can better detect a dog in a position it's never seen because the CNN has built into it the relationship between 3d space and the 2d representation of that space.
Contrast that example with any sufficiently large, uniform sampling of the n-dimensional unit cube. The data is inherently high dimensional, and any attempt (attempts from a certain class of allowed methods -- no need to burden ourselves with the details) to reduce from n to k<n dimensions will throw away some important information about the structure of the data.
Interestingly, it's very possible to have a low dimensional representation of a high dimensional process. One of the other comments mentioned the example of a photograph representing a 3D scene (also containing a week, implicit view into some other variables like temperature). That transformation from 3D->2D is inherently lossy, but for some kinds of problems a finite sampling of a low dimensional representation allows you to uniquely reconstruct a high dimensional data representation. I haven't read the article yet, but the other comments seem to indicate something of that flavor happening here.
The differences comes in the fact that higher dimensional tuples may contain data that is independent of other fields. Say, if you have a tuple that's: (name, dob, address,) and you have a projection function that accepts such a 3-tuple and returns a 2-tuple of (name, dob,). For that function, the address dimension has no relationship at all to the other fields, meaning, that there is not unproject function that a person could create such that 3_tuple == unproject(project(3_tuple)).
With manifolds, the higher dimensions can have a relationship with lower dimensional data, and such a relationship can be encoded into a function. What the research appear to have designed is a system that, given a priori knowledge of task and enough n-tuples for learning, can produce and approximation of such an unproject function.
Thus, after learning, they have a system where, unproject(project(n_tuple)) ~= n+1_tuple. Because they were able to inform the learning system about the nature of the relationship between the two dimensions.
It would be trivial to make any ML model satisfy dimensional homogeneity. Just use only dimensionless variables consistent with the Buckingham Pi theorem. Other symmetries would probably have to be baked in from the start.
As I recall some engineers are developing ML-like models that satisfy all sorts of physical constraints under the name "model order reduction".
"Now, researchers have delivered, with a new theoretical framework for building neural networks that can learn patterns on any kind of geometric surface. These “gauge-equivariant convolutional neural networks,” or gauge CNNs, developed at the University of Amsterdam and Qualcomm AI Research by Taco Cohen, Maurice Weiler, Berkay Kicanaoglu and Max Welling, can detect patterns not only in 2D arrays of pixels, but also on spheres and asymmetrically curved objects. “This framework is a fairly definitive answer to this problem of deep learning on curved surfaces,” Welling said."
"A Mathematical Theory of Deep ConvolutionalNeural Networks for Feature Extraction":
https://arxiv.org/pdf/1512.06293.pdf
"Understanding Convolutional Neural Networks with A Mathematical Model":
A) Studying an existing technique with math.
B) Coming up with a new technique.
You could get a modern engineering consultancy to review your steam engine, but it would still be a steam engine.
- Huh ?
- I build websites types on an air keyboard.
I write manuals for computers.
"So, like, for how to use them?"
No, the computer reads it so it knows what to do.
Ordinary convnets are a way of building in translational symmetry, which is the group R^2 (in the plane). The work being described extends this to larger symmetry groups, such as rotations of a molecule in 3D (which is SO(3)).
For either of these, you can work in Fourier space instead of real space, where convolutions become products. For ordinary convnets means ordinary FFT, but nobody does that as translating to neighbouring pixels is simple enough. Rotations aren't so simple, and so working in Fourier space can be an efficient way to do things. And the connection to physics is really just that the representation theory of SO(3) is a bread-and-butter exercise there, the basis of atomic theory.
It seems to be a way to lessen inductive bias by making decisions about available ML algos. That is, it vastly increases the solution space but remains effective by omitting unlikely solutions.
This is how distributed training often works, for example. Data parallelism.
I don’t understand why it still works in higher dimensions, but it seems to.
The intuition is that the multiple models are “spinning” around the true solution, so averaging gives the final result more quickly. But it works even early in the training process.
What works is averaging similar networks and averaging your networks a lot of times.
w2tanh(w1x) = -1w2tanh(-w1*x)
But if you average the weights in those two equivalent models you get 0.
If you're talking about asynchronous data parallelism, then there can be some averaging of weights, but they all start with the same weights and are re-synched often enough that weights are never too different to break it.
We average the weights themselves, and the efficiency seems to be similar to gradient gathering.
It’s also averaging in slices, not the full model. There’s never a full resync.
SWA is the theoretical basis for why it works, I think.
Another way of thinking about it: If the gradients can be averaged, then so can the weights.
For inference, I dont think there are many papers that claim direct average of weights perform better than any single model. It is usually the output that is accumulated in some way.
"Averaging Weights Leads to Wider Optima and Better Generalization"
Weirdly, the averaging doesn't have to be synchronous.
For ensembling, the mathematical justification for why this surprising result is true (e.g just averaging many weak but different models gives a better model) is pretty interesting: https://en.wikipedia.org/wiki/Condorcet%27s_jury_theorem
I always thought the preferred method was to average the gradient updates, and pass that to update the single mother-model.
http://jmlr.csail.mit.edu/papers/volume17/16-137/16-137.pdf
The math there is above my paygrade, but it describes a way to structurally combine a certain class of ML models for very significant performance gains. More importantly, it describes why it works.
I think this is one of the least strange things in AI. All you're doing is taking N overfitted models (unlikely to be overfit in the same way) and then asserting that the average of those predictions is probably not overfitted as much (regularization). Overfitting as a concept is not restricted to some number of dimensions.
It doesn't really matter though, because the parent comment is nonsense and you can't just average the weights at the end and get a working neural net.
If you think about what sort of brain you'd get if you "averaged" a few hundred geniuses' brains by blending them into a soup and adding some gelatin to a brain-sized sample, yeah, that's about the level of intelligence I saw in the "average neural net." It was basically in a constant state of seizure.
Neural networks already see in "higher dimensions" (whatever that means). Anyone who's ever used neural networks already knows each neuron's branch (i.e. dendrite) of an N-sized vector can already be though of as a "dimension" of a data set. CNN (convolutions) flatten that data (reduce it or seeing the same pattern over less "dendrites", much like PCA, etc.).
CNNs only make sense when working with image data anyways.
Not true, CNNs are used for audio and text as well.
I don't think the title is clickbait, you may be misinterpreting it. Its referring to using CNNs on higher dimensional inputs, not that the layer has multiple dimensions (which has been done since the creation of convnets)
Not true, N-dimensional convnets, 1-d convnets (for NLP and time series analysis), spatially sparse convnets, graph and non-Euclidean space convnets, ... exist and are used.
CNNs are akin to multiscale wavelet transforms. They can be applied on different spaces (just as graph wavelet transforms exist).
Edit: I should say it is a big step in the direction of my prediction.
Figured once there it would be pretty trivial to search for my name.
anyway linked above you, so know you have something else to lol about
>My big prediction is in physics: Thanks to Einstein we live in a 4D reality, 3 spatial dimensions and the dimension of time.
In the next 10 years I predict our understanding of physics will evolve (with a confirmation through observation) from us being 3D beings living in a 4D reality to...something more.