Training computer vision models on random noise instead of real images
unite.ai
unite.ai
The meat of the paper is that the scribbles can be remarkably shitty and still get you decent pretraining.
Compared to alternatives:
imagenet: exhaustively annotated, disk-space-consuming, nonfreeforcommercialuse real world photos - huge and has licensing issues
rotnet: take random images, rotate them and ask the network whether they're right side up or not - easy, but still takes up disk space
3d renders: render realistic 3d scenes on the fly and use those as input - nvidia has enough of my money already, thx
fractaldb: generate fractal images and pretrain by predicting the fractal's parameters. - why bother with fractals if you don't need them
It is basically an experiment how the mid-to-low-level features in NNs generalize from random abstract generated images to real images.
This is not a very surprising result, because it is known NNs generalize fairly well and at these levels you mostly only have blobs and edges that do not look much different from artificial ones.
[1]: https://proceedings.neurips.cc/paper/2021/hash/14f2ebeab937c...
However, this is now the second time this week I have spotted shoddy writing coming from unite.ai – I am referring to the conspiracy-esque spin from [2] that we saw just a few days ago. Apparently they come from the same author even, that seems to put a great deal of emphasis on their articles reaching the Hacker News front page [3]. I am not sure how I feel about this, I would like to believe that “we” as a community are better than to fall for hype and click bait; I am also very uncomfortable with the idea of there being professional prestige in getting onto the front page.
[2]: https://www.unite.ai/a-cartel-of-influential-datasets-are-do...
ANNs generalise atrociously badly, hence the need to train with ever bigger data to cover every nook and cranny of the instance space of each class of interest.
Incidentally, when people say that ANNs "generalise" they mean many different things, for example that they "genealise on the test set" which is usually only observed when the tests set is known in advance (and so it has been used in tuning hyperparameters and the like) or even, incredibly that they "generalise on the training set" (i.e. the validation partitions in cross-validation). Conversely, there is a glut of novel terminology like "out-of-sample" or "out-of-distribution" to describe generalisation beyond the test set, but this kind of generalisation is typically held up as a weakness of ANN, because they're genreally really bad at it.
In any case, strong evidence of robust generalisation on out-of-sample data from few examples and with no or little pre-training, in ANNs, would be a surprise, indeed.
The difference between interpolation and extrapolation is almost the most important concept in all of machine learning practice.
It's infinitesimally rare (from what I've seen so far) that a practical machine learning model can perform high quality extrapolation, for many different metrics of quality.
There's almost always far, far too many confounding variable.
In high dimensional spaces basically everything is extrapolation including in pixel space and embedding space.
> The notion of interpolation and extrapolation is fundamental in various fields from deep learning to function approximation. Interpolation occurs for a sample x whenever this sample falls inside or on the boundary of the given dataset's convex hull. Extrapolation occurs when x falls outside of that convex hull. One fundamental (mis)conception is that state-of-the-art algorithms work so well because of their ability to correctly interpolate training data. A second (mis)conception is that interpolation happens throughout tasks and datasets, in fact, many intuitions and theories rely on that assumption. We empirically and theoretically argue against those two points and demonstrate that on any high-dimensional (>100) dataset, interpolation almost surely never happens. Those results challenge the validity of our current interpolation/extrapolation definition as an indicator of generalization performances.
https://arxiv.org/abs/2110.09485
Also:
> The location of decision boundaries inside the convex hull of training set can be investigated in relation to the training samples. However, our analysis shows that in standard image classification datasets, all testing images are considerably outside that convex hull, in the pixel space, in the wavelet space, and in the internal representations learned by deep networks. Therefore, the performance of a trained model partially depends on how its decision boundaries are extended outside the convex hull of its training data.
Edit: or take photos of brown cats indoors from the front and see if the model recognizes albino cats from the side outside.
And it does not matter what space you are using as long as you operate under the convex hull definition of interpolation vs. extrapolation you will need exponentially more samples as intrinsic dimensionality of the space increases.
This means that even under the manifold hypothesis, as long as intrinsic dimensionality is reasonably high i.e. in low hundreds, models will be doing extrapolation.
The thing is also, any individual hyper-dimensional case can be outside of the training set's convex hull itself and be correctly classified.
However, you would still have to quantify what the relationship of the dimensions with the highest feature importance were to said space. Which is why the second paper is so fascinating.
From the perspective of issues product/engineering teams face in the field, I'd definitely maintain that fire alarms should start sounding once you see any sort of extrapolation and you should dive deeper.
Unfortunately the maturity level of this space is still at the point where peer review of data-set transformations before deploying to production and committing Jupyter Notebooks to GitHub is a heated in-office discussion.
The majority of the commercial world is a long way from that kind of best practice.
That is a big flaw in machine learning research in general, but that's for another conversation, I guess. My point above is that if neural nets could generalise well, they wouldn't need so much data. In a sense, even if trained neural net models can generalise to instances outside the dense region of instance space circumscribed by their training set that is not that important, if that region has to be gigantic for this generalisation to be possbile in the first place. For one thing, at that point it becomes difficult to separate what is "training" and what is "test", especially so when test sets are four times the size of training sets as in typical practice.
This statement is meaningless without controlling for model complexity and data type. For their simplicity, ANNs generalize well on a wide variety of data. GPT-3 yields almost human-level generalization ability for some tasks.
I also clarified that the generalization is probably not far. There is not much complex "realism" to be found in low-to-mid-level features; they're almost mathematical in their simplicity, similar to basis functions.
That's an extravagant claim. There's no machine learning or other system, or algorithm, or technique, that can approach the ability of humans to "generalise" in any task, no matter how you want to define "generalisation". The models built by ANNs in particular are shallow and over-specialised and have none of the depth or complexity of whatever "models" of the world and the entities in the world that humans build in our heads.
Evaluations that show "superhuman" ability are poorly designed. Machine learning research is following benchmarks and metrics that mean nothing and show nothing, beyond the ability to beat said benchmarks with said metrics, which is then blithely taken to mean "progress" towards the approximation of human intelligence. This then leads to hyperbole like in your comment.
... and how you want to define "task". For some prompts/"tasks", GPT-3 does generate impressive (more than trivial) outputs that cannot be found on the internet and that are indistinguishable from what a human would respond, so it generalizes in that sense. Maybe human ingenuity and generalization is also just slightly perturbed interpolation? It is very difficult to produce something truly novel, so we are also rather tightly limited by prior experience. Who knows? Also, who cares if submarines swim? Anyhow, it seems 50% of the internet is bikeshedding about definitions.
So, suppose you say that a particular bit of text generated by GPT-3 is "indistinguishable from what a human would respond". If I say it isn't, how can we decide who is right in a way that we can both agree on?
And that's all before we try to figure out "generalisation".
You simply hand the text to someone and ask to guess if it was produced by a human or not.
Or hand them two (or ten, or ten thousand) texts and ask them to label the human and AI texts without knowing the actual distribution.
Wednesday was RSU vesting day. Your purchase is appreciated. <3
The abstract seems to say that, but again, I could be misinterpreting: “Our findings show that it is important for the noise to capture certain structural properties of real data but that good performance can be achieved even with processes that are far from realistic.”
edited: Typo,clarity
For example, to classify a hotdog it might be useful to first generate an intermediate representation of the image (think "cylindrical, brown, meaty thing"). Such a representation can then fairly easily be mapped to the concept "hot dog".
These representations can be learned from large image datasets alone (they do not require labels!). In our work we show that you don't even need real images, but that images that are generated from noise processes are enough to train such representations, and that these representations are surprisingly good for classification.
Hope this clarifies things a bit, and happy to answer any other questions!
Back before deep learning, people used to make recognizers for features like that as a lower level of feature recognition. Now it's expected that features will be derived automatically from real imagery. This is kind of a return to that level.
A useful training set might be a big texture library used for game development or animation. Those are easily available.
I would expect it to be somewhere in the ballpark of our StyleGAN images, which also look very "textural", but lack these effects that are an result of imaging the 3D world. Interestingly, modelling these effects without realistic textures seems to result in worse performance - this is for example the case for images taken from CLEVR or generated from Minecraft, and both perform worse than the StyleGAN images.
It is then possible to generate arbitrary amounts of these images as samples from the stochastic process - these images exhibit certain image-like structures (such as oriented edges), but are as a whole still random and extremely varied, which is good and necessary for the representation learning.
In terms of helping, though, it is important to note that we do not achieve state-of-the-art performance yet, and when looking at absolute performance for a task like image classification, using real images is still better. That being said, something that is in the paper but generally seems to get lost is that our representations work very well when analyzing data that is very different from normal images, such as medical images or satellite images.
But collecting lots of image data is hard so someone comes along and points out that the most important characteristic of your images is that they're appear drawn from an autoregressive process... That is that each pixel is (say) 95% correlated to the pixel 1 away, 95%^2 to the pixel two away, 95^3 three away and so on. So to compute the PCA you only need the covariance matrix anyways, so you just generate an AR0.95 covariance matrix and use the transform derived from that. And you find this works pretty well too.
This work is along those likes, they're generating random images with some simple natural-like statistical properties and training the first part of the classifier using them and getting useful results.
One interesting promising part is that this line of thinking may result in better network designs or insights that allow skipping training these initial layers: Going back to the block transform example, the next step would be to notice that the PCA of the AR0.95 matrix is the discrete cosign transformation, which is agreeable to an extremely efficient implementation.
I'm not enjoying this article. Effective at what exactly?
Actual: various detection tasks, as compared to other image networks.
Simple version: Many nn tasks work by feeding a lot of data into a network, then “refining” it with a task you want it to actually do.
So let’s say I want to detect corgis.
I take a network that can detect all kinds of shit, and feed it a bunch of images of dogs and tell it, no, only these ones please.
Why?
…because you don’t have 10TB of dog images, you have 10TB of pictures of random crap, and 1GB of corgi images.
…and fair intuitively, this works. If something knows how to tell dogs from cats from cars, it’s not a big stretch to become more specific and only detect corgis.
Now, this paper is showing that instead of feeding the initial network labelled images of cats, dogs, etc … you just feed it with procedural noise, it still works.
That’s a) surprising (wtf, why does telling the difference between a squiggly and another squiggly help you tel corgis from cats?) …
And b) really important, because it means you don’t need to spend hundreds (thousands) of hours or $$$ collecting the initial datasets.
Practically, what does that mean?
Well, here’s some food for thought: Google / Amazon etc are considered to have a fairly defensible moat for their voice recognition tech, because they are the only ones with enough data to train good models.
…but that moat vanishes in a blinding flash of steam if you can get comparable results from just feeding enough generated noise into a network.
So this is pretty interesting stuff.
https://en.m.wikipedia.org/wiki/Retinal_waves
[edit: I see the authors have already made this connection in the paper]
An image of IID pixels is a very unusual one, and I assume that picture to picture their process is IID.
Another way of looking at it, say you image an image with iid exponentially distributed pixels. The bits in the image file, however, would not be iid. So just because you can point to some part and say it's not iid doesn't make it wrong to call it random, it's just a question of what scale you're operating at.
Similarly, if you made an image that was 1/f instead of totally spectral flat, it wouldn't be IID (looking at the pixels alone, again)-- but I don't think anyone would fail to call a such an image "random noise".
The paper itself seems ok. It's certainly not ultra groundbreaking, but I think the research is useful and presents a pretraining step that could be used in many applications.
https://arxiv.org/abs/2107.10356
> However, our findings that AI can trivially predict self-reported race -- even from corrupted, cropped, and noised medical images -- in a setting where clinical experts cannot, creates an enormous risk for all model deployments in medical imaging: if an AI model secretly used its knowledge of self-reported race to misclassify all Black patients, radiologists would not be able to tell using the same data the model has access to.
Potentially the method in OP's paper could be used as an adversarial critic or similar counterbalancing effort to eliminate "secret knowledge" when it's not desirable.
Somehow this seems a little bit similar. Neural networks have a great many parameters. If training on very abstract images gives good results, then it seems like it indicates that there are certain symmetries that makes the neural network better, irrespective of dataset or task.
If that hypothesis is true, then it may be possible to change the architecture of the network to directly provide those properties without any training at all.
Consider that dimension reduction with the Locality Sensitive Hashing algo, or something like this [], has proven utility. You could make some hand-wavey argument that the approach in the article is similar in extracting features from randomness.
[] https://scikit-learn.org/stable/modules/random_projection.ht...
I expect this result to be not that interesting for most AI programmers, because you'll use a pretrained ResNet preprocessing anyway. But it is a very elegant solution for the licensing issues that come with large unsupervised image collections like ImageNet.
The researchers suggest that the current crop of machine learning architectures may be inferring something far more fundamental (or, at least, unexpected) from images than was previously thought...
Is it more plausible that this shows they are inferring something fundamental, rather than that they are differentiating images on the basis of some of their accidental (i.e. non-essential) features?