Improving Deep Learning Performance with AutoAugment
ai.googleblog.com
ai.googleblog.com
For example, you might use it to formulate a set of wavelets that when combined judiciously would effectively span a well-defined distribution of shapes generated from a small grammar. In so doing, you could quantify the shape variance and identify which augmentation transformations added most value for training (minimally modeling that variance) and which added least.
Maybe you could also combine this with t-SNE to gain some intuition of which 'wavelet' manifested where in the trained net, which resonated most, and in concert with which other wavelets. You could explore this across different CNN sizes and designs, looking for evidence of wavelet ensemble or hierarchy.
With some careful engineering, you could try to force emergent autoencoders to reveal themselves and then explore their interactions.
If you prespecify what data augmentation you would do, like preregistering the details of a clinical trial, you’ll be less susceptible to a spurious result from this.
It seems like especially things like color distribution manipulation would have a potentially very adverse effect that counters any gains from clamping the supervised learning to be “robust” to that color variation.
I’m thinking in the spirit of: < https://arxiv.org/abs/1711.11561 >.
That said I can see how if done incorrectly you could easily end up overfitting instead.
I would be quite surprised to learn that, especially for the experimental result they have with the low-pass Fourier filter on the training set, and also because the Bengio paper is quite recent.
An opposite sort of strategy is to apply every augmentation that doesn't increase human error rate on a test sample.