Also, it kind of defeats the purpose of neural networks to do substantial feature engineering like that.
Also, it kind of defeats the purpose of neural networks to do substantial feature engineering like that.
Blanking out areas that are "not of interest" I would consider substantial feature engineering (unless the task you were training a net for was explicitly to find interesting vs. uninteresting areas).
Our human eyes have a lot of filters (hardware and software-based) before recognition takes place.
I currently work on estimating emphysema extent in CT lung scans. Emphysema can be very diffuse and it is not possible to label individual pixels, so instead we try to learn the local emphysema pattern from a global label. Neural networks are interesting for this problem because the learn the features, but it is also a "problem" because the features might not make physically sense, which could make it hard to transfer the model and convince clinicians that they should use it.
We should just be realistic. We want to take real image, except it might be tinkered with, and make neural net tell us what we see on it, except we also want it to see what we can't see, and we want it to answer as accurate as possible, except we also want short and definitive answer.
We also kind of want it to admit that image always contains more than one thing, but kind of don't.
You won't win much by making every neural network learn stuff from scratch that can be done once, good.