Myths in Machine Learning Research
crazyoscarchang.github.io
crazyoscarchang.github.io
Even more damning is the recent BagNet paper (nice summary here: https://blog.evjang.com/2019/02/bagnet.html), which indicates that ImageNet can likely be solved with no global features (i.e. model doesn't have to learn anything truly abstract, just configurations of textures, shapes, colors). I thought the author of that blog post put it nicely:
"As someone who is deeply interested in AGI, I find ImageNet much less interesting now, precisely because it can be solved with models that have little global understanding of images."
For years, Middlebury was what you tested on, and for years that's what got you published. Nowadays Middlebury is viewed as solved by the top algorithms. If you try those algorithms on your own data, good luck getting similar performance; at least I've not seen any kind of advantage in using anything other than SGM (outside of specific research contexts like my PhD).
I'm more concerned that everyone is using KITTI as a (often the only) benchmark for deep-learning based stereo matchers, since those are all images of roads. At least with classical stereo you have some idea what* your cost function is. The other one people are increasingly using is Scene Flow, which is (entirely?) synthetic. Not a great situation.
* KITTI is a widely used dataset of driving imagery
As I read #3, the hypothesis doesn't depend on individual researchers acting unethically by validating against the test set. Instead, I read it as an analogy to significance bias in other sciences: machine learning models that don't perform as well on the validation set simply aren't published, so the field as a whole over-fits, as if validation is performed on the test set.
In the paper you link, the authors themselves do explicitly note some test-set shenanigans:
> But this assumption [that the models are independent of the test set] is undermined by the common practice of tuning model hyperparameters directly on the test set, which introduces dependencies between the model $\ˆf$ and the test set S. In the extreme case, this can be seen as training directly on the test set
However, it absolutely does not answer the question that is asked in the title of the paper, and the process they use is incapable of answering that question.
If you go back and read the original CIFAR10 paper, you'll see that the process they carefully went through meant that they curated the most suitable images for each category. By definition, what's left over (which is what the Recht et al paper chose from) is less good images, which are of course therefore harder to classify.
All the experiment measures is how good they are at matching the distribution of the original dataset. The answer, they discovered, is: not very.
What #3 has described is a totally normal workflow:
1. Conceptualize a new model idea
2. Implementing and training is totally legal without involving test set.
3. However, once finished, the model is evaluated on test set, the performance of which will decide whether this idea is worthy of publishing or not. If not, they go to the first step and repeat.
Such loop essentially makes the test, the actual validation set, if you think human as his/her own optimizer, and he/she takes a look at the test set periodically, and decide whether to pursue the current idea or not. Sounds like early stopping, isn't it?
Remember back in 2015, there is a debacle from Baidu, where a researcher had fabricated multiple accounts to run unlimited tests against ImageNet's own reserved test set, which the competition straightly forbade.
If a test set is a 'true' test set, then it should work like a test in a real world: be kept secret before revealing to the public, once evaluated, the same problems/examples shall never appear in the later tests ever. But such approach would not be accepted because the cost is simply too high.
In the Recht et al. study, the reason the new test accuracy is wildly outside of a binomial confidence interval around the original test set accuracy is that the distribution is different. The CI only applies to data drawn from the same distribution.
ML research still suffers from replication issues; such is the nature of the scientific incentive structure. However, these issues generally come in the form of poorly tuned baselines, buggy code, and claims with insufficient experimental/theoretical justification. Outside of some isolated cases, publication bias and cheating at hyperparameter tuning do not seem to be major factors.
----
[1] Statistically speaking, to compare two models on the same dataset, one does not care about the accuracy numbers but instead about the number of examples model A gets right that model B does not and vice versa; see McNemar's test.
Also the footnote on the first page cracks me up: "Authors ordered alphabetically. Ben did none of the work."
Browsing through papers from a few years ago some of these myths in 2011 might have been:
1. You need to pre-train large networks so that they converge
2. You need GPUs to train deep networks efficiently
#1 did indeed turn out to be false; pre-training has pretty convincingly been debunked. But if anything #2 turned out very right; there's no serious training happening today on CPUs.
A bit of conjecture here but I suspect word embeddings are going to turn to be the next big thing that turns out not to be all that useful.
Transfer learning using the first few layers pertained on imagenet or a related task have consistently given 1-2% improvements in scores... As recently as mid 2018.
This is especially for complex tasks like VQA
I also think the "do not trust saliency maps" is too strongly worded. The authors of that paper used adversarial techniques to change the saliency maps. Not just random noise or slight variation, but carefully crafted noise to attack saliency feature importance maps.
> For example, while it would be nice to have a CNN identify a spot on an MRI image as a malignant cancer-causing tumor, these results should not be trusted if they are based on fragile interpretation methods.
Interpretation methods are as fragile as the deep learning model itself, which is susceptible to adversarial images too. If you allow for scenario's with adversarial images, not only should you not trust the interpretation methods, but also the predictions themselves, destroying any pragmatic value left. It is hard to imagine a realistic threat scenario where MRI's are altered by an adversary, _before_ they are fed into a CNN. When such a scenario is realistic, all bets are off. It is much like blaming Google Chrome exposing passwords during an evil maid attack (when someone has access to your computer, they can do all sorts of nasty stuff, it is nearly impossible to guard against this). [3]
[1] https://www.technologyreview.com/s/538111/why-and-how-baidu-...
[3] https://www.theguardian.com/technology/2013/aug/07/google-ch...
EDIT: meta(I liked the article. I do not want to argue it is wrong. It is difficult for me to start a thread without finding the one or two things to nitpick at, or to expand upon a point, but this article was already very resourceful)
I’ve seen tons of papers doing that and getting published, especially on cifar10. Not saying it’s a good practice, just that it’s fairly common.
To a certain extent, saliency maps can be perturbed even with random noise, but the more dramatic attacks (and certainly the targeted attacks, in which we move the saliency map from one region of the image to a specified another region of the image) require carefully-crafted adversarial perturbations.
what about when people in the hospital who have a patient that they suspect has cancer use the best machine to create that patients scans and tend to push patients that they think are ok to the older less good instrument? Or if they choose to utilise time on the best instrument for children?
What about when the MRI's done at night are done by one technician who uses a slightly different process from the technicians who created the MRI data set?
At the very least there is a significant risk of systematic error being introduced by these kind of bias, and as you say, it's really hard to guard against this, but if a classifier that I produce is used where this happens and people die... Well, whatever I feel I would be responsible.
Edit: After some wikipediaing, "bases" might be a better word than "units."
Additionally, you've sort of got it backwards. A matrix with units (and a set of basis vectors) attached is one representation of a rank (1, 1) tensor. But it's not really a unique representation of the tensor - you could choose a different set of basis vectors and come up with a different matrix representation of the exact same tensor. The tensor is an entity, while the matrix is a representation of an entity within a given coordinate system.
In physics, however, a tensor has a more specific meaning. In this context, certain 2-dimensional tensors can be represented as matrices, but a matrix is a distinct concept. A bit more precisely, in physics a tensor is an object that transforms a particular way during coordinate transformations. Intuitively this means that a tensor must be some physical "thing".
A classical example of a tensor is the moment of inertia tensor. Every 3-d object has a moment of inertia tensor. This tells you how the torque relates to angular acceleration, and it will in general be different across different axes of the object. Now, you can choose any three (non-collinear) directions you want and write down a matrix which represents the tensor in that basis, but this representation is fundamentally coordinate dependent. The moment of inertia tensor, by contrast is a coordinate-independent entity. Just like a vector, it will have certain values in certain reference frames, but the vector itself transcends any coordinate system. (Though this is a bit of tautology since a vector is a 1-dimensional tensor.)
For those interested, the first chapter of Kip Thorne's book has a good, though idiosyncratic, explanation of tensors: http://www.pmaweb.caltech.edu/Courses/ph136/yr2012/1201.1.K....
Consider you have a vector and a bunch (let's say q) of matrices, and you take the matrix product of the vector with all those matrices. You will get q vectors, which you can stick together to form a matrix. This act of multiplying a vector by a bunch of matrices is clearly linear with respect to the input: If we multiply the input by X, every vector will be multiplied by X, so the resulting matrix will be multiplied by X. Suppose you do this operation, T on v to get vectors T1(v), T2(v) ... Tq(v) and on w to get T1(w), T2(w), ... Tq(w). Since all T's are matrix products (linear) then if we do the operation on v + w we will get T1(v) + T1(w), T2(v) + T2(w) ... Tq(v) + Tq(w). Which is essentially T(v) + T(w). So now we know T is linear with respect to the input. Now, this all took a long time to describe, so let's simplify it: How about instead of a group of matrices, we just call this thing a 3-d tensor? we can let i and j index the regular matrix dimensions and make up a new dimension for the matrix we're on, call it k. Now at any coordinate we get a value so it's basically a 3-d array, but it represents something much more specific than that. You can guess how this might generalize to mapping matrices x matrices to 3d tensors or 3d matrices x 3d tensors to 4d tensors and so on.
So now the question is, does TensorFlow conflate these? I think it does - somewhat. A convolution can be viewed as a tensor (a single filter maps matrices (images) x 3d-tensor (kernel) to matrices (another image)) so I'd call that a Tensor operation. But consider the input image itself. Is this truly a tensor? If we consider a simple situation, say we have some data vector and we're doing a matrix multiply to get the output of a linear model. Is the input a matrix? I would say no, because we don't think of it as acting on the model, we thing of the model as acting on it, even though what we are doing is really equivalent to multiplying two matrices. Equivalently, I would not call the input image, or any activation in a neural network a true tensor, even though it is numerically equivalent. There are true tensors in TensorFlow, but if you're using high level functions (dense, conv2d) they are usually hidden from the user.
No offence, but that's a hideous definition :)
For me a (real) tensor is a function that takes an ordered set of N row vectors and M column vectors as arguments, and spits back a real number as a result. It has to be linear in its arguments. That's all folks!
By this token a matrix A is a tensor: it takes one row vector x, and one column vector y, and returns a real number xAy.
Similarly, a row vector x is a tensor: feed it a column vector y and you get the real number xy.
You can dress all this up in the language of linear functionals or n-forms, but at core that's what's going on.
y_i = A_ij x_j.
If I understand correctly, the paper points out that an algorithm for computing derivatives based on this notation is faster for taking higher-order derivatives compared to using TensorFlow.
Mathematically, tensors are more complicated objects. Basically, they are what you get when you take higher-order derivatives of a function. In particular, the first-order derivative of a function f: R^n -> R^m at a point x \in R^n is the best linear function A_x \in R^{m X n} that approximates the original function, i.e.,
f(x + dx) ~= f(x) + A_x dx.
A linear function is represented by a matrix, so a first-order derivative is a matrix. If I take the second-order derivative, I get a more complicated object B_x, which represents the quadratic term in the Taylor expansion:
f(x + dx) ~= f(x) + A_x dx + B_x(dx, dx)
where B_x(a, b) is a linear function (or more precisely, a "multilinear" function) of two vectors a, b (which are the same in the above formula). That is, whereas A_x is a (linear) function R^n -> R^m, B_x is a (multilinear) function R^n X R^n -> R^m. This mathematical object B_x is an example of a tensor. In R^n and R^m, tensors are pretty boring, but they become more interesting when dealing with functions on manifolds.
The question, then, is: a) whether the space of problems where you have good algebraic notation lines up well with the total scope of TF problems, and b) whether the extra complexity of supporting the full computer algebra system is 'worth it.'
For the latter, keep in mind that algebraic derivatives can get cumbersome/expensive when you have an exponentially complex piecewise linear space (eg: https://arxiv.org/pdf/1711.02114.pdf); the linked paper makes no mention of ReLUs... things might be fine with sigmoid activations, but they're the exception, these days...
I didn't think "Attention > Convolution" was a prevalent myth, given how integral convolutions are to SOTA image classifiers and GANs (if anything, I believe attention is unde-utilised here and due to grow in usage a lot)
Would have been better titled, "Things to avoid when doing ML research"
Damn this hits the nails. People are essentially using 'test' set as validation set, the validation set as the early stopping helper set.
As I understand it, you fit() with training, then do parameter tuning with validation and the best parameter tuned model is used on test.
Now I'm still a little confused as to why we don't just fit() then do hyperparameter tuning with the test set (best-tuned model wins, no need for test). Why would calling predict() on a model cause it to update its weights and overfit?
It would be interesting if someone would see whether they could sneakily (Sokal-style) publish a paper like the following: "We took (popular model X) and augmented it with an additional lexicon of specific lookup data, and the result blows away all the competition. This is deeply profound and implies that built-in lexicons could be the key to true general intelligence!" (When in fact all they did was hard-code the test set or part of the test set into their model.) Then see how many popular presses churn out sensational articles.
Is that backed up by any data?